Works above the source technology
Because comparison occurs on the records returned to the pipeline, the same approach can be used across APIs, databases, files, and application data.
Identify meaningful inserts, updates, and deletions from business data even when the source cannot provide a native change stream or transaction log.
Native CDC is powerful, but it is not always available. SaaS APIs, files, reports, legacy applications, and restricted databases may expose only the current state. Synthetic CDC compares that state with the last known state and emits the logical changes that matter downstream.
Because comparison occurs on the records returned to the pipeline, the same approach can be used across APIs, databases, files, and application data.
One or more key columns identify a unique business record. DataZen uses those identities to determine whether a record is new, changed, unchanged, or deleted.
Volatile values such as retrieval timestamps can be excluded from change identification so they do not create false updates on every execution.
When changes exist, the capture result is stored with schema and execution metadata so downstream readers can process and replay it consistently.
The first keyed capture has no previous state to compare against, so the returned records form the initial capture. A reinitialization intentionally resets that baseline. Without keys, capture does not perform the differential and the complete retrieved data set is captured.
A high watermark reduces what must be requested from a cooperative source. Synthetic CDC determines what actually changed within the returned records. They solve different problems and can be used together to reduce source reads and downstream processing.
Synthetic CDC turns snapshots into a governed stream of business changes, independent of native source CDC support.
Back to DataZen