What should an enterprise CDC proof of concept test?
An enterprise CDC proof of concept should test four things: data correctness, recoverability after failure, impact on source and target systems, and sustainable security and operations. The RFP should define guarantee boundaries and known limitations. The PoC should deliberately introduce network, schema, target, credential, and failover faults instead of proving only that one low-volume table can replicate.
Why does a happy-path CDC demo prove so little?
A feature demo proves that a prepared path can move data. It does not prove that the path can survive the conditions that make production replication difficult: large transactions, uneven change rates, schema changes, a slow target, expiring credentials, network partitions, source failover, or a recovery that starts from an ambiguous checkpoint.
Separate three stages. A feature demo confirms basic compatibility. A technical proof of concept tests defined requirements under controlled stress. A production pilot validates operating ownership, support, change control, and real downstream use. Treating these stages as interchangeable is how a polished demonstration becomes an unplanned production experiment.
A credible PoC uses a reduced but representative version of the intended topology. It includes realistic distributions of inserts, updates, deletes, large transactions, hot keys, wide rows, and schema changes. It also records tail latency, backlog recovery, duplicates, gaps, and source impact. A single average latency figure cannot answer those questions.
What requirements belong in an enterprise CDC RFP?
Start with the workload and operating boundary, not a connector count. The RFP should make every vendor answer the same questions using the same definitions. Require each response to be labelled Supported, Conditional, Planned, or Not supported. A bare yes or no hides version restrictions, topology dependencies, and manual steps.
| Requirement area | Questions the RFP must answer | Required evidence |
|---|---|---|
| Environment | Which source versions, editions, topologies, table types, and targets are in scope? | Supported matrix and architecture diagram |
| Semantics | How are inserts, updates, deletes, transaction boundaries, ordering, keyless tables, and DDL handled? | Documented behavior plus PoC test |
| Service objectives | How are freshness, RPO, RTO, availability, and source impact defined? | Measurement method and limits |
| Operations | Who owns upgrades, monitoring, replay, reconciliation, failover, and 2 a.m. escalation? | RACI and runbook |
| Security | Where do payloads, metadata, telemetry, keys, logs, and support access cross boundaries? | Data-flow and access-control evidence |
| Lifecycle | How are state, history, and downstream contracts migrated or removed at exit? | Migration and decommission plan |
Hard requirements should cover correctness, source safety, recoverability, security, and residency. Preferences such as user-interface style, an existing team skill, or procurement model can be weighted, but they should not compensate for a failed hard gate.
How should teams establish a fair PoC baseline?
Before enabling CDC, record the state of the source and target. Measure source CPU and database time, I/O, transaction latency, log generation, network use, and storage growth. Measure target write capacity and queryability. Without a baseline, a team cannot distinguish CDC overhead from an existing bottleneck.
Freeze the dataset, workload generator, concurrency, test duration, versions, instance sizes, network path, and durability settings. Run at least four workload phases: steady state, expected peak, short burst, and recovery from backlog. Include the largest normal transaction and a deliberately oversized transaction. Record clock synchronization and the exact points at which timestamps are taken.
- Use the same workload and measurement method for every shortlisted option.
- Preserve raw metrics and logs, not only screenshots or summary slides.
- Document every tuning change and rerun the baseline after material changes.
- Report failed runs and configuration limitations alongside successful runs.
How should source impact, latency, and throughput be tested?
Break end-to-end freshness into stages: source commit, capture, processing or queueing, target commit, and target queryability. Capture latency may remain low while a target queue grows for hours. Report p50, p95, and p99 freshness, the age of the oldest outstanding event, steady-state throughput, backlog growth, and backlog drain rate.
Measure source impact after all required logging and identity settings are enabled. Watch source transaction latency, CPU, I/O, log volume, storage headroom, and contention. The acceptable threshold is a business and DBA decision; there is no universal percentage that makes a test pass.
| Test | Measure | Pass condition | Common trap |
|---|---|---|---|
| Steady state | Tail freshness and source overhead | Business SLO and DBA guardrails met | Reporting only an average |
| Peak burst | Backlog growth and saturation | No unsafe source-log pressure | Ending the test before queues grow |
| Slow target | Backpressure and isolation | Defined degradation path works | Monitoring only the connector process |
| Catch-up | Drain rate and time to normal | Recovery objective met | Calling restart time the RTO |
How should snapshot-to-stream handoff be validated?
An initial snapshot and an incremental stream must share a defined consistency boundary. Depending on the source, that boundary may be represented by a log position or another database-specific marker. The PoC must show that transactions around the boundary are neither missed nor applied twice.
Keep writes active while the snapshot runs. Insert, update, and delete boundary records before and after the consistency point. Restart capture, interrupt the network, change a schema, and generate a large transaction during the copy. Determine whether a large-table restart resumes from a checkpoint, reruns one table, or rebuilds the entire snapshot.
Reconcile in layers: row counts, key coverage, key-range checksums, control totals, and selected business invariants. Counts alone can match even when the wrong rows or values arrived.
Which failure, replay, and reconciliation tests are mandatory?
Failure injection turns recovery claims into observed evidence. Terminate the capture process, interrupt the control path, partition the network, stop the target, expire a credential, exhaust a buffer, and perform a source switchover where the intended topology supports it. Inject failures at different points, including immediately before and after a checkpoint or target commit.
| Failure | Record during the test | Verify after recovery |
|---|---|---|
| Capture crash | Detection and restart point | No gap; duplicates handled as designed |
| Target outage | Buffer growth and source-log headroom | Controlled catch-up and final reconciliation |
| Network partition | Retry behavior and alert timing | Checkpoint continuity and ordering scope |
| Credential expiry | Error classification and ownership | Rotation procedure and audit trail |
| Source failover | New source position and manual steps | Continuity or documented resnapshot path |
| Disk pressure | Throttle, pause, or shutdown behavior | Safe recovery without silent loss |
For every run, retain the claimed behavior, test steps, observed result, artifacts, operator, open risk, actual RPO, and actual time to restored freshness. Distinguish replaying retained events from backfilling historical state. A target with external side effects may require a different replay design from a database that supports idempotent upserts.
How should schema change and data quality be scored?
Do not score “DDL support” as one checkbox. Test detection, representation, propagation, application, approval, and rollback separately. Include added, removed, renamed, and type-changed columns; key changes; partition operations; incompatible changes; and at least one unsupported type.
Data-quality tests should cover nulls, numeric precision, time zones, character encoding, large objects, truncation, deletes, duplicates, out-of-order events, and transaction boundaries. Record whether a problematic change is synchronized, paused, ignored, quarantined, or requires manual action. The correct behavior is policy-dependent; silent behavior is the risk.
What security, residency, and audit evidence should a vendor provide?
Ask for a complete flow of business data, schema metadata, telemetry, logs, credentials, keys, and support traffic. Identify the component, location, direction, protocol, retention, owner, and whether the path can be disabled. A compliance badge cannot replace a product-level architecture review.
- Verify least-privilege source and target accounts, role separation, and account revocation.
- Rotate a secret or certificate during the PoC and observe continuity and audit events.
- Export audit records and confirm they identify the actor, change, time, and result.
- Test the intended private-network or disconnected operating mode where applicable.
- Document remote-support approval, duration, visibility, and removal.
How should a weighted scorecard and exit criteria work?
Set weights before testing. A useful starting structure is 30 percent correctness and recovery, 20 percent source impact and performance, 15 percent operations, 15 percent security and governance, 10 percent deployment fit and limitations, and 10 percent commercial ownership. Change the weights to reflect the workload, but never let weighted convenience offset a failed hard gate.
Use three outcomes: Go, when all gates pass and risks are accepted; Go with conditions, when named owners can close bounded issues before production; and No-go, when correctness, source safety, recoverability, security, or residency remains unproven. Every condition needs an owner, due date, evidence requirement, and consequence if it remains open.
How should operator readiness be tested?
A platform can pass engineering tests and still fail its operating handoff. During the first run, let the vendor or implementation team establish a correct reference configuration. During the second run, remove that dependency: the customer on-call team should detect a seeded incident, identify the failure domain, execute the approved runbook, restore service, and produce the evidence needed to close the event.
Score the documentation and operating surface, not only the outcome. Operators should be able to find the affected pipeline, source and target positions, oldest outstanding event, recent configuration changes, credential state, queue headroom, and the action that last changed the system. Diagnostic exports should be usable without granting unnecessary access or sending sensitive payloads outside the approved boundary.
| Operator test | Evidence | Failure signal |
|---|---|---|
| Detect a seeded outage | Alert timestamp and affected consumer | Only a generic process alarm appears |
| Identify the failure domain | Source, path, target, or identity diagnosis | Recovery begins by trial and error |
| Execute recovery | Versioned runbook and command log | Undocumented vendor-only action is required |
| Verify data | Freshness and reconciliation evidence | Incident closes when the process turns green |
| Explain the incident | Timeline, owner, root cause, and follow-up | No durable audit trail exists |
Include routine work as well: rotate a credential, add a table, approve a schema change, deploy an upgrade in non-production, roll it back, and export audit evidence. Record the number of teams and privileged roles required. A system that works only when its original designer is present is not operationally ready.
Test communication as part of the incident. The alert should identify the affected business service and current freshness, while the incident channel should preserve technical detail for responders. Define who may pause capture, extend retention, fail over a source, replay a range, or accept a temporary SLO breach. If authority is unclear, operators will either wait too long or take a high-risk action without the right approval.
What changes between a PoC and a production pilot?
The PoC proves bounded technical claims. The production pilot proves that those claims survive real ownership, change control, downstream use, and support boundaries. Before promotion, replace temporary accounts and permissive network rules, connect enterprise identity and secrets, define maintenance and escalation windows, test backup or recovery of platform state, and confirm monitoring retention.
Carry every condition from the scorecard into a production-readiness register. For each item, name the control, owner, due date, evidence, and consequence if it remains open. Re-run tests affected by production topology: regional routing, source standby or cluster behavior, target quotas, encryption keys, firewall rules, and the final support-access path. Performance results from a single-node laboratory should not be treated as evidence for a redundant deployment.
Promotion should require sign-off from the source owner, target owner, platform operations, security, and the business data owner. The decision packet should include the frozen requirements, configurations, raw test results, unresolved risks, runbooks, SLOs, rollback plan, and the exact capability status at approval time. This prevents a conditional PoC result from becoming an unconditional production claim.
Fit, non-fit, and the next step
This method fits enterprise CDC programs with persistent workloads, meaningful recovery obligations, regulated boundaries, or several downstream consumers. It is intentionally heavier than a connectivity check. A short demo may be sufficient for a disposable experiment or one-time transfer, but it should not be presented as production proof.
If Deltaplex is part of the shortlist, apply the same workload, failure plan, evidence log, and hard gates used for every other option. The next step is to convene the source DBA, target owner, platform SRE, security reviewer, and business data owner; freeze the scorecard; and run the first baseline before any tuning begins.
Start with the enterprise CDC evaluation guide, then use the CDC architecture guide and schema evolution guide to define the test boundary.
Frequently asked questions
How long should an enterprise CDC proof of concept run?
Long enough to complete baseline, steady-state, peak, failure, recovery, reconciliation, and independent operator runs. Calendar length matters less than completing the agreed workload phases and evidence set. A complex estate may need several weeks; a bounded workload may need less. Do not shorten the exercise by omitting the target-outage, source-failover, or backlog-catch-up phases. A short feature demo that proves only connectivity is not a production PoC. Run the critical recovery procedure once with vendor guidance and again with the customer team following the proposed documentation. Include enough continuous runtime to observe log growth, scheduled jobs, credential behavior, monitoring noise, and at least one planned change window. Freeze the exit criteria before testing begins so the calendar cannot be used to declare success while hard evidence remains incomplete.
Which failure tests should every CDC PoC include?
At minimum, test a capture-process crash, target outage, network interruption, credential failure, buffer or storage pressure, and the source failover path intended for production. Trigger failures near checkpoint and target-commit boundaries, not only while the pipeline is idle. Record detection, ownership, manual steps, restart, backlog catch-up, actual RPO and RTO, duplicate or missing effects, and final reconciliation. A recovered process without a reconciled target is incomplete evidence. Also test an outage long enough to expose buffer growth and source-log retention risk, not only a brief restart.
How should CDC latency be measured fairly?
Measure from a defined source commit or business-event time to the point at which the target result is committed and queryable. Report p50, p95, p99, oldest-event age, and backlog drain time rather than one average. Use the same dataset, workload, hardware class, durability settings, transformations, network path, clock method, and target commit behavior for every option. Publish the measurement points with the result so reviewers know what the number includes. Separate successful and failed operations, and document clock synchronization or heartbeat methods used to control timestamp error.
What source-system impact is acceptable for CDC?
There is no universal acceptable percentage. Establish a pre-CDC baseline, enable every required logging and identity setting, and let the source owner set workload-specific guardrails for transaction latency, CPU, I/O, locks or waits, log generation, storage headroom, and operational risk. Measure steady state, peak, large transactions, and target slowdown. A connector that uses little host CPU can still cause unacceptable database or log-retention pressure. The DBA should own the guardrails and the authority to stop a run before the source enters an unsafe state.
How do teams test exactly-once claims?
First define the claim’s start point, end point, ordering scope, failure assumptions, and sink behavior. Then inject crashes around source-position persistence, publish acknowledgements, target commits, and pipeline checkpoints. After each recovery, search for missing keys, duplicate effects, partial transactions, stale versions, and order violations. Repeat the test because timing-sensitive windows may not appear once. A connector-level label does not prove an end-to-end result at an external sink. Include a side-effecting consumer only when it has an explicit idempotency or deduplication contract; otherwise narrow the claim.
Should pricing be part of a technical PoC scorecard?
Commercial fit belongs in the final decision, but it should not distort the technical evidence. Keep correctness, source safety, recoverability, and security as separate hard gates. After those are assessed, add a comparable lifecycle-cost model covering software, infrastructure, storage, network, non-production environments, platform and DBA labor, on-call, support, migration, dual run, rollback, and exit. Use formal quotes and explicit assumptions rather than unverified public prices. Record how each option meters usage and how growth, regional redundancy, or longer retention changes the model.
