Tags: Altinity/clickhouse-sink-connector
Tags
fix: history-mode DELETE lost -- _operation dropped from INSERT, and … …the delete races its own batch (#1411) * fix: _operation dropped from history-mode INSERTs; over-eager missing-column error Two defects in replication-history (bitemporal SCD Type 2) mode, both reproducing on released 2.10.0 and on the 1410 image. 1. _operation never reached ClickHouse. createColumns() drops any target column the incoming change event does not carry, unless the connector populates that column itself. The connector-managed list -- isConnectorManagedColumn -- named _version, is_deleted, _sign, _valid_to and _valid_from, but not _operation. No source record ever carries _operation, so it was filtered out of every history-mode INSERT. The result is silent: the SCD Type 2 rows land with an empty _operation, so a DELETE cannot be told apart from an insert. Row counts stay plausible, and a count-based checksum reports the table clean. Proven failing-first: on unmodified develop the generated statement is INSERT INTO `process`(`id`,`appkey`,`_valid_from`,`_valid_to`, `_version`,`is_deleted`) with no _operation at all. 2. The missing-column error fired for columns that were never missing. The error added in #1410 could not distinguish a column the record does not carry -- where omission is deliberate and ClickHouse applies the DEFAULT -- from a column the record does carry but that has no placeholder, which is a real dropped value. A history-mode run emitted 15 errors for a `comment` column that was simply NULL at source. recordCarries() now separates them, so the error only fires when a value is genuinely lost. Verified end to end on an isolated MySQL 8.0.36 -> connector -> ClickHouse 24.8.14 stack, target pre-created exactly as the production host has it. After the fix _operation is populated (C/U/D), the tombstone rows appear (is_deleted=1), and MySQL and ClickHouse agree value-for-value at 13 rows where ClickHouse previously held 15. sink-connector: 226 run, 0 failures, 20 errors sink-connector-lightweight: 223 run, 0 failures, 1 error Error counts unchanged from unmodified develop measured in a pristine worktree; all are testcontainers "Could not find a valid Docker environment" on this build host. Known-remaining, not addressed here: the delete is timing-sensitive. Across repeated runs of the same workload the tombstone was written in one run and not in two others, with no error logged either way. That is a separate scheduling/flush defect in the batch path, not the column-list defect fixed above, and it needs its own investigation rather than being bundled in. * fix: history-mode DELETE races the batch it depends on and is silently lost The SCD Type 2 delete reads the row it is closing straight back out of the target: INSERT INTO <table>(...) SELECT ... FROM <table> FINAL WHERE <pk>=? AND `_valid_to` = <sentinel> AND `is_deleted` = 0 UNION ALL SELECT ... (the tombstone) and it runs INLINE, through its own connection. Plain inserts in the same batch are only STAGED on the PreparedStatement; they do not reach ClickHouse until executeBatch() at the end of the batch. So when one batch carries a row's CREATE and its DELETE -- ordinary for any busy source -- the delete's SELECT executes BEFORE the insert it depends on. It matches nothing, both branches of the UNION produce no rows, and the statement succeeds having written zero rows. Nothing is logged: the delete is dropped and the row stays visible in ClickHouse forever. This is the production zero-row-INSERT signature: on the affected host 77-92% of INSERTs on the busy table wrote zero rows every day, while a sibling table in the same database sat at 0%. The fix flushes the staged batch before the inline delete, so the delete observes the same state the source did at that binlog position. It is the same ordering rule the TRUNCATE branch a few lines above already applies, for the same reason. Why it looked intermittent: it only bites when the CREATE and the DELETE land in the same batch. Repeating one workload against a clean target: released 2.10.0-lt 0/3 pass (ch=0, total loss -- Code:57 loop) this branch minus this fix 0/4 pass (ch=15 vs mysql=13, no tombstones) this branch 5/5 pass (ch=13 = mysql=13, 2 tombstones) Standard mode re-verified for regressions on the same build: 17/17 tables at 20,002 rows, plus kill -9 mid-stream with writes during the outage replaying exactly once (20,002 rows, 20,002 distinct ids, zero duplicates). sink-connector: 226 run, 0 failures, 20 errors sink-connector-lightweight: 223 run, 0 failures, 1 error Error counts unchanged from unmodified develop measured in a pristine worktree; all are testcontainers "Could not find a valid Docker environment" on this build host. --------- Co-authored-by: OmniWatcher <omniwatcher@jumptrading.com>
PreviousNext