Rank a capture's packs by their first publish so a replayed pack cannot flip a pinned read - #161
Merged
Merged
Conversation
A pass that re-indexes an already-published pack (crash before commit_packs, an outcome-unknown publish that landed, a conflict whose commit_packs failed, a rebuild beside the live indexer) writes that pack's descriptor rows at a fresh, higher version before publishing anything. The reader ranked a capture's candidate rows by the row's own index_version, and membership bounds packs rather than rows, so those rows outranked a newer pack describing the same capture inside snapshots that were already pinned. If the replay never published, that stayed true. Found by TLA+ model checking (VersionPublish Replay/ReplayCrash/ ReplayCrashPinned). The Python and C++ readers now join the snapshot's packs with the newest version their publish reached the watermark at (clickhouse_sql. member_versions, the same paired-manifest rows as membership_predicate) and resolve on (member_version, store_id, pack_id, index_version). The public view's DDL is unchanged. The C++ indexer also gets Python's guard on the conflict path: a commit_packs failure no longer replaces kPublishConflict. On top of the two-phase search page, both of its queries now read the snapshot join: the inner key query and the outer argMax query take the same FROM and the same caller filters, and the inner one adds WHERE only when there are filters. Membership left `clauses`, so an inner query kept on the raw table would either render "WHERE GROUP BY" when unfiltered or pick page keys from unpublished and replayed rows and end a walk early.
The reader ranked a capture's packs on the newest paired publish at or below the pin, max(index_version) over the manifest. A replay that does publish moves that rank: pass A publishes P1 at v1 and dies before commit_packs, pass B publishes P2 at v2 describing the same capture, and a later pass that re-indexes P1 publishes it again at v3. From v3 on every head resolved the capture back to P1, the older pack, on the normal crash-recovery path. Pins stayed correct; heads did not. Both readers now take min(index_version): a pack's rank is fixed once it is first published, so a replay, published or not, cannot move it. Pins are stable either way, since a later publish lands above the pin. A genuine re-capture is a new pack and a mirror is another store, so both still get a fresh first publish. The Python and C++ member_versions subqueries change together and stay identical. The live replay test and the native/Python parity test now publish the replay and require P2 at the new head, before and after a merge; both failed with max. EXPLAIN indexes=1 plans for the probe corpus (2M rows, 20k packs) are identical under min and max, as are rows read, and timings are within run-to-run noise.
NativeIndexer::commit() records a conflicted publish's packs in the inventory and rethrows kPublishConflict, and when that commit_packs fails the conflict still wins (catalog.py's `raise conflict from commit_failure`). Only the Python oracle's test covered it: the conformance driver had no way to produce a same-version conflict, so removing the guard passed every native suite. The `index` op now takes conflict_at_publish: a foreign watermark row lands at the pass's version after the pass's own and before its owners read-back, from the request hook RequestDeadline runs before each request, which is the hook the storage service's lease scope uses. after_conflict="transport" then fails the inventory INSERT that follows. The live test checks both: a plain conflict records the packs, and a failed record still reports SnapshotPublishConflictError with the failure in its message. With the guard reverted to main's shape the transport case reports ClickHouseError and fails.
…ts packs The conflict-path guard in NativeIndexer::commit() rethrew any commit_packs failure as kPublishConflict. Since #159 that INSERT runs under index_bounded's second LeaseScope, whose before-request hook renews the lease once a renewal is due. A rival holding the lease, which is the situation a second writer creates, refuses that renewal with kHeld, and the guard relabelled the refusal as a conflict. index_bounded rethrows only a lease refusal (is_lease_refusal), so it treated the lost lease as an ordinary failed batch. When a reconcile was due, run_cycle went on to it under no lease: it read a page's missing packs from the object store before commit() refused with kLease, and those packs inflated index_failures. If this happened on the last page of the start-time reconcile that had missing packs, reconcile_owed_ stayed false. The conflicted batch then stayed out of the inventory until something reconciled again, which with the default interval of 0 is the next start. Readers were unaffected, because the packs are members. A kHeld or kLease from commit_packs now keeps its kind and carries the conflict in its message. Any other failure still becomes kPublishConflict. is_lease_refusal moves from storage_service.cpp's anonymous namespace to lease_coordinator.h, so the service and the indexer share one definition. The Python oracle's commit_packs never renews, so it has no such case and does not change. The conformance `index` op's after_conflict="lease_refused" plants a rival claim at the lease's term and renews before the inventory INSERT, as keep_lease_in_pass does. Before this change the new live case got SnapshotPublishConflictError; it now gets PublisherLeaseHeldError.
The resolution order ends in index_version so that when one pack's rows exist at two versions, a read resolves to the row a merge will keep (ReplacingMergeTree(index_version)). No native test covered this. Replays write byte-identical rows, so the pick never showed, and dropping index_version from kResolutionOrder passed every live suite. The parity suite now publishes a pack at 7 and writes it again at 9 with payload_offset moved, then requires the v9 rows from native search and get_by_ids, and agreement with the Python reader, before and after OPTIMIZE FINAL. An undefined tie resolves by the order parts are read in, so the v9 rows go once into a part written after the v7 part and once into one written before it. With the component dropped, the second case returns the v7 rows (3 of 3 runs).
A pinned read now joins capture_raw to a members subquery that supplies each pack's member_version, in place of the old (store_id, pack_id) IN membership. Both queries of the two-phase page read the join, so a page builds the members set twice. docs/benchmarks.md now records the cost against main (8b7991d), measured on a 2.1M-row synthetic corpus with superseding packs and unpublished and republished replays. Each build's captured statement was re-run 9 times, interleaved, with use_query_condition_cache=0. - 20k packs: selective pages 19-28% slower (+11 to +14 ms), about two members builds (11 ms each, against 2 ms for main's IN set). Unfiltered pages unchanged. - 100k packs: 0.88x to 1.08x. An independent review on another corpus shape measured 1.06x to 1.34x here. - 1M packs: 2.7x to 3.7x faster, because main's primary-key condition carries a 2M-element pack set. Rows read and capture_raw granules are identical in every case. Page contents differ from main only where main mis-ranks a replayed pack. The design doc's two-phase section points to the numbers. Sharing one members build between the two queries has not been measured.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #156, which the Claude GitHub app opened. Its two commits are re-landed on current main (8b7991d) and authored by Alan Liu. An earlier merge commit on #156's branch was dropped, and its conflict resolution redone and verified. Follow-up commits from an independent review are added on top.
Original description (#156)
Requested by Alan Liu · Slack thread
Before: Sometimes an indexing pass re-indexes a pack that is already published. That happens after a crash between publishing and
commit_packs, or after an outcome-unknown publish that actually landed, or on the C++ conflict path whencommit_packsfails. A replay like that could silently change which pack a pinned reader resolves a capture to. The pinned read went back to an older pack that a newer one had superseded. If the replaying pass never finished, the flip was permanent, and a merge made sure of it. This is the opposite of what the reader promises: "a pin still resolves to the pack it was taken over".After: A capture's packs are ranked by when each pack's publish reached the watermark, at or below the pin. That comes from the manifest paired with the watermark log. The version a descriptor row happened to be written at no longer counts. A replay now changes nothing a pinned reader sees, before or after a merge. A pack's rank is fixed by its first publish at or below the pin. A replay that goes on to publish the same pack again does not re-promote it over a pack that superseded it. The C++ indexer also no longer buries
kPublishConflictwhencommit_packsfails after a conflict.How
Found by TLA+ model checking of the publish protocol (
formal/tla/VersionPublish; configsReplay,ReplayCrash,ReplayCrashPinned). The scenario:commit_packs.argMax(..., (index_version, store_id, pack_id))ranks (3,P1) above (2,P2).Fix chosen: reader-side ranking (candidate (a)).
clickhouse_sql.member_versionsreturns each member pack withmin(index_version)over its paired manifest rows ≤ W, i.e. its first paired publish. It is built from the same manifest-rows fragment asmembership_predicate, so the two cannot drift. The public view's DDL is byte-identical to before.INwith anINNER JOIN ... USING (store_id, pack_id)against it.(member_version, store_id, pack_id, index_version).index_versiononly orders one pack's own rows. It picks the row aReplacingMergeTree(index_version)merge keeps, so pre-merge and post-merge reads agree. An existing live test (test_hydration_rejects_a_re_described_catalog_row) depends on this.Why not the writer-side skip (b):
SKIP_MEMBER_PACKS). It fixesReplayCrash/ReplayCrashPinnedbut still failsReplay(ResolvesNewest, 31-state trace). That is the rebuild-beside-the-live-indexer route: the second pass checks membership before the first pass's publish lands. No check before re-indexing can close that window.Design note (changed after review, 59c28ce):
member_versionis the pack's first paired publish, not its newest. With the newest publish, this sequence resolves X back to the older pack at every head ≥ v3, on the normal crash-recovery path:commit_packs.With the first publish, a pack's rank is fixed once it is first published. A real re-capture is a new pack, and a mirror is a different store, so both still get a fresh version. Pinned reads are stable either way.
C++ conflict guard:
indexer.cppnow catches acommit_packsfailure inside thekPublishConflicthandler and rethrows it askPublishConflict, with the commit error appended. This matchescatalog.py'sraise conflict from commit_failure.Docs updated:
clickhouse_reader.pymodule docstring (the pinned-resolution paragraph),_snapshot,_projection._publishdocstring incatalog.py. The retry's descriptor rewrite is no longer what supersession relies on.capture-storage-design.md(Phase 5 ordering).catalog-differential-review-2026-09-01.md, next to "Reader-visible corruption: none".Model re-check (measured before the switch to first-publish ranking; the scratch model was not committed and has not been re-run under
min): a scratch copy ofVersionPublish.tlawith aRANK_BY_MEMBERSHIPflag, plus an optionalMERGESaction that collapses a pack's rows to the highest version. The spec was not committed. TLC results:RANK_BY_MEMBERSHIPMERGESSkipRewritealso passes under the fix, which confirms the retry rewrite is no longer load-bearing. The candidate (b) results are in the previous section.Test evidence
New tests:
tests/test_clickhouse_capture_reader.py::test_a_replayed_pack_ranks_by_its_publish_not_by_its_rewritten_rows. It pins the join and the ordering at both query sites. It fails ona987dfeand passes with the fix.tests/test_clickhouse_snapshot_live.py::test_a_replayed_pack_does_not_flip_a_pinned_read[×2]. It runs the full scenario: pinnedget_by_idsandsearch, a cursor walk, a forced merge, then publishing the replay.tests/test_native_reader_parity_live.py::test_a_replayed_pack_does_not_flip_a_pinned_read_on_either_side. The same scenario against the native reader and the Python reader.Existing CPU tests were adjusted for the new SQL shape. The fake client now identifies the head read by
SELECT max(index_version) FROM.Commands and results:
python -m pytest -m cpu -q→ 1944 passed, 324 skipped. The skips are native conformance drivers that are not built here.minchange): the four live suites (snapshot, native reader parity, native capture storage, native catalog lease) gave 143 passed, andpytest -m cpugave 2301 passed. Before the fix, the replay tests fail on both the native and the Python reader.EXPLAIN indexes=1shows the same pruning forminandmax, with timings within noise on a 2M-row corpus.clickhouse_driver.Clientshim that is not committed:pytest tests/*_live.py -m "clickhouse and manual and not garage"→ 64 passed, 1 failed. The failure istest_a_role_that_cannot_see_one_object_is_told_to_grant_it_not_to_rebuild, which also fails on main under the shim because chdb has no roles. Ona987dfethe new live test fails with "the replay flipped the pinned read"; with the fix it passes.OPTIMIZE FINAL:a987dferesolves P2 → P1 → P1, and this branch resolves P2 → P2 → P2, for bothget_by_idsandsearch.g++ -std=c++20 -fsyntax-only -Wall -Wextrais clean onreader.cppandindexer.cpp. The full native build and the conformance driver could not be built here; this relies on CI's native-backend-compile and clickhouse-live jobs.test_native_reader_parity_live.pycould not run here.commit_packsfailure. This is a follow-up.Cost of the hot query: I measured an A/B of the
INform against the join form on the same data, interleaved, 9 trials, median, on chdb 26.7. The machine was shared, so timings are noisy.No regression stood out above the noise.
benchmarks/bench_capture_search.pyalso ran on both trees through the shim, three runs each: medians overlap and the noise was too large to resolve a difference. Primary-key pruning survives the join:test_selection_resolve_prunes_to_the_tenant_rangepasses under the shim.Re-landed on main, plus review fixes
How it composes with main:
reader.cppsearch()): both the inner key query and the outer argMax query readFROM snapshot(), and both carry the caller's filters. The innerWHEREis omitted when there are no filters. With the inner query switched back toFROM capture_raw, Resolve a search page's argMax for its own keys, not every row past the cursor #139'stest_a_page_walk_skips_unpublished_keys_without_ending_earlyfails, so the silent snapshot-drop trap is covered.index()split: the conflict-path guard is inNativeIndexer::commit()'s publish-retry loop, the only place wherecommit_packsruns on thekPublishConflictpath.Added commits:
kHeld/kLease) raised bycommit_packson the conflict path keeps its kind; it is no longer relabelled as a publish conflict. This matters after Bound how long the publisher lease can go unrenewed: a deadline for every request under it #159, whose renewal hook can refuse there.index_versiontiebreak within one pack.docs/benchmarks.md.INset, which is the price of pinned-read stability.Evidence:
maxinstead ofmin, or the inner query without the snapshot, each turns a test red.