Rank a capture's packs by publish version so a replayed pack cannot flip a pinned read - #156
Closed
claude[bot] wants to merge 3 commits into
Closed
claude[bot] wants to merge 3 commits into
claude[bot] wants to merge 3 commits into
Conversation
A pass that re-indexes an already-published pack (crash before commit_packs, an outcome-unknown publish that landed, a conflict whose commit_packs failed, a rebuild beside the live indexer) writes that pack's descriptor rows at a fresh, higher version before publishing anything. The reader ranked a capture's candidate rows by the row's own index_version, and membership bounds packs rather than rows, so those rows outranked a newer pack describing the same capture inside snapshots that were already pinned. If the replay never published, that stayed true. Found by TLA+ model checking (VersionPublish Replay/ReplayCrash/ ReplayCrashPinned). The Python and C++ readers now join the snapshot's packs with the newest version their publish reached the watermark at (clickhouse_sql. member_versions, the same paired-manifest rows as membership_predicate) and resolve on (member_version, store_id, pack_id, index_version). The public view's DDL is unchanged. The C++ indexer also gets Python's guard on the conflict path: a commit_packs failure no longer replaces kPublishConflict.
zaoxing
force-pushed
the
claude/reindex-pinned-reader
branch
from
September 24, 2026 22:59
721bbc9 to
439f465
Compare
zaoxing
marked this pull request as ready for review
September 24, 2026 23:37
The reader ranked a capture's packs on the newest paired publish at or below the pin, max(index_version) over the manifest. A replay that does publish moves that rank: pass A publishes P1 at v1 and dies before commit_packs, pass B publishes P2 at v2 describing the same capture, and a later pass that re-indexes P1 publishes it again at v3. From v3 on every head resolved the capture back to P1, the older pack, on the normal crash-recovery path. Pins stayed correct; heads did not. Both readers now take min(index_version): a pack's rank is fixed once it is first published, so a replay, published or not, cannot move it. Pins are stable either way, since a later publish lands above the pin. A genuine re-capture is a new pack and a mirror is another store, so both still get a fresh first publish. The Python and C++ member_versions subqueries change together and stay identical. The live replay test and the native/Python parity test now publish the replay and require P2 at the new head, before and after a merge; both failed with max. EXPLAIN indexes=1 plans for the probe corpus (2M rows, 20k packs) are identical under min and max, as are rows read, and timings are within run-to-run noise.
Contributor
Author
|
Re-checked the VersionPublish TLA+ model with 59c28ce's first-publish ranking (
The model's old Generated by Claude Code |
claude Bot
pushed a commit
that referenced
this pull request
Sep 25, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019oKb8SWwCwPRAtxuWLQSXt
Brings in #149-#152 and #139 (two-phase native search page). One conflict, native/csrc/catalog/reader.cpp search(): #139 split the page into an inner key-only query and an outer argMax query over the same filters, while this branch moved snapshot membership out of `clauses` into the FROM clause (snapshot(), the join that supplies member_version). Resolved by running BOTH phases over snapshot() with the same caller `clauses` (possibly empty), so the inner LIMIT only counts member keys (#139's walk-ends-early guard) and the outer argMax still ranks on (member_version, store_id, pack_id, index_version). The design doc's two-phase paragraph now says both queries read the same snapshot join. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019oKb8SWwCwPRAtxuWLQSXt
zaoxing
added a commit
that referenced
this pull request
Sep 28, 2026
…ot flip a pinned read (#161) * Rank a capture's packs by publish version, not descriptor version * Rank a pack by its first publish, not its newest * Test the native indexer's conflict-path guard * Keep a lease refusal's kind when a conflicted publish cannot record its packs * Test that a pinned read picks, within one pack, the row a merge keeps * Record what the snapshot join costs a native search page Replaces #156.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Alan Liu · Slack thread
Before: Sometimes an indexing pass re-indexes a pack that is already published. That happens after a crash between publishing and
commit_packs, or after an outcome-unknown publish that actually landed, or on the C++ conflict path whencommit_packsfails. A replay like that could silently change which pack a pinned reader resolves a capture to. The pinned read went back to an older pack that a newer one had superseded. If the replaying pass never finished, the flip was permanent, and a merge made sure of it. This is the opposite of what the reader promises: "a pin still resolves to the pack it was taken over".After: A capture's packs are ranked by when each pack's publish reached the watermark, at or below the pin. That comes from the manifest paired with the watermark log. The version a descriptor row happened to be written at no longer counts. A replay now changes nothing a pinned reader sees, before or after a merge. A pack's rank is fixed by its first publish at or below the pin. A replay that goes on to publish the same pack again does not re-promote it over a pack that superseded it. The C++ indexer also no longer buries
kPublishConflictwhencommit_packsfails after a conflict.How
Found by TLA+ model checking of the publish protocol (
formal/tla/VersionPublish; configsReplay,ReplayCrash,ReplayCrashPinned). The scenario:commit_packs.argMax(..., (index_version, store_id, pack_id))ranks (3,P1) above (2,P2).Fix chosen: reader-side ranking (candidate (a)).
clickhouse_sql.member_versionsreturns each member pack withmin(index_version)over its paired manifest rows ≤ W, i.e. its first paired publish. It is built from the same manifest-rows fragment asmembership_predicate, so the two cannot drift. The public view's DDL is byte-identical to before.INwith anINNER JOIN ... USING (store_id, pack_id)against it.(member_version, store_id, pack_id, index_version).index_versiononly orders one pack's own rows. It picks the row aReplacingMergeTree(index_version)merge keeps, so pre-merge and post-merge reads agree. An existing live test (test_hydration_rejects_a_re_described_catalog_row) depends on this.Why not the writer-side skip (b):
SKIP_MEMBER_PACKS). It fixesReplayCrash/ReplayCrashPinnedbut still failsReplay(ResolvesNewest, 31-state trace). That is the rebuild-beside-the-live-indexer route: the second pass checks membership before the first pass's publish lands. No check before re-indexing can close that window.Design note (changed after review, 59c28ce):
member_versionis the pack's first paired publish, not its newest. With the newest publish, this sequence resolves X back to the older pack at every head ≥ v3, on the normal crash-recovery path:commit_packs.With the first publish, a pack's rank is fixed once it is first published. A real re-capture is a new pack, and a mirror is a different store, so both still get a fresh version. Pinned reads are stable either way.
C++ conflict guard:
indexer.cppnow catches acommit_packsfailure inside thekPublishConflicthandler and rethrows it askPublishConflict, with the commit error appended. This matchescatalog.py'sraise conflict from commit_failure.Docs updated:
clickhouse_reader.pymodule docstring (the pinned-resolution paragraph),_snapshot,_projection._publishdocstring incatalog.py. The retry's descriptor rewrite is no longer what supersession relies on.capture-storage-design.md(Phase 5 ordering).catalog-differential-review-2026-09-01.md, next to "Reader-visible corruption: none".Model re-check (measured before the switch to first-publish ranking; the scratch model was not committed and has not been re-run under
min): a scratch copy ofVersionPublish.tlawith aRANK_BY_MEMBERSHIPflag, plus an optionalMERGESaction that collapses a pack's rows to the highest version. The spec was not committed. TLC results:RANK_BY_MEMBERSHIPMERGESSkipRewritealso passes under the fix, which confirms the retry rewrite is no longer load-bearing. The candidate (b) results are in the previous section.Test evidence
New tests:
tests/test_clickhouse_capture_reader.py::test_a_replayed_pack_ranks_by_its_publish_not_by_its_rewritten_rows. It pins the join and the ordering at both query sites. It fails ona987dfeand passes with the fix.tests/test_clickhouse_snapshot_live.py::test_a_replayed_pack_does_not_flip_a_pinned_read[×2]. It runs the full scenario: pinnedget_by_idsandsearch, a cursor walk, a forced merge, then publishing the replay.tests/test_native_reader_parity_live.py::test_a_replayed_pack_does_not_flip_a_pinned_read_on_either_side. The same scenario against the native reader and the Python reader.Existing CPU tests were adjusted for the new SQL shape. The fake client now identifies the head read by
SELECT max(index_version) FROM.Commands and results:
python -m pytest -m cpu -q→ 1944 passed, 324 skipped. The skips are native conformance drivers that are not built here.minchange): the four live suites (snapshot, native reader parity, native capture storage, native catalog lease) gave 143 passed, andpytest -m cpugave 2301 passed. Before the fix, the replay tests fail on both the native and the Python reader.EXPLAIN indexes=1shows the same pruning forminandmax, with timings within noise on a 2M-row corpus.clickhouse_driver.Clientshim that is not committed:pytest tests/*_live.py -m "clickhouse and manual and not garage"→ 64 passed, 1 failed. The failure istest_a_role_that_cannot_see_one_object_is_told_to_grant_it_not_to_rebuild, which also fails on main under the shim because chdb has no roles. Ona987dfethe new live test fails with "the replay flipped the pinned read"; with the fix it passes.OPTIMIZE FINAL:a987dferesolves P2 → P1 → P1, and this branch resolves P2 → P2 → P2, for bothget_by_idsandsearch.g++ -std=c++20 -fsyntax-only -Wall -Wextrais clean onreader.cppandindexer.cpp. The full native build and the conformance driver could not be built here; this relies on CI's native-backend-compile and clickhouse-live jobs.test_native_reader_parity_live.pycould not run here.commit_packsfailure. This is a follow-up.Cost of the hot query: I measured an A/B of the
INform against the join form on the same data, interleaved, 9 trials, median, on chdb 26.7. The machine was shared, so timings are noisy.No regression stood out above the noise.
benchmarks/bench_capture_search.pyalso ran on both trees through the shim, three runs each: medians overlap and the noise was too large to resolve a difference. Primary-key pruning survives the join:test_selection_resolve_prunes_to_the_tenant_rangepasses under the shim.