Skip to content

Cluster/Calvin: barrier failure proposes no abort vote — peers hold locks until restart, client burns its deadline #393

Description

@EnRaiha

Version / build tested against

origin/main @ e235fe5 (2026-09-29)

Deployment mode

Origin — distributed / cluster

Engine(s) involved

Not engine-specific / unsure

Summary

A cross-shard dependent-read transaction whose passive barrier times out (or whose pre-stage dispatch fails) never proposes an abort vote to the sequencer group. The vote tally for that txn never completes, so no Verdict is emitted and every already-staged peer keeps holding its locks and staged buffer until process restart. The client burns its full deadline and reports a generic timeout.

Steps to reproduce

Reproduction path (runtime repro pending):

  1. rg -n "dependent_barrier" nodedb/src/control/cluster/calvin — locate the barrier sweep.
  2. Force a barrier timeout (fault injection or paused peer) with at least one peer already staged.
  3. Observe: locks released on the owning vShard, no SchedulerProposal::Vote { abort } proposed; peers stay in CommitState::AwaitingVerdict.
    Tests to extend: cargo test -p nodedb -- calvin.

Expected behavior

On barrier timeout or pre-stage dispatch failure, the owning vShard proposes an abort vote (leader-gated, kept in OwedEntries, re-proposed on the stall tick). The tally completes and peers leave AwaitingVerdict, releasing locks.

Actual behavior

read_result.rs timeout sweep releases locks and proposes nothing; active_dispatch.rs failures propose RoutingFailed/release locks only; the tally rule requires votes.len() == expected_participants, so the vote never completes.

What actually happened? (severity facts)

  • Acknowledged/committed data was lost, corrupted, or silently wrong
  • The server crashed, hung, or failed to start
  • A security or isolation boundary was crossed
  • Core functionality is broken with no acceptable workaround
  • A workaround exists (restart the affected node(s))

Proposed severity

SEV-2 — High: a cross-shard transaction wedges peer lock overlays until restart; stored data intact.

Reproducibility

Intermittent — some attempts (requires the timeout/dispatch-failure path)

Last known-good version / commit (if a regression)

(unknown / not a regression)

Environment & logs

Linux x86_64; Verified by static code reading at the pin above; runtime reproduction pending.
Code references:

  • (see prior-art line below)

Before submitting


Additional evidence (origin/main @ e235fe55c)

  • What: When a cross-shard dependent-read transaction's passive barrier times out, or its active dispatch fails before staging, the owning vShard never proposes a SequencerEntry::Vote { commit: false } / AbortVote. The tally never reaches expected_participants, so no Verdict is emitted. Every peer participant already staged and parked in CommitState::AwaitingVerdict keeps holding its locks and staged buffer indefinitely, because the stall sweep never aborts. The client burns its deadline and reports a generic timeout.
  • Where: nodedb/src/control/cluster/calvin/scheduler/driver/core/read_result.rs:67-84 (releases locks, proposes nothing); .../core/dispatch/active_dispatch.rs:41-51 (plan-decode failure — releases locks only) and :78,90 (routing failure — proposes RoutingFailed, not a vote); .../core/commit_resolve/verdict.rs:143-180 (holds locks, never aborts); vote cast site .../commit_resolve/vote.rs:58-65.
  • Evidence: quotes above; tally rule nodedb-cluster/src/calvin/completion_verdict.rs:101-104; module doc .../owed/entry.rs:8-11.
  • Impact: a cross-shard dependent-read txn that loses a passive response or fails local dispatch on one vShard wedges peer vShards' lock overlays until process restart; the stalled client waits out its full deadline instead of a fast abort.
  • Fix: on barrier timeout and on any pre-stage local dispatch failure, immediately propose_sequencer_entry(txn_id, SchedulerProposal::Vote { abort: Some(<reason>) }) (leader-gated; kept in OwedEntries — owed/entry.rs:144 — and re-proposed by retry_owed_sequencer_entries() on the stall tick, owed/retry.rs:70; the repropose-until-tally tests owed/retry.rs:162,312 are the pattern to extend), so the tally completes and the verdict releases peers. Add a test asserting that a timed-out barrier leaves a complete abort tally.
  • Prior-art: adjacent items examined — #392 (merged; does not add the vote), #353 (file-size refactor), #352/cluster: a saturated dispatch queue has no documented retryable class — clients cannot back off correctly; add one and a counter #371 (dispatch queue capacity/retry class). No open issue covers this path.
    Why: a cross-shard dependent-read txn whose barrier times out (or whose pre-stage dispatch fails) never proposes an abort vote, so the tally for that txn never completes and every already-staged peer keeps holding locks and its staged buffer until restart. Client sees only a generic timeout.
    Steps to verify: (1) run the existing calvin dependent-read tests with a forced barrier timeout (rg -n "dependent_barrier" nodedb/src/control/cluster/calvin to find the knob), (2) assert OwedEntries receives a Vote { abort } for the timed-out txn and peers leave AwaitingVerdict, (3) run cargo test -p nodedb -- calvin.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions