Skip to content

Enforce open-mode rules and add rekey_into_writer for restored stores - #45

Merged
farhan-syah merged 2 commits into
mainfrom
fix/observer-retry-scope
Sep 27, 2026
Merged

farhan-syah merged 2 commits into
mainfrom
fix/observer-retry-scope

Conversation

@farhan-syah

@farhan-syah farhan-syah commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

What changed

  • Every public open applies its mode's rules. Db::open_existing_with_counterpart_kek now opens through open_with_mode. It holds the writer sentinel, reports NotFound for a missing store, refuses a restored directory, and reads a page that fails authentication once instead of 4 times. The retry count is set where the pager is configured, so no open path can skip it. Only an Observer retries.
  • Restored directories keep their restore mode. snapshot_to and restore_from stamp both A/B header slots READ_ONLY. promote_to_follower and the new Db::open_follower record FOLLOWER. A Standalone open of either reports RestoredNotPromoted.
  • Header writers keep persisted fields. Commit, compaction, quota, and apply-journal headers carry restore_mode, flags, and the retention policy from state instead of writing 0.
  • New Db::open_follower reopens a Follower directory and replays an interrupted apply journal.
  • New Db::rekey_into_writer forks a ReadOnly or Follower handle into an independent Standalone writer. It re-encrypts every page and segment under a fresh file_id, a fresh kek_salt, and the supplied KEK into main.db.fork, and renames that over main.db. Realm, commit ids, keys, values, segment page ids, counters, and quotas carry over. Commit history starts empty. Fork segments wait in seg/.fork, which the orphan scan skips, and the first Standalone open adopts them into staging.
  • Advisory locks exclude every handle in the process. src/vfs/oslock.rs keeps one process-wide table, keyed by resolved lock-file path, and one OS lock per file. Before, each VFS instance had its own table, and macOS F_SETLK never conflicts within a process, so two handles on one directory could both take the writer sentinel. The GCD and IOCP backends drop their own lock copies and call the shared one.
  • Cleanups. Header-slot authentication is shared in pager/header.rs (authenticate_slot, authenticate_slot_with_kek). vfs::remove_if_present replaces three copies. STAGING_DIR and staging_path replace literal staging paths.

Why

  • restore_mode was never written, so RestoredNotPromoted never fired. A restored copy shares its source's file_id, kek_salt, and mk_epoch: one key and one nonce space. Opened as a Standalone writer, it sealed pages with nonces its source also issues.
  • open_existing_with_counterpart_kek bypassed open_with_mode. A second writer could attach beside its handle, and it kept the Observer retry count.
  • A fork must never seal a byte under the source key. Building the replacement file under a fresh identity and publishing it by one rename meets that, and a fork that fails before the rename leaves the directory restored and readable.

How to check it

  • cargo nextest run --all-features: 841 passed, 10 skipped.
  • New suites: tests/open_mode_discipline.rs, tests/restored_store_modes.rs, tests/restored_store_fork.rs. They cover each open path, the restore-mode lifecycle, the fork's fresh salt, and a fork interrupted before and after the rename.
  • cargo clippy --all-targets --all-features -- -D warnings, RUSTDOCFLAGS='-D warnings' cargo doc, the wasm32-unknown-unknown --features opfs and wasm32-wasip1 checks, and cargo check --all-targets for aarch64-apple-darwin and x86_64-pc-windows-gnu are clean.

@farhan-syah farhan-syah changed the title Enforce restored store isolation: block unsafe Standalone opens Enforce open-mode rules and add rekey_into_writer for restored stores Sep 27, 2026
Each native backend (Tokio, io_uring, GCD, IOCP) kept its own in-process
lock table, so distinct VFS instances rooted at the same directory could
each believe they held the OS lock. GCD and IOCP also each carried their
own bespoke fcntl/LockFileEx implementation.

Replace the per-VFS tables with one process-wide table in oslock, keyed
by the resolved (canonicalized parent + file name) lock path, so every
instance over one directory contends correctly regardless of path
spelling. A per-domain gate serializes acquiring the OS lock, and the
single OS lock descriptor is now owned by the shared entry instead of
each handle, so an early handle close can no longer drop the F_SETLK
lock out from under a remaining in-process holder. GCD and IOCP now
route through the same implementation instead of their own copies.
A snapshot directory and a directory `restore_from` fills copy
`main.db` byte for byte, so they share the source's DEK and nonce
space with it. Opening one Standalone previously proceeded anyway,
letting independent writes on both directories repeat nonces under one
key.

Record a restore mode (STANDALONE, READ_ONLY, FOLLOWER) in the main.db
header. A Standalone open now refuses any non-STANDALONE mode with
RestoredNotPromoted. A restored directory can still open ReadOnly, or
promote to Follower and track its source. `rekey_into_writer` forks a
ReadOnly or Follower handle into a Standalone writer under a fresh
identity: it rekeys the tree and quota catalogs and rewrites segments
into a fork directory the open-time orphan scan skips, then adopts
those segments into staging and publishes a fresh header once the fork
takes the writer sentinel.

Thread the KEK-changing-rekey resume path through open_with_mode via
an optional counterpart KEK, and remove_if_present to make crash
cleanup of scratch files idempotent.
@farhan-syah
farhan-syah force-pushed the fix/observer-retry-scope branch from fc2a5cb to 7badf9a Compare September 27, 2026 21:14
@farhan-syah
farhan-syah merged commit 629aee0 into main Sep 27, 2026
19 checks passed
@farhan-syah
farhan-syah deleted the fix/observer-retry-scope branch September 27, 2026 21:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant