Skip to content

fix(wal): write the batch on wasm32-wasip1 instead of reporting it written - #388

Closed
EnRaiha wants to merge 1 commit into
NodeDB-Lab:mainfrom
EnRaiha:fix/wal-wasi-write
Closed

EnRaiha wants to merge 1 commit into
NodeDB-Lab:mainfrom
EnRaiha:fix/wal-wasi-write

Conversation

@EnRaiha

@EnRaiha EnRaiha commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Problem

On wasm32-wasip1, WalWriter::flush_buffer wrote nothing and returned
Ok(()). The batch is written by a #[cfg(unix)] arm with no other arm, and
cfg(unix) is false on this target (target_family is wasm), so the write was
not compiled at all while the statements after it still ran: file_offset
advanced by the batch length, the buffer was cleared, record_flush() was called
and the flush returned success. sync() reported a durable append for bytes that
never reached the file.

fsync_directory had the same shape of problem: it opened the directory and
called sync_all, which wasi preview1 cannot do, so every caller that renames and
then fsyncs a directory failed on that target.

Both are pre-existing at 1ff35512b: the wasip1 test job that would have
exercised them is the subject of a separate change, not this one.

Change

writer/flush.rs — the unix arm is cfg(all(unix, not(target_arch = "wasm32")))
now, and wasm gets its own arm: seek to file_offset, then write. It covers every
wasm32 target, including wasm32-unknown-unknown where libc defines no pwrite
at all, and std's wasi positional API is still unstable; the segment is only
ever appended to, so seeking to the end and writing is equivalent to a
positional write. A failed write returns before the shared bookkeeping, so the buffer
and the offset survive for a byte-for-byte retry — the property the pwrite arm
documents. write_error became classify_write_error, which takes the failed
write's own io::Error instead of re-reading the thread's errno, and its
full-device branch is keyed on ErrorKind::StorageFull rather than
libc::ENOSPC: libc defines no constants at all for
wasm32-unknown-unknown, which this crate is compiled for by the existing
wasm32-decoders job, and std maps the full-device errno to that kind on every
target with a filesystem. A full device therefore stays WalError::OutOfSpace
on wasi instead of degrading to a transient Io. A compile_error! guard
refuses a target family that is neither unix nor wasm32, which would otherwise
compile no arm at all.

segment/atomic_io.rs — fsync_directory is a documented no-op on wasm32.
The rename still succeeds; what the target cannot provide is the assurance that
the directory entry survives a host crash. The crash-injection failpoint is still
evaluated first, so wal::fsync_directory keeps working there.

double_write/raw_io.rs — the reason pwrite_all reports Unsupported on
wasm rather than falling back to a seek-and-write: a fallback would mirror the
slot and the writer would report DwbProtection::Active for protection that
recover_record cannot read back on that target. That is the Direct path.
Buffered still reports protection it cannot deliver — that is a separate defect
and needs a read half, not a write gate.

Cargo.toml — tokio (workspace features = ["full"]) and fluxbench are
native-only dev-dependencies now. Tokio rejects fs/io-std/net/process/
rt-multi-thread/signal on wasm, Cargo unifies dev-dependency features across the
whole cargo test invocation, and neither is reachable from a wasm test: this
crate has no tokio call sites, and fluxbench is used only by
benches/wal_throughput.rs, which cargo test does not build. Without that gate
no test target in this crate builds for wasm, which is why the probe below could
not run at all before it.

Evidence

On 51e17220 (fix/wal-wasi-write, based on 1ff35512b), rustc/cargo 1.96.1,
wasmtime 35.0.0. The wasip1 probe is run with --nocapture on purpose: a panic
aborts the process on this target, so without it a failed assertion arrives as a
wasm trap with no message.

Check Command Result
wasip1 probe CARGO_TARGET_WASM32_WASIP1_RUNNER="wasmtime --dir=." cargo test -p nodedb-wal --target wasm32-wasip1 --test wasi_append -- --nocapture 2 passed, exit 0
wasip1 failure path same, plus --features failpoints 3 passed, exit 0
native regression cargo test -p nodedb-wal 270 passed, 0 failed (218 unit + 52 wal_suite), exit 0
format cargo fmt --all -- --check exit 0
lints, native cargo clippy -p nodedb-wal --all-targets -- -D warnings exit 0
lints, wasip1 cargo clippy -p nodedb-wal --target wasm32-wasip1 --lib -- -D warnings exit 0
wasip1 build cargo check -p nodedb-wal --target wasm32-wasip1 --lib exit 0
wasm32-unknown-unknown cargo check --target wasm32-unknown-unknown -p nodedb-codec -p nodedb-columnar -p nodedb-strict (the wasm32-decoders job this PR triggers) exit 0
lints, wasm32-unknown-unknown cargo clippy -p nodedb-wal --target wasm32-unknown-unknown --lib -- -D warnings exit 0

Red, with the two source files reverted to 1ff35512b and the probe unchanged:

$ ... --test wasi_append -- an_appended_record_reaches_the_file --nocapture
thread 'main' (1) panicked at nodedb-wal/tests/wasi_append.rs:59:5:
assertion `left == right` failed: sync() reported success with 73 bytes written, but the segment holds 0 bytes
  left: 0
 right: 73
error: test failed                                                        (exit 134)

$ ... --test wasi_append -- a_checkpoint_can_fsync_its_directory --nocapture
thread 'main' (1) panicked at nodedb-wal/tests/wasi_append.rs:75:33:
wasi preview1 has no directory fsync, not a failure: Io(Os { code: 8, kind: Uncategorized, message: "Bad file descriptor" })
error: test failed                                                        (exit 134)

And the arm itself is the chokepoint, not the wrapper above it: with the wasi
arm compiled out (#[cfg(any())]) and the rest of the change intact, the armed
test fails —

thread 'main' (1) panicked at nodedb-wal/tests/wasi_append.rs:123:22:
an armed write failpoint must not be reported as a flush: ()
                                                                          (exit 134)

Green, with the change in place: 2 passed by default, 3 passed with
--features failpoints.

Limits, stated rather than implied

  • The wasi arm's failure path is entered and classified under
    --features failpoints: wal::wasm_flush_write sits inside the arm, the error
    it raises goes through the same classify_write_error call the real failure
    does, and the test asserts OutOfSpace, an unchanged file_offset and an
    empty segment. Compiling the arm out makes the injection unfireable and that
    test fails, so it is a chokepoint for the arm rather than for the wrapper above
    it. Without the feature the test is compiled out rather than weakened — the
    default run is 2 tests, the failpoint run 3.
  • What is not covered: a real write_all failure on wasm. The errno comes from
    the runtime and cannot be injected here, so the arm's own map_err closure is
    exercised only through the same classifier the failpoint path calls, and the
    real full device is verified by the mapping std owns — the classifier keys on
    ErrorKind::StorageFull, and the probe's injected error carries that kind
    rather than a raw errno, so what it pins is the classifier's rule, not the
    runtime's errno translation.
  • fsync_directory on wasi is a weaker guarantee than the name suggests, and the
    doc now says so. Callers on that target keep the rename but not the
    survives-a-crash assurance.
  • The padded-slice path is untested on wasm: use_direct_io defaults to true in
    the writer config, and on targets without O_DIRECT the open ignores the flag
    while flush_buffer still writes the padded aligned slice. The probe uses
    open_without_direct_io, so only the unpadded path runs.
  • The wasip1 test job itself is a separate change; this one only makes the
    defect it would have caught impossible. Until that job exists, the probe runs
    only when someone invokes it as above, with --features failpoints for the
    full set.

Fixes #387.

…itten

`WalWriter::flush_buffer` wrote the batch with a `#[cfg(unix)]` arm and had no
other arm. `cfg(unix)` is false on `wasm32-wasip1` — `target_family` is `wasm` —
so on that target the write was not compiled at all, while everything after it
still ran: `file_offset` advanced by the batch length, the buffer was cleared,
`durability.record_flush()` was called and the flush returned `Ok(())`. `sync()`
therefore reported a durable append for bytes that never reached the file, and
`file_offset()` agreed with it.

The wasm arm seeks to `file_offset` and writes. It covers every wasm32 target,
including `wasm32-unknown-unknown` where libc defines no `pwrite` at all, and
std's wasi positional API is still unstable; the segment is only ever appended
to, so seeking to the end and writing is equivalent to a positional write. A
failed write returns before the shared bookkeeping, so the buffer and the offset
survive for a byte-for-byte retry — the property the `pwrite` arm documents.
`write_error` is now `classify_write_error`, taking the failed write's own
`io::Error` instead of re-reading the thread's errno. The full-device branch is
keyed on `ErrorKind::StorageFull` rather than `libc::ENOSPC`, for two reasons:
`libc` defines no constants at all for `wasm32-unknown-unknown`, which this crate
is compiled for by the existing `wasm32-decoders` CI job, and std maps the
full-device errno to that kind on every target that has a filesystem. So a full
device stays `WalError::OutOfSpace` on wasi, which is the rule the function's own
doc states, rather than degrading to a transient `Io`. A `compile_error!` guard
refuses a target family that is neither unix nor wasm32: that target would
compile no arm at all and reinstate the silent success this change removes.

`fsync_directory` had the same shape of problem: it opened the directory and
called `sync_all`, which wasi preview1 cannot do, so every caller that renames
and then fsyncs a directory failed on that target. It is a documented no-op
there now, with the crash-injection failpoint still evaluated first.

`tests/wasi_append.rs` decides both from the file rather than the return value:
the segment's length must equal the offset the writer reported, and the payload
must be in it. Against the code above it fails with

    assertion `left == right` failed: sync() reported success with 73 bytes
    written, but the segment holds 0 bytes

and, filtered to the other test,

    wasi preview1 has no directory fsync, not a failure:
    Io(Os { code: 8, kind: Uncategorized, message: "Bad file descriptor" })

Both pass with the change. Run it with `--nocapture`: a panic aborts the process
on this target, so otherwise the failure arrives as a wasm trap with no message.

A third test arms `wal::wasm_flush_write`, which sits inside the wasi arm, and
asserts that the failed write is reported as a failure, is classified
`OutOfSpace` — the injected error carries the kind a real full device produces,
so the branch is exercised on wasm rather than argued — leaves the segment empty
and does not advance `file_offset`. That injection is the only way to reach
`classify_write_error` from the arm on a runtime whose writes cannot be made to
fail on demand, and it is a chokepoint for the arm: compiling the arm out makes
the injection unfireable and the test fails with "an armed write failpoint must
not be reported as a flush". It needs `--features failpoints`; without the
feature the injection expands to nothing and the assertions would describe a
successful flush, so it is compiled out instead of weakened.

The crate's `tokio` (workspace `features = ["full"]`) and `fluxbench`
dev-dependencies are native-only now. Tokio rejects
fs/io-std/net/process/rt-multi-thread/signal on wasm, Cargo unifies
dev-dependency features across the whole `cargo test` invocation, and neither is
reachable from a wasm test: this crate has no tokio call sites, and `fluxbench`
is used only by benches/wal_throughput.rs, which `cargo test` does not build.
Without that gate no test target in this crate builds for wasm, which is why the
probe above could not run at all before it.

`double_write/raw_io.rs` gains the reason `pwrite_all` reports `Unsupported` on
wasm rather than falling back to a seek-and-write: a fallback would mirror the
slot successfully and the writer would report `DwbProtection::Active` for
protection that `recover_record` cannot read back on that target. That is the
`Direct` path; `Buffered` still reports protection it cannot deliver, which is a
separate defect and needs a read half rather than a write gate.

wasm32-wasip1 under wasmtime 35.0.0: 2 passed by default, 3 with
`--features failpoints`. `cargo check --target wasm32-unknown-unknown -p
nodedb-codec -p nodedb-columnar -p nodedb-strict` (the `wasm32-decoders` job)
still exits 0. Native unchanged at 270 passed, 0 failed. fmt and clippy clean for
this crate on native, on wasm32-wasip1 and on wasm32-unknown-unknown.
Copilot AI lite review requested due to automatic review settings September 27, 2026 15:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@farhan-syah

Copy link
Copy Markdown
Member

Thanks for the careful analysis. The defect is real: on wasm32 the flush advanced the offset and reported success without writing.

We are closing this as superseded. The Origin WAL writer does not target wasm. nodedb-wal on wasm32 now builds only crypto, secure_mem, error and record::header, the pieces Lite and the shared crates use. The writer, segment I/O and double-write modules are gated #[cfg(not(target_arch = "wasm32"))], so the silent-success arm no longer exists. #378 and #387 are closed on the same basis.

The change is in fix/calvin-dispatch-and-typed-errors. If a shared crate needs another WAL piece on wasm, please open an issue that names the caller.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

nodedb-wal: every append on wasm32-wasip1 reports success with nothing written

3 participants