Skip to content

cranelift: aarch64: zero-extend expected value in i32 atomic_cas loop - #14525

Open
kevaundray wants to merge 2 commits into
bytecodealliance:mainfrom
kevaundray:fix-aarch64-atomic-cas-i32
Open

kevaundray wants to merge 2 commits into
bytecodealliance:mainfrom
kevaundray:fix-aarch64-atomic-cas-i32

Conversation

@kevaundray

Copy link
Copy Markdown
Contributor

This PR description and commits were generated with the use of claude code. Feel free to take the test case and write another PR with another fix :)

On aarch64 without LSE, atomic_cas.i32 is lowered to AtomicCASLoop, which compared the loaded value against the expected value with a full 64-bit cmp x27, x26. The ldaxr zero-extends the loaded value, but the upper 32 bits of the register holding the i32 expected value are unspecified in the aarch64 backend. For example, ireduce.i32 of an i64 reuses the same register. When those bits are nonzero, the compare fails even though the low 32 bits match, so the exchange is skipped and the old value is returned. CLIF semantics, the interpreter and the other backends all perform the exchange.

I8/I16 already compare with uxtb/uxth. This PR does the same for I32 and emits cmp x27, w26, uxtw.

Commits

  1. Tests showing the bug:
    • runtests/atomic-cas-i32-upper-bits.clif (test interpret + test run, same targets as atomic-cas.clif): uses ireduce.i32 of an i64 with nonzero upper bits as the expected value.
    • A new function in isa/aarch64/atomic-cas.clif (precise output), blessed with CRANELIFT_TEST_BLESS=1. It shows the buggy cmp x27, x26.
  2. Fix in inst/emit.rs, which also updates the sequence comment. The AtomicCASLoop I32 encoding in emit_tests.rs is updated, and isa/aarch64/atomic-cas.clif is re-blessed, so both functions now show cmp x27, w26, uxtw.

Reduced repro:

function %cas32(i64, i32, i32) -> i32, i32 {
    ss0 = explicit_slot 8
block0(v0: i64, v1: i32, v2: i32):
    v3 = stack_addr.i64 ss0
    store.i32 v1, v3
    v4 = ireduce.i32 v0
    v5 = atomic_cas.i32 v3, v4, v2
    v6 = load.i32 v3
    return v5, v6
}
; run: %cas32(0x100000005, 5, 7) == [5, 7]   ; aarch64 (no LSE) returned [5, 5]

Evidence

I cross-built clif-util for aarch64-unknown-linux-musl on an x86_64 host at each commit and ran it under qemu-aarch64-static:

cargo build --release -p cranelift-tools --no-default-features --features all-arch \
  --target aarch64-unknown-linux-musl     # linker = rust-lld

Commit 1 (tests only):

$ qemu-aarch64-static clif-util-before test cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif
FAIL cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif: run

Caused by:
    Failed test: run: %atomic_cas_i32_ireduce(5, 4294967301, 7) == [5, 7], actual: [5, 5]
1 tests
Error: 1 failure

Commit 2 (fix):

$ qemu-aarch64-static clif-util-after test cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif
1 tests
$ qemu-aarch64-static clif-util-after test cranelift/filetests/filetests/runtests/atomic-cas{,-little,-subword-little,-subword-big}.clif
4 tests

Host (x86_64):

$ cargo run -p cranelift-tools -- test cranelift/filetests/filetests/isa/aarch64/ \
    cranelift/filetests/filetests/runtests/atomic-cas-i32-upper-bits.clif \
    cranelift/filetests/filetests/runtests/atomic-cas.clif \
    cranelift/filetests/filetests/runtests/atomic-cas-little.clif \
    cranelift/filetests/filetests/runtests/atomic-cas-subword-little.clif
138 tests
$ cargo test -p cranelift-codegen --features all-arch
test result: ok. 232 passed; 0 failed; ...
test result: ok. 22 passed; 0 failed; 4 ignored; ...

The new runtest passes in the interpreter and on x86_64 both before and after the fix. I checked the riscv64 (has_a) and s390x lowerings with clif-util compile -D: riscv64 zero-extends both operands before bne, and s390x uses cs, which compares only 32 bits. Neither is affected.

On aarch64 without LSE, `atomic_cas.i32` is lowered to an
`AtomicCASLoop` whose comparison is `cmp x27, x26`: a 64-bit compare
of the zero-extended loaded value against the whole 64-bit register
holding the expected value. The upper 32 bits of a register holding an
`i32` are unspecified in the aarch64 backend (e.g. `ireduce.i32` of an
`i64` reuses the same register), so when they are nonzero the compare
fails even though the low 32 bits match: the exchange is not performed
and the old value is returned.

Add a runtest exercising this (`ireduce.i32` of an `i64` with nonzero
upper bits as the expected value) and a precise-output test showing the
current, incorrect `cmp x27, x26` sequence. The runtest fails on
aarch64 (without `has_lse`); it passes in the interpreter and on the
other backends.
The `AtomicCASLoop` sequence compared the loaded value against the
expected value with `cmp x27, x26` for `I32`, relying on the `ldaxr`
zero-extending the loaded value. However, the upper 32 bits of the
register holding the `i32` expected value are unspecified, so if they
were nonzero the compare failed even when the low 32 bits matched, and
the exchange was not performed.

Use the extended-register form `cmp x27, w26, uxtw` for `I32`, as is
already done with `uxtb`/`uxth` for `I8`/`I16`.
@kevaundray
kevaundray requested a review from a team as a code owner October 4, 2026 18:06
@kevaundray
kevaundray requested review from cfallin and removed request for a team October 4, 2026 18:06
@github-actions github-actions Bot added cranelift Issues related to the Cranelift code generator cranelift:area:aarch64 Issues related to AArch64 backend. labels Oct 4, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cranelift:area:aarch64 Issues related to AArch64 backend. cranelift Issues related to the Cranelift code generator

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant