cranelift: mask shift amount in (x << k) >> k mid-end rules - #14526
Merged
alexcrichton merged 2 commits intoOct 4, 2026
Merged
alexcrichton merged 2 commits into
alexcrichton merged 2 commits into
Conversation
The mid-end rules rewriting `(x << k) >> k` into `sextend`/`uextend` of an `ireduce` compute the narrow type from `ty_bits(ty) - k` using the raw shift constant. Shift amounts are taken modulo the bit width of the shifted type, so for example `iconst.i64 -8` shifts an `i8` by 0, but `8 - 0xffff_ffff_ffff_fff8` wraps to 16 and the rule builds `sextend.i8 (ireduce.i16 v0)` with `v0: i8`, which is ill-typed. With `opt_level=speed` compilation then aborts with a verifier error, even though the original expression is just `v0`. Add a runtest with such out-of-range amounts (plus in-range and congruent-to-in-range amounts) and `test optimize` cases showing the expected optimized output. The `sshr` cases currently fail with a verifier error. The `ushr` cases happen to pass because elaboration picks `v0` over the ill-typed `uextend` node in the same e-class.
The rules rewriting `(x << k) >> k` into `sextend`/`uextend` of an `ireduce` used the raw shift constant to compute the narrow type as `ty_bits(ty) - k`. Shift amounts are taken modulo the bit width of the shifted type, so an out-of-range `k` could select a narrow type that is wider than `ty` (e.g. `k = -8` on an `i8` yields `i16`), producing ill-typed IR that fails the verifier. Mask the shift amount with `ty_shift_mask` before using it, as the neighbouring shift rules do. This also lets the rules fire directly on amounts congruent to an in-range amount, such as `56` on an `i32`.
Subscribe to Label Actioncc @avanhatt, @cfallin, @fitzgen, @mmcloughlin DetailsThis issue or pull request has been labeled: "cranelift", "isle"Thus the following users have been cc'd because of the following labels:
To subscribe or unsubscribe from this label, edit the |
alexcrichton
approved these changes
Oct 4, 2026
alexcrichton
left a comment
Member
There was a problem hiding this comment.
Thanks!
cc @avanhatt and @mmcloughlin for an ISLE opt bug which wasn't caught through verification
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bug
In
cranelift/codegen/src/opts/shifts.isle, the two rules after;; (x << N) >> N == x as T_SMALL as T_LARGEcompute the narrow type asty_bits(ty) - Nfrom the raw shift constant. CLIF shift amounts are taken modulo the bit width of the shifted type, soiconst.i64 -8shifts ani8by 0, but8 - 0xffff_ffff_ffff_fff8wraps to16.shift_amt_to_typethen returnsi16, and the rule buildssextend.i8 (ireduce.i16 v0)withv0: i8, which is ill-typed. Withopt_level=speedcompilation aborts in the verifier, even though the original expression is justv0.The bug affects every target because it's in the mid-end. Other affected amounts:
-24oni8(picksi32) and-16oni16(picksi32). Theushrrule builds the same ill-typeduextend (ireduce ..)node. In the tests here, elaboration happens to pickv0from the same e-class, so the bad node never reaches the verifier.Before (commit 1, without the fix)
When each
sshrfunction is run on its own, they fail the same way:-24/i8buildsireduce.i32 v0/sextend.i8, and-16/i16buildsireduce.i32 v0/sextend.i16.After (commit 2)
All of these pass on x86_64. The runtest runs natively on x86_64, on the pulley targets and in the interpreter. After the fix, the
ushrcases optimize toreturn v0. Thesshrcases leave the shifts by0in place: there is nosshr x, 0simplification, which is a separate issue.