Skip to content

jitv2: one branch in step_jit's scheduling path; drop the hot lane - #149

Closed
atomchild411 wants to merge 1 commit into
techomancer:mainfrom
atomchild411:jitv2-single-branch
Closed

atomchild411 wants to merge 1 commit into
techomancer:mainfrom
atomchild411:jitv2-single-branch

Conversation

@atomchild411

Copy link
Copy Markdown
Contributor

As discussed: one branch in step_jit's scheduling path instead of two.

A page whose compile request was already queued took a second branch on
every arrival: count the arrival, and past 256 put a copy of the request on
a hot lane the workers drained first. It was built for the first program
after boot on a machine with few compile threads, waiting behind a long
queue. With more threads it buys nothing, and it put an atomic add on the
interpreted path of every page waiting for its compile.

Now scheduling is one test: if this arrival wins try_schedule_page, push
the request (and clear the flag if the queue is full); either way,
interpret. HOT_QUEUE, push_hot_request, note_waiting_arrival, the
page's hot_wait/promoted fields and their resets, and the worker's
hot-first pop are gone (1 line added, 67 removed).

Testing

  • cargo test --lib --features jitv2: 1116 passed, 0 failed, on current
    main (95b7ad0).
  • Indy (R4400, IRIX 6.5.22, lightning,rex-jit,jitv2, 4 compile threads),
    best of six runs of your bench binaries: whetstone 100000 6.233 s with
    the hot lane, 6.310 s without; dhrystone 10000000 8.803 s with, 8.775 s
    without. The same within noise on this host (M4 Pro, 16 KB pages); your
    Threadripper with 16 threads is the more interesting measurement.

🤖 Generated with Claude Code

A page whose compile request was already queued took a second branch on
every arrival: count the arrival, and past 256 put a copy of the request
on a hot lane the workers drained first. It was built for the first
program after boot on a machine with few compile threads, waiting behind
a long queue. With more threads it buys nothing, and it put an atomic
add on the interpreted path of every page waiting for its compile.

Now scheduling is one test: if this arrival wins `try_schedule_page`,
push the request (and clear the flag if the queue is full); either way,
interpret. HOT_QUEUE, push_hot_request, note_waiting_arrival, the page's
hot_wait and promoted fields and their resets, and the worker's hot-first
pop are gone.

Measured on an Indy (R4400, IRIX 6.5.22, lightning,rex-jit,jitv2, 4
compile threads), best of six runs: whetstone 100000 6.233 s with the hot
lane, 6.310 s without; dhrystone 10000000 8.803 s with, 8.775 s without --
the same within noise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@techomancer

Copy link
Copy Markdown
Owner

ok i found the problem that caused regression, defaulting to r5k in iris.toml. which is slower.
but i do haves some ideas about maybe simplifying the compile dispatch with single atomic counter.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants