Repository navigation
server: /train claims a run and shows it 'preparing' before rendering its examples, so a timed-out caller reconciles by the run's id - #55
Merged
Conversation
… its examples, so a timed-out caller reconciles by the run's id
The prepare (render and tokenize every example) ran before the run was claimed or its state
set, ~13 s per lived turn beside serving on the 5090 (continuum, 2026-10-11). A caller whose POST
timed out during it saw the previous run's state, failed a job the engine then trained (job
97e081e2), and nothing adopted the gene (continuum card 93ffc555). Now the cheap checks (out,
repacked weights, parse_special) run first, the run is claimed and GET /train reports
{state: preparing, out} at once, and every refusal during the prepare releases the claim and
leaves that run's error under its out.
joelteply
commented
Oct 11, 2026
joelteply
left a comment
Author
There was a problem hiding this comment.
Approved at 6416304, pending green. Claim verified, by reading the diff: the run is claimed (running CAS) and its state set to 'preparing' with its out before prepare_examples, and every refusal after the claim goes through refuse(), which writes state=error under the same out and releases running. One ordering nit, non-blocking: worker.join() sits between the claim and state=preparing, so for that instant GET reports the previous run's state; setting state before the join closes it. The core side has a blocker, on ggml-org#4946.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
On the 5090 (continuum, 2026-10-11 10:08Z),
/trainrendered and tokenized all of a run's examples, ~13 s per lived turn beside serving, before it claimed the run or set its state. The core's POST timed out at 120 s during that prepare.GET /trainstill showed the previous run, so the core failed job 97e081e2, the engine then trained it, and nothing adopted the gene (continuum card 93ffc555). Kimi noticed it herself: her own job record says failed while her losses come in.What
out, repacked weights,parse_special) run first. Then the run is claimed (therunningCAS) andGET /trainreports{state: "preparing", out}straight away, before the prepare.refuse(). It releases the claim and leaves that run's error under itsout, so a caller reconciling by id sees the engine's refusal of its own run, not silence.preparingtostarting.Measured
On BigMama (Windows, CUDA build at 6416304, Qwen3.5-0.8B on the 5090, 2026-10-11 ~10:30Z):
/trainwith one example of 3,016 tokens into a 256 window, no fit, returns the refusal.GET /trainthen reports{state: error, out: ack-test.gguf}with that reason. Before this change it reported the previous run.starting, outack-test2.gguf), so the refusal released the claim.test-chat: all pass.The
preparingwindow itself is only observable with a slow prepare. The core half (continuum, same card) reconciles a POST transport error byGET /train'sout; its receipt is the 5090 run that follows.🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc