Skip to content

server: /train claims a run and shows it 'preparing' before rendering its examples, so a timed-out caller reconciles by the run's id - #55

Merged
joelteply merged 1 commit into
masterfrom
train/ack-before-prepare
Oct 11, 2026
Merged

joelteply merged 1 commit into
masterfrom
train/ack-before-prepare

Conversation

@joelteply

Copy link
Copy Markdown

Why

On the 5090 (continuum, 2026-10-11 10:08Z), /train rendered and tokenized all of a run's examples, ~13 s per lived turn beside serving, before it claimed the run or set its state. The core's POST timed out at 120 s during that prepare. GET /train still showed the previous run, so the core failed job 97e081e2, the engine then trained it, and nothing adopted the gene (continuum card 93ffc555). Kimi noticed it herself: her own job record says failed while her losses come in.

What

  • The cheap checks (out, repacked weights, parse_special) run first. Then the run is claimed (the running CAS) and GET /train reports {state: "preparing", out} straight away, before the prepare.
  • Every refusal during the prepare goes through one refuse(). It releases the claim and leaves that run's error under its out, so a caller reconciling by id sees the engine's refusal of its own run, not silence.
  • The later lock section just moves preparing to starting.

Measured

On BigMama (Windows, CUDA build at 6416304, Qwen3.5-0.8B on the 5090, 2026-10-11 ~10:30Z):

  • POST /train with one example of 3,016 tokens into a 256 window, no fit, returns the refusal. GET /train then reports {state: error, out: ack-test.gguf} with that reason. Before this change it reported the previous run.
  • A second POST straight after is accepted (starting, out ack-test2.gguf), so the refusal released the claim.
  • test-chat: all pass.

The preparing window itself is only observable with a slow prepare. The core half (continuum, same card) reconciles a POST transport error by GET /train's out; its receipt is the 5090 run that follows.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc

… its examples, so a timed-out caller reconciles by the run's id

The prepare (render and tokenize every example) ran before the run was claimed or its state
set, ~13 s per lived turn beside serving on the 5090 (continuum, 2026-10-11). A caller whose POST
timed out during it saw the previous run's state, failed a job the engine then trained (job
97e081e2), and nothing adopted the gene (continuum card 93ffc555). Now the cheap checks (out,
repacked weights, parse_special) run first, the run is claimed and GET /train reports
{state: preparing, out} at once, and every refusal during the prepare releases the claim and
leaves that run's error under its out.

@joelteply joelteply left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved at 6416304, pending green. Claim verified, by reading the diff: the run is claimed (running CAS) and its state set to 'preparing' with its out before prepare_examples, and every refusal after the claim goes through refuse(), which writes state=error under the same out and releases running. One ordering nit, non-blocking: worker.join() sits between the claim and state=preparing, so for that instant GET reports the previous run's state; setting state before the join closes it. The core side has a blocker, on ggml-org#4946.

@joelteply
joelteply merged commit 243f3a4 into master Oct 11, 2026
10 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant