Turn the logs your agent loop already writes into a small model that does its most repetitive job.
If you run an agent, a pipeline, or any loop that calls a large model over and over on the same narrow task, you are already paying for the training data to replace it. Every call is a labelled example: your input, the big model's answer. This repo is the machinery for collecting those pairs, turning them into a dataset, training a small student on them, and — the part everybody skips — finding out whether the student can actually do the job before you believe it can.
It is the pipeline described in The 47-hour day: training a small model to replace a big one.
Here: the machinery. Harvest engine, redaction, dataset builder, exam harness, the training recipe, and the canary methodology.
Not here: our training pairs, our extraction prompt, or our production quality
gate. Those are specific to our operation and would not help you anyway — the
whole point is that the pairs come from your logs and the prompt is your
callsite's. Where those belong, you will find a placeholder in
harvest/task_example.py with comments explaining what it has to do and how it
goes wrong.
- Log the loop. Your pipeline already calls a big model. Keep the transcripts: the requests, and what the machine observed in response.
- Harvest pairs. Replay chunks of those logs through a teacher, redacting before anything leaves the machine, and keep only the answers that pass a quality gate. Append-only, resumable, N calls in parallel.
- Build the dataset. De-duplicate by input, apply validity filters and a length cap, split with a fixed seed, write messages-format JSONL.
- Train an adapter. A LoRA over a small open model. Overnight on a laptop- class machine, with a supervisor that survives crashes.
- Exam, then canary. Score the student on held-out pairs — then run it through your real pipeline before you believe the score.
Step 5 is two steps on purpose. CANARY.md is the most important file in this repository. Our student scored 400/400 on the exam and then completed 1 of 25 real episodes.
- A loop that logs. Any format. You write ~20 lines that yield records out of your files; the engine handles everything else.
- A teacher. Any model you can reach from a shell command — a vendor CLI, a curl one-liner, a local server. There is no SDK dependency in this repo.
- A LoRA rig. TRAINING.md is written for
mlx-lmon Apple silicon because that is a machine many people already own. Any LoRA trainer works; the supervisor pattern is the transferable part.
Python 3.9+. The harvester and dataset builder use the standard library only.
The exam needs mlx-lm (or your equivalent) at generation time; rescore.py
needs nothing at all.
| Path | What it does |
|---|---|
harvest/harvest_pairs.py |
The harvest engine: discovery, chunking, parallel teacher calls, one retry, gate, append-only output, resume cursor, live STATS.md. |
harvest/redaction.py |
Runs before anything leaves the machine. Drops sensitive lines rather than scrubbing them. |
harvest/task_example.py |
The file you edit. Your prompt, your log format, your teacher command, your gate. All placeholders. |
dataset/build_dataset.py |
De-duplicate, filter, cap, seeded split, messages format, BUILD-REPORT.json. |
exam/exam.py |
Greedy-decode a student over held-out pairs. JSON validity, schema validity, set agreement, density, speed. |
exam/rescore.py |
Re-score saved raw outputs with no model and no GPU. |
exam/scoring.py |
The metrics both share. |
TRAINING.md |
The LoRA recipe, every flag explained, and the auto-resume supervisor. |
CANARY.md |
The methodology. Read this one. |
examples/pairs.sample.jsonl |
Six synthetic pairs so the dataset builder runs out of the box. |
Prove the dataset builder works on the synthetic sample:
git clone https://github.com/h3ro-dev/loop-distillery.git
cd loop-distillery
python3 dataset/build_dataset.py \
--task harvest/task_example.py \
--pairs examples/pairs.sample.jsonl \
--out /tmp/demo --valid-n 1 --test-n 1Then wire up your own:
cp harvest/task_example.py harvest/task.py
$EDITOR harvest/task.py # your prompt, your logs, your teacher, your gate
# five chunks into isolated smoke- paths, so a mistake is cheap
python3 harvest/harvest_pairs.py --task harvest/task.py --smoke 5
cat harvest-out/smoke-STATS.md
# then let it run
python3 harvest/harvest_pairs.py --task harvest/task.py --target 15000 --hours 8Read STATS.md while it runs. It tells you the keep rate, what the gate is
rejecting and why, teacher latency, and how many hours are left at the current
rate. If the gate is rejecting almost everything for one reason, the gate is
usually the thing that is wrong.
Redact before, not after. redaction.py runs on every record on the way out
of your logs, and again on the assembled chunk. If a chunk still trips a pattern
it is skipped, never cleaned harder. Skipping costs one training pair.
One retry, then change the door. A teacher that cannot produce JSON twice is a door problem, not a prompt problem. The engine watches a rolling window of hard failures, steps to the next teacher when it trips, and stops when it runs out. An unattended harvest hammering a broken endpoint all night is worse than one that stopped at 2am.
Checkpoint more often than feels necessary. In the harvester the output is
append-only with a resume cursor. In training, --save-every is literally your
blast radius: it is what a crash costs you.
Valid JSON is not a passing answer. Most frameworks parse the response and
then read the key they wanted. A well-formed object with the wrong keys is not
an error anywhere in the stack — it silently becomes "nothing found". Pass
--required-keys to the exam and grade on schema_valid_pct. This is how a
student scores 100% and produces nothing.
MIT. See LICENSE.
Built at Utlyze while replacing a 27B extraction engine with a 4B student. Write-ups live on the notes index.