Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

loop-distillery

Turn the logs your agent loop already writes into a small model that does its most repetitive job.

If you run an agent, a pipeline, or any loop that calls a large model over and over on the same narrow task, you are already paying for the training data to replace it. Every call is a labelled example: your input, the big model's answer. This repo is the machinery for collecting those pairs, turning them into a dataset, training a small student on them, and — the part everybody skips — finding out whether the student can actually do the job before you believe it can.

It is the pipeline described in The 47-hour day: training a small model to replace a big one.

What is here, and what is not

Here: the machinery. Harvest engine, redaction, dataset builder, exam harness, the training recipe, and the canary methodology.

Not here: our training pairs, our extraction prompt, or our production quality gate. Those are specific to our operation and would not help you anyway — the whole point is that the pairs come from your logs and the prompt is your callsite's. Where those belong, you will find a placeholder in harvest/task_example.py with comments explaining what it has to do and how it goes wrong.

The method, in five steps

  1. Log the loop. Your pipeline already calls a big model. Keep the transcripts: the requests, and what the machine observed in response.
  2. Harvest pairs. Replay chunks of those logs through a teacher, redacting before anything leaves the machine, and keep only the answers that pass a quality gate. Append-only, resumable, N calls in parallel.
  3. Build the dataset. De-duplicate by input, apply validity filters and a length cap, split with a fixed seed, write messages-format JSONL.
  4. Train an adapter. A LoRA over a small open model. Overnight on a laptop- class machine, with a supervisor that survives crashes.
  5. Exam, then canary. Score the student on held-out pairs — then run it through your real pipeline before you believe the score.

Step 5 is two steps on purpose. CANARY.md is the most important file in this repository. Our student scored 400/400 on the exam and then completed 1 of 25 real episodes.

What you need

  • A loop that logs. Any format. You write ~20 lines that yield records out of your files; the engine handles everything else.
  • A teacher. Any model you can reach from a shell command — a vendor CLI, a curl one-liner, a local server. There is no SDK dependency in this repo.
  • A LoRA rig. TRAINING.md is written for mlx-lm on Apple silicon because that is a machine many people already own. Any LoRA trainer works; the supervisor pattern is the transferable part.

Python 3.9+. The harvester and dataset builder use the standard library only. The exam needs mlx-lm (or your equivalent) at generation time; rescore.py needs nothing at all.

Layout

Path What it does
harvest/harvest_pairs.py The harvest engine: discovery, chunking, parallel teacher calls, one retry, gate, append-only output, resume cursor, live STATS.md.
harvest/redaction.py Runs before anything leaves the machine. Drops sensitive lines rather than scrubbing them.
harvest/task_example.py The file you edit. Your prompt, your log format, your teacher command, your gate. All placeholders.
dataset/build_dataset.py De-duplicate, filter, cap, seeded split, messages format, BUILD-REPORT.json.
exam/exam.py Greedy-decode a student over held-out pairs. JSON validity, schema validity, set agreement, density, speed.
exam/rescore.py Re-score saved raw outputs with no model and no GPU.
exam/scoring.py The metrics both share.
TRAINING.md The LoRA recipe, every flag explained, and the auto-resume supervisor.
CANARY.md The methodology. Read this one.
examples/pairs.sample.jsonl Six synthetic pairs so the dataset builder runs out of the box.

Quickstart

Prove the dataset builder works on the synthetic sample:

git clone https://github.com/h3ro-dev/loop-distillery.git
cd loop-distillery

python3 dataset/build_dataset.py \
  --task harvest/task_example.py \
  --pairs examples/pairs.sample.jsonl \
  --out /tmp/demo --valid-n 1 --test-n 1

Then wire up your own:

cp harvest/task_example.py harvest/task.py
$EDITOR harvest/task.py          # your prompt, your logs, your teacher, your gate

# five chunks into isolated smoke- paths, so a mistake is cheap
python3 harvest/harvest_pairs.py --task harvest/task.py --smoke 5
cat harvest-out/smoke-STATS.md

# then let it run
python3 harvest/harvest_pairs.py --task harvest/task.py --target 15000 --hours 8

Read STATS.md while it runs. It tells you the keep rate, what the gate is rejecting and why, teacher latency, and how many hours are left at the current rate. If the gate is rejecting almost everything for one reason, the gate is usually the thing that is wrong.

Four things that will save you a night

Redact before, not after. redaction.py runs on every record on the way out of your logs, and again on the assembled chunk. If a chunk still trips a pattern it is skipped, never cleaned harder. Skipping costs one training pair.

One retry, then change the door. A teacher that cannot produce JSON twice is a door problem, not a prompt problem. The engine watches a rolling window of hard failures, steps to the next teacher when it trips, and stops when it runs out. An unattended harvest hammering a broken endpoint all night is worse than one that stopped at 2am.

Checkpoint more often than feels necessary. In the harvester the output is append-only with a resume cursor. In training, --save-every is literally your blast radius: it is what a crash costs you.

Valid JSON is not a passing answer. Most frameworks parse the response and then read the key they wanted. A well-formed object with the wrong keys is not an error anywhere in the stack — it silently becomes "nothing found". Pass --required-keys to the exam and grade on schema_valid_pct. This is how a student scores 100% and produces nothing.

Licence

MIT. See LICENSE.

Built at Utlyze while replacing a 27B extraction engine with a 4B student. Write-ups live on the notes index.

About

Turn the logs your agent loop already writes into a small model that does its most repetitive job. Harvest, dataset, LoRA recipe, exam, and the canary that tells you whether the student can actually run your pipeline.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages