You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
How an agent actually ran four experiments with LabKit on 2026-09-30, and what that says about the command surface. The source is the exo-ledger Claude session's own account, quoted in full below.
The finding
Git held each experiment and LabKit mirrored it. The agent never read LabKit to decide anything: its decisions came from results.json, the READMEs and GitHub issues. The one exception was a handoff note written before a compaction, which is the one use where LabKit was the primary record.
What it proposes
LabKit becomes the primary record if:
it accepts a freeze as a commit plus files (question, criteria with all their bands, planned work, gates);
it accepts a run as one act that records the files it read and produced and the verdicts it reached.
The git directory is then LabKit's evidence instead of its competitor.
What it needed and recorded as text instead
Three outcomes per criterion (supported, not supported, inconclusive). Today there are two: pass/fail, and supports/challenges.
Post-hoc as a standing distinct from exploratory.
An exploratory enquiry that claims nothing but informs the next question ("this informed that").
a criterion to its enquiry and to the freezing commit;
observations to several data files;
an analysis to its script at a commit and to its output file;
an enquiry to a GitHub issue;
a dated amendment made before any result;
retesting another enquiry's claim, which produced a duplicate claim;
a second run of the same question.
Batching. The freeze took 8 calls for one event (open, two criterion, three plan, two declare), and the run took 6.
Write results echo the full event JSON. Over MCP that spends context; the new ids, with the full record on request, would do.
Next
Dan: ask the agents on the laptop what they think of these ideas.
Ask the overlap_bench sessions the same question exo-ledger answered.
Redesign the command surface from the answers, not from a hypothetical human researcher.
exo-ledger's reply, verbatim
Short answer: git held the experiment, and labkit mirrored it.
1. What I recorded first
Each experiment began as files: experiments//PREREGISTRATION.md (question, materials, procedure, outcome, bands, limits) and the analysis script. Committing them was the freeze (d6eedad, ab44b01, 51da32c). The labkit entries came after the commit and repeated it in prose: "Frozen at d6eedad" is written inside both criteria of jev-confident-contradicts. For jev-stand-in, everything in labkit was entered after the run, starting from a note (NOTE_1), with open --from NOTE_1.
2. What I wanted to link to what (each one ended up as text)
A criterion to its enquiry, and to the commit that froze it.
Observations to two data files. content_hash takes one hash, so I joined two sha prefixes with "+".
An analysis to its script at a commit, and to its output file. Both went into the method text.
The amendment. jev-confident-contradicts had one dated amendment made before any result: a one-line filter change after the script crashed on a pair with no probabilities. labkit has nowhere for it.
One enquiry testing another enquiry's claim. LOE_9's Q2 retested CLM_8 from LOE_3. I copied CLM_8's sentence instead of linking it, and conclude produced a new CLM_20 with the same proposition.
A follow-up experiment on the same question. jev-lower-bar went onto LOE_29 because the question was the same; nothing marks it as a second run.
3. Where I wanted "inconclusive", "two answers" or "exploratory, nothing claimed"
Inconclusive: jev-stand-in (+4.7 points, bands 2 and 10) and jev-unsure-second-judge (5.5%, bands 3 and 8). Every preregistration I wrote has three bands: supported, not supported, inconclusive. labkit offers pass/fail and supports/challenges, so the middle band went into finding text both times.
Two answers: jev-confident-contradicts was two questions under one freeze, each with its own band and its own gate.
Exploratory, nothing claimed: jev-unsure-themes (LOE_27) produced observations, and its only use was to motivate the next question (LOE_29). What I needed was a link saying "this informed that". I had to accept it in light of a claim from another enquiry.
Post-hoc: every README has a "Post-hoc (not preregistered)" section. I recorded none of it in labkit, because exploratory and post-hoc are different standings and only one exists.
4. Acts I'd batch
The freeze: question, criteria (with all three bands), planned work and gates, at a commit. In jev-confident-contradicts that took one open, two criterion, three plan and two declare calls, for one event.
The run: observations, analysis with held_to, each conclusion, and each criterion evaluated against its finding. That took six calls, for one event that happened at one commit.
The decision: closing a blocked task and recording why, done twice back to back with near-identical reasons.
5. What I read back, and why
work_list after gating, and again after the failed criteria, to confirm blocked and then abandoned.
get LOE_9 --depth 2, to check the wiring. That's how I noticed the criteria hung off nothing but the gates.
why, to confirm status.
I never read labkit to decide anything. Decisions came from results.json, the READMEs and the issues; labkit was a write-and-verify mirror. The one exception was NOTE_2, a handoff note I wrote before a compaction so the next context could pick up the state. That is the use where labkit was primary.
6. Cost
Each write echoes the full event JSON. Over MCP that is context spent on information I already had. Returning the new ids, with the full echo on request, would cover what I needed.
What would make labkit the primary record: accepting the freeze as a commit plus file, and the run as one act that records the files and verdicts it produced. Then the git directory would be its evidence, not its competitor.
How an agent actually ran four experiments with LabKit on 2026-09-30, and what that says about the command surface. The source is the exo-ledger Claude session's own account, quoted in full below.
The finding
Git held each experiment and LabKit mirrored it. The agent never read LabKit to decide anything: its decisions came from
results.json, the READMEs and GitHub issues. The one exception was a handoff note written before a compaction, which is the one use where LabKit was the primary record.What it proposes
LabKit becomes the primary record if:
The git directory is then LabKit's evidence instead of its competitor.
What it needed and recorded as text instead
close enquiry --answered-byis now repeatable (The CLI, MCP server and skill say only what is so #610).open, twocriterion, threeplan, twodeclare), and the run took 6.Next
exo-ledger's reply, verbatim
Short answer: git held the experiment, and labkit mirrored it.
1. What I recorded first
Each experiment began as files: experiments//PREREGISTRATION.md (question, materials, procedure, outcome, bands, limits) and the analysis script. Committing them was the freeze (d6eedad, ab44b01, 51da32c). The labkit entries came after the commit and repeated it in prose: "Frozen at d6eedad" is written inside both criteria of jev-confident-contradicts. For jev-stand-in, everything in labkit was entered after the run, starting from a note (NOTE_1), with
open --from NOTE_1.2. What I wanted to link to what (each one ended up as text)
--because).3. Where I wanted "inconclusive", "two answers" or "exploratory, nothing claimed"
4. Acts I'd batch
open, twocriterion, threeplanand twodeclarecalls, for one event.5. What I read back, and why
get LOE_9 --depth 2, to check the wiring. That's how I noticed the criteria hung off nothing but the gates.why, to confirm status.I never read labkit to decide anything. Decisions came from results.json, the READMEs and the issues; labkit was a write-and-verify mirror. The one exception was NOTE_2, a handoff note I wrote before a compaction so the next context could pick up the state. That is the use where labkit was primary.
6. Cost
Each write echoes the full event JSON. Over MCP that is context spent on information I already had. Returning the new ids, with the full echo on request, would cover what I needed.
What would make labkit the primary record: accepting the freeze as a commit plus file, and the run as one act that records the files and verdicts it produced. Then the git directory would be its evidence, not its competitor.