Repository navigation
Proposal: AgenticReplayBackend — replay mined tasks with a real agentic CLI in a throwaway worktree, gate on the diff #155
Description
Activity
Thanks for the thoughtful proposal. We agree that there is a real gap between text-only replay and the multi-step behaviors users often want coding agents to improve. A tiered design — keeping cheap text replay as the default and invoking agentic replay only for candidates that require an interactive environment — is a sensible direction to explore.
The most important boundary is that a throwaway Git worktree is disposable version-control state, but it is not a security sandbox. A real agentic replay backend would still need isolation from the host, user credentials and configuration, unrelated files, network access, and shared Git metadata. It would also need hard limits on turns, time, processes, concurrency, and provider spend, with reliable cleanup after every outcome.
We would also treat task verification as the primary gate rather than diff similarity. Builds, tests, task-specific verifiers, allowed-file boundaries, and regression checks provide stronger evidence of correctness. Similarity to the change accepted in the original session can be useful context for a judge or human reviewer, but it should remain secondary because multiple substantially different diffs may solve the same task correctly.
We are interested in exploring this direction and would welcome a community design or draft PR. A safe first slice could be:
- fully opt-in and limited to one agent backend and one narrow task category;
- a pinned starting commit and reproducible fixture/environment;
- execution inside a real container or OS sandbox with isolated home/configuration and network disabled by default;
- deterministic tests or a task verifier as the main acceptance signal;
- explicit token, turn, time, and replay-count budgets; and
- a staged diff/report for review, without automatic adoption.
Starting with a small end-to-end prototype would let us evaluate the evidence quality, cost, and safety model before generalizing it to arbitrary mined tasks. We would be happy to work with you or other interested contributors on refining that scope.
Hi, I'd like to help implement this. I'm newer to the codebase — could I take a first pass at a small piece, like the sandboxed worktree setup (step 1)? Happy to follow whatever pattern you'd prefer.
Agreed with the ordering in the reply above — task verification as the primary gate, similarity secondary. One thing the first slice sketched there does not yet name, and it is much cheaper to settle before the report schema exists than after.
The acceptance signal needs three outcomes, not two.
That slice already shrinks the problem: a pinned starting commit and a reproducible fixture remove most environment drift. It does not close it. Under the same constraints, a replay run can still end in a state that is neither pass nor fail:
- a turn, time, or spend cap fires mid-task and the CLI never reaches a final state — and those caps exist to bound cost, not to bound tasks, so they will fire;
- a build or test step needs a fetch, and network is disabled by default;
- the container or fixture does not come up at the pinned commit;
- a test flakes.
None of that is evidence about the candidate rule, but each one still produces a run that has to be recorded as something.
Collapsing them is bad in both directions:
- Into
fail: good candidates get rejected by infrastructure weather, and the gate's measured precision starts tracking harness health instead of rule quality. That is the exact number you need to trust later, when deciding whether the expensive tier is earning its budget. - Into
pass: an unvalidated rule advances carrying the expensive gate's endorsement. That defeats the tiering the proposal is built on — the whole argument for spending here is that this tier certifies something the cheap one structurally cannot.
Concretely: emit a third value —
inconclusive, or whatever fits the codebase — and have the tier policy treat it as not validated yet. The candidate does not advance, the run stays eligible for retry while the nightly budget lasts, and it counts as evidence for neither side. Today that is an enum value and a branch. Once runs have accumulated under a two-valued schema, it is a migration plus a re-reading of every historical result.It also buys a diagnostic for free. A nightly inconclusive rate separates "the harness is degrading" from "this week's mined candidates are worse," which is otherwise hard to tell apart from outside, and it makes a mis-set budget visible: a cap that keeps firing mid-task shows up as a rising inconclusive rate rather than as a phantom drop in rule quality.
One follow-on for the staged report, taking similarity as a secondary signal rather than a gate: the ordering that earns a reviewer's attention is verifier passed, diff diverged. Those are the runs where the agent solved the task a different way — the case the reply gives for keeping similarity secondary — and they are where a human actually learns something about the candidate rule. Surfacing them first turns the secondary signal into a review queue instead of a number nobody acts on.
Concrete reference, if it is useful: I maintain Mergen, a small independent verifier that made this call. Its verdict set is
pass/conditional_pass/fail/unverifiable; a single check returningunverifiableholds the whole decision atblock, andunverifiablecan never resolve topass(mergen_supervise.py,_decision). Each check also carries an evidence class —independently_executedthroughagentically_inferredandunavailable— so the report records how a result was obtained, not only whether it passed.AI-assistance disclosure: this comment was drafted with Claude Code under my authorization, and every claim it makes about my own project was checked against the repository before posting.
Thank you — we would be very happy to have you take a first pass at this.
A small Draft PR is a good way to start. The initial scope could focus on the disposable workspace lifecycle and its tests, while keeping agent execution disabled by default. One important distinction is that a Git worktree is disposable state, but not a security boundary, so it should not yet be treated or named as a sandbox. Actual agent execution will eventually need a container or OS-level sandbox with isolated credentials, configuration, filesystem access, and network access.
We also agree with the suggestion above that replay outcomes should distinguish
pass,fail, andinconclusive. Timeouts, budget exhaustion, fixture failures, or unavailable dependencies should not be counted as evidence that a candidate skill failed.Please feel free to open a Draft PR early so we can review the architecture and help keep the first implementation small and safe. Thank you both for helping move this direction forward.
On the three-outcome question: the four producers of
inconclusivenamed above — a cap firing mid-task, a disabled network, a fixture that will not come up at the pinned commit, a flaky test — are all infrastructure weather. There is a fifth: the run completes cleanly, the verifier returns a real result, and there is still not enough of it to decide.We hit that one on real data. Our gate replays a candidate rule against the user's own recorded transcript in time order and returns PASS / FAIL / INSUFFICIENT, where INSUFFICIENT means fewer than three post-
t0decidable actions. One candidate came back INSUFFICIENT at 0/2: nothing timed out, nothing failed, the rule was simply scoped narrowly enough that the record held two decidable actions. Collapsing it either way is wrong in the manner described above, and no timeout or budget flag catches it, so a reason code of "below the evidence floor" belongs in the schema alongside the four weather ones.Against the three binding constraints, including where we are no use to you:
- Three-valued outcome — satisfied and shipped, the INSUFFICIENT branch exercised on real data rather than only in tests.
- Execution off by default — satisfied trivially, because our gate executes nothing: zero model calls, reading a record that already exists.
- Sandboxing — not satisfied. We have proposer/evaluator separation and hash-chained receipts; real isolation (mount namespace or container) is not implemented and our README says so. On the constraint governing the first slice we have nothing.
One number for the tier policy, since it decides whether the expensive tier earns its budget: the cost of a wrong-but-accepted candidate. One of ours passed every structural check, and sweeping it over the user's own record priced it at 68 legitimate interruptions across 17,842 recorded tool calls. We retired it on that number alone. Re-scoped to one project it survived the structural checks and the gate returned PASS. It has fired zero times since: 27 matches across the whole 19,878-call record, every one of them predating the correction that created the rule.
Offline, deterministic, zero model calls; N=1 machine, 8 sessions, one rule enforced live — not a multi-user study. (Correction to my own comment: the repository link I first posted here is not public, so I have removed it rather than leave a 404.) Happy to write the INSUFFICIENT branch and its reason codes as a small draft PR if wanted, without touching the workspace-lifecycle slice.
AI-assistance disclosure: drafted with Claude Code under my authorization; every number was re-derived on my machine before posting.
Correction, 2026-09-18. Two errors in the paragraph above, both mine, both checkable.
(1) The re-scoped rule did not return INSUFFICIENT — it returned PASS. The INSUFFICIENT at 0/2 is real but belongs to a different candidate entirely (a
require_regexrule overWorkflowscripts,after-t0 false 0/2). I grafted one rule's verdict onto another's story.(2) INSUFFICIENT is not "the common verdict here" — it is the rarest: across every gate verdict on this machine, PASS 4, FAIL 4, INSUFFICIENT 1. One observation out of nine.
The substantive point survives and is why I am correcting rather than deleting: a verdict for "below the evidence floor" still belongs in the schema, and we did hit it on real data. But "common" was an overclaim, and the comment ends by saying every number was re-derived before posting, which obliges me to say so when one was not.
Motivation
Replay currently scores a single text completion: the candidate skill text + memory + task description go in as one prompt, one block of text comes out, and it is checked programmatically. (The only tool access is the small sanctioned
attempt_with_toolsloop fortool_calledchecks.)That is honest for text-expressible behaviors — formatting, structure, phrasing, conventions — but it structurally cannot evaluate agentic behaviors: "run the tests, read the failure, fix the config" needs an environment that reacts. Much of what users actually want improved in coding agents lives in that multi-step space, and it also connects to #154: agentic tasks are exactly the ones whose success can't be expressed as a regex.
Proposal: AgenticReplayBackend
A replay backend that drives a real open agentic CLI, following the existing backend/attempt pattern:
This gives real ground truth with no simulator-fidelity questions. (The heavier alternative — a language world model like Qwen-AgentWorld simulating environment feedback — is interesting for offline multi-step replay but adds a large unvalidated link to the evidence chain; probably a research direction rather than a v-next feature.)
Cost alignment
An obvious objection: SkillOpt's premise is cheap, token-efficient improvement of frozen agents, and a real agentic replay is orders of magnitude heavier than a single text completion (full CLI session, tool calls, builds, test runs). To be clear, this proposal is not "make every replay agentic." It's a tiered gate:
Two further points on alignment with the original premise:
Bigger picture
This is one of the key steps toward SkillOpt functioning as on-the-job training for a specific job role: learning from real work, evaluated on real work.
Happy to help spec this out if there's maintainer interest.