Skip to content

Proposal: AgenticReplayBackend — replay mined tasks with a real agentic CLI in a throwaway worktree, gate on the diff #155

Description

Motivation

Replay currently scores a single text completion: the candidate skill text + memory + task description go in as one prompt, one block of text comes out, and it is checked programmatically. (The only tool access is the small sanctioned attempt_with_tools loop for tool_called checks.)

That is honest for text-expressible behaviors — formatting, structure, phrasing, conventions — but it structurally cannot evaluate agentic behaviors: "run the tests, read the failure, fix the config" needs an environment that reacts. Much of what users actually want improved in coding agents lives in that multi-step space, and it also connects to #154: agentic tasks are exactly the ones whose success can't be expressed as a regex.

Proposal: AgenticReplayBackend

A replay backend that drives a real open agentic CLI, following the existing backend/attempt pattern:

  1. Give the CLI the mined task in a throwaway git worktree/branch (sandboxed, disposable).
  2. Let it actually run tools and edit files.
  3. Gate on the diff rather than on output text: does it build? do tests pass? does the diff resemble the change the user eventually accepted in the original session?
  4. Surface the diff in the report/dashboard for human review.

This gives real ground truth with no simulator-fidelity questions. (The heavier alternative — a language world model like Qwen-AgentWorld simulating environment feedback — is interesting for offline multi-step replay but adds a large unvalidated link to the evidence chain; probably a research direction rather than a v-next feature.)

Cost alignment

An obvious objection: SkillOpt's premise is cheap, token-efficient improvement of frozen agents, and a real agentic replay is orders of magnitude heavier than a single text completion (full CLI session, tool calls, builds, test runs). To be clear, this proposal is not "make every replay agentic." It's a tiered gate:

  1. Cheap text replay stays the default for everything text-expressible — no change for the vast majority of candidates.
  2. Agentic replay is invoked only for the candidate-rule category that cheap replay structurally cannot validate (the agentic behaviors from Mining is limited to programmatic success checks — intent-level improvements get dropped or Goodharted into shallow proxy rules #154), and only for candidates that have already survived the cheap filters.
  3. It runs under an explicit per-night budget (e.g., at most N agentic replays per night, token-capped), opt-in per user.

Two further points on alignment with the original premise:

  • SkillOpt's cheapness claim is fundamentally about the deployed artifact: a small text file improving a frozen model at zero inference-time cost. This proposal doesn't touch that — the skill that comes out is exactly as cheap to deploy. The added cost is one-time validation expense, amortized over every future session, which is consistent with the existing design (the optimizer itself already uses a frontier model).
  • The pitch is to spend expensive validation only where cheap validation is worthless. A wrong-but-accepted agentic rule costs more — in degraded future sessions — than the validation would have. An unvalidated rule that passes a regex gate but breaks real workflows is the truly expensive outcome.

Bigger picture

This is one of the key steps toward SkillOpt functioning as on-the-job training for a specific job role: learning from real work, evaluated on real work.

Happy to help spec this out if there's maintainer interest.

Activity

  1. Yif-Yang commented on Jul 21, 2026

    @Yif-Yang
    Contributor

    Thanks for the thoughtful proposal. We agree that there is a real gap between text-only replay and the multi-step behaviors users often want coding agents to improve. A tiered design — keeping cheap text replay as the default and invoking agentic replay only for candidates that require an interactive environment — is a sensible direction to explore.

    The most important boundary is that a throwaway Git worktree is disposable version-control state, but it is not a security sandbox. A real agentic replay backend would still need isolation from the host, user credentials and configuration, unrelated files, network access, and shared Git metadata. It would also need hard limits on turns, time, processes, concurrency, and provider spend, with reliable cleanup after every outcome.

    We would also treat task verification as the primary gate rather than diff similarity. Builds, tests, task-specific verifiers, allowed-file boundaries, and regression checks provide stronger evidence of correctness. Similarity to the change accepted in the original session can be useful context for a judge or human reviewer, but it should remain secondary because multiple substantially different diffs may solve the same task correctly.

    We are interested in exploring this direction and would welcome a community design or draft PR. A safe first slice could be:

    • fully opt-in and limited to one agent backend and one narrow task category;
    • a pinned starting commit and reproducible fixture/environment;
    • execution inside a real container or OS sandbox with isolated home/configuration and network disabled by default;
    • deterministic tests or a task verifier as the main acceptance signal;
    • explicit token, turn, time, and replay-count budgets; and
    • a staged diff/report for review, without automatic adoption.

    Starting with a small end-to-end prototype would let us evaluate the evidence quality, cost, and safety model before generalizing it to arbitrary mined tasks. We would be happy to work with you or other interested contributors on refining that scope.

  2. siddhivmd commented on Jul 22, 2026

    @siddhivmd

    Hi, I'd like to help implement this. I'm newer to the codebase — could I take a first pass at a small piece, like the sandboxed worktree setup (step 1)? Happy to follow whatever pattern you'd prefer.

  3. OnourImpram commented on Jul 26, 2026

    @OnourImpram

    Agreed with the ordering in the reply above — task verification as the primary gate, similarity secondary. One thing the first slice sketched there does not yet name, and it is much cheaper to settle before the report schema exists than after.

    The acceptance signal needs three outcomes, not two.

    That slice already shrinks the problem: a pinned starting commit and a reproducible fixture remove most environment drift. It does not close it. Under the same constraints, a replay run can still end in a state that is neither pass nor fail:

    • a turn, time, or spend cap fires mid-task and the CLI never reaches a final state — and those caps exist to bound cost, not to bound tasks, so they will fire;
    • a build or test step needs a fetch, and network is disabled by default;
    • the container or fixture does not come up at the pinned commit;
    • a test flakes.

    None of that is evidence about the candidate rule, but each one still produces a run that has to be recorded as something.

    Collapsing them is bad in both directions:

    • Into fail: good candidates get rejected by infrastructure weather, and the gate's measured precision starts tracking harness health instead of rule quality. That is the exact number you need to trust later, when deciding whether the expensive tier is earning its budget.
    • Into pass: an unvalidated rule advances carrying the expensive gate's endorsement. That defeats the tiering the proposal is built on — the whole argument for spending here is that this tier certifies something the cheap one structurally cannot.

    Concretely: emit a third value — inconclusive, or whatever fits the codebase — and have the tier policy treat it as not validated yet. The candidate does not advance, the run stays eligible for retry while the nightly budget lasts, and it counts as evidence for neither side. Today that is an enum value and a branch. Once runs have accumulated under a two-valued schema, it is a migration plus a re-reading of every historical result.

    It also buys a diagnostic for free. A nightly inconclusive rate separates "the harness is degrading" from "this week's mined candidates are worse," which is otherwise hard to tell apart from outside, and it makes a mis-set budget visible: a cap that keeps firing mid-task shows up as a rising inconclusive rate rather than as a phantom drop in rule quality.

    One follow-on for the staged report, taking similarity as a secondary signal rather than a gate: the ordering that earns a reviewer's attention is verifier passed, diff diverged. Those are the runs where the agent solved the task a different way — the case the reply gives for keeping similarity secondary — and they are where a human actually learns something about the candidate rule. Surfacing them first turns the secondary signal into a review queue instead of a number nobody acts on.

    Concrete reference, if it is useful: I maintain Mergen, a small independent verifier that made this call. Its verdict set is pass / conditional_pass / fail / unverifiable; a single check returning unverifiable holds the whole decision at block, and unverifiable can never resolve to pass (mergen_supervise.py, _decision). Each check also carries an evidence class — independently_executed through agentically_inferred and unavailable — so the report records how a result was obtained, not only whether it passed.

    AI-assistance disclosure: this comment was drafted with Claude Code under my authorization, and every claim it makes about my own project was checked against the repository before posting.

  4. Yif-Yang commented on Aug 7, 2026

    @Yif-Yang
    Contributor

    Thank you — we would be very happy to have you take a first pass at this.

    A small Draft PR is a good way to start. The initial scope could focus on the disposable workspace lifecycle and its tests, while keeping agent execution disabled by default. One important distinction is that a Git worktree is disposable state, but not a security boundary, so it should not yet be treated or named as a sandbox. Actual agent execution will eventually need a container or OS-level sandbox with isolated credentials, configuration, filesystem access, and network access.

    We also agree with the suggestion above that replay outcomes should distinguish pass, fail, and inconclusive. Timeouts, budget exhaustion, fixture failures, or unavailable dependencies should not be counted as evidence that a candidate skill failed.

    Please feel free to open a Draft PR early so we can review the architecture and help keep the first implementation small and safe. Thank you both for helping move this direction forward.

  5. Linxiushen commented on Sep 17, 2026

    @Linxiushen

    On the three-outcome question: the four producers of inconclusive named above — a cap firing mid-task, a disabled network, a fixture that will not come up at the pinned commit, a flaky test — are all infrastructure weather. There is a fifth: the run completes cleanly, the verifier returns a real result, and there is still not enough of it to decide.

    We hit that one on real data. Our gate replays a candidate rule against the user's own recorded transcript in time order and returns PASS / FAIL / INSUFFICIENT, where INSUFFICIENT means fewer than three post-t0 decidable actions. One candidate came back INSUFFICIENT at 0/2: nothing timed out, nothing failed, the rule was simply scoped narrowly enough that the record held two decidable actions. Collapsing it either way is wrong in the manner described above, and no timeout or budget flag catches it, so a reason code of "below the evidence floor" belongs in the schema alongside the four weather ones.

    Against the three binding constraints, including where we are no use to you:

    • Three-valued outcome — satisfied and shipped, the INSUFFICIENT branch exercised on real data rather than only in tests.
    • Execution off by default — satisfied trivially, because our gate executes nothing: zero model calls, reading a record that already exists.
    • Sandboxing — not satisfied. We have proposer/evaluator separation and hash-chained receipts; real isolation (mount namespace or container) is not implemented and our README says so. On the constraint governing the first slice we have nothing.

    One number for the tier policy, since it decides whether the expensive tier earns its budget: the cost of a wrong-but-accepted candidate. One of ours passed every structural check, and sweeping it over the user's own record priced it at 68 legitimate interruptions across 17,842 recorded tool calls. We retired it on that number alone. Re-scoped to one project it survived the structural checks and the gate returned PASS. It has fired zero times since: 27 matches across the whole 19,878-call record, every one of them predating the correction that created the rule.

    Offline, deterministic, zero model calls; N=1 machine, 8 sessions, one rule enforced live — not a multi-user study. (Correction to my own comment: the repository link I first posted here is not public, so I have removed it rather than leave a 404.) Happy to write the INSUFFICIENT branch and its reason codes as a small draft PR if wanted, without touching the workspace-lifecycle slice.

    AI-assistance disclosure: drafted with Claude Code under my authorization; every number was re-derived on my machine before posting.

    Correction, 2026-09-18. Two errors in the paragraph above, both mine, both checkable.

    (1) The re-scoped rule did not return INSUFFICIENT — it returned PASS. The INSUFFICIENT at 0/2 is real but belongs to a different candidate entirely (a require_regex rule over Workflow scripts, after-t0 false 0/2). I grafted one rule's verdict onto another's story.

    (2) INSUFFICIENT is not "the common verdict here" — it is the rarest: across every gate verdict on this machine, PASS 4, FAIL 4, INSUFFICIENT 1. One observation out of nine.

    The substantive point survives and is why I am correcting rather than deleting: a verdict for "below the evidence floor" still belongs in the schema, and we did hit it on real data. But "common" was an overclaim, and the comment ends by saying every number was re-derived before posting, which obliges me to say so when one was not.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions