Skip to content

feat(bench): add condition matrix, tracing, oracle grading, and report to the agent benchmark - #460

Merged
Teakowa merged 1 commit into
mainfrom
feat/414-benchmark-harness
Sep 30, 2026
Merged

Teakowa merged 1 commit into
mainfrom
feat/414-benchmark-harness

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Refs #414. Implements the harness side of SPEC-414 (merged as #459). The spec is not complete: see "Not included".

What it adds (benchmarks/agent, split into small modules beside agent_bench.py):

  • Conditions: run --wright-level none|bin|bin+skill --knowledge none|wiki|web --network off|on, and matrix (agents x cells x scenarios x trials, shuffled by a seed, resumable, infrastructure exit 75 retried). Each run gets a fresh HOME, an allowlisted environment, and canaries; a failed canary marks the run invalid.
  • Adapter contract: the cell reaches the agent as BENCH_* variables; the adapter reports per-turn usage, a transcript, and its loaded context. adapters/claude_code.py is a reference adapter. A pi adapter only needs to implement the same three files.
  • Grading: validity per authority (wright check, wright compile, emitted Workshop text, and the pinned upstream OverPy 9.7.10 oracle under oracle/), with disagreement output as an owner-issue reproducer. New check kinds: oracle, wright-compile, compiled-contains/compiled-absent. validate also calibrates negative/ overlays. Results carry a grader hash.
  • Trace: the shim records envelope fields (ok, exit, diagnostic codes, input_identity, selection) and tees wright serve sessions; the entry file is snapshotted after each write and every snapshot is graded, giving first-valid index and regressions.
  • Analysis: expectation detectors E01-E04, E06, E08, E11, E12 and friction counters; token, context and to-first-valid metrics from adapter usage; estimated Wright output tokens per command. E05, E07, E09, E10 report unavailable until a normalized transcript exists.
  • Report: report [--regrade] gives usable rates with Wilson intervals, tokens per usable result, paired comparison (tokens only where both are usable), per-scenario/split tables, expectations, friction, and eval diagnostics from the automated eval-design article (headroom, variance, infrastructure failures, grader consistency).
  • Result contract is now wright-agent-bench/v2; the old baseline/wright conditions are replaced by the levels above. docs/agent-benchmark.md and the spec (REQ-025 to REQ-027, Q-004) are updated.

Verification

  • python3 -m unittest discover -s benchmarks/agent: 19 tests pass against a wright 0.4.0 binary, including the real oracle after setup-oracle; without node they skip the oracle test only. validate passes for all four scenarios.
  • One real end-to-end trial through the Claude Code adapter (bin+skill, understand-score-flow): usage, loaded context, and the guide were recorded, and the report rendered. That is a smoke check, not a benchmark result.
  • CI runs validate and unittest as before; neither needs node.

Not included

  • New OPY scenarios (ana-paintball and the others in REQ-005), so no scenario uses oracle or compiled-contains yet. They need reference solutions and negatives first.
  • The guide-tuning loop (REQ-027).
  • The offline network sandbox: the harness runs the canary and invalidates on failure, but enforcement is the adapter's job, and the reference adapter only removes web tools.
  • Transcript normalization for E05, E07, E09, E10.

…t to the agent benchmark

Refs #414. Implements the harness parts of SPEC-414: wright/knowledge/network cells with canaries and a scrubbed environment, a matrix runner with retries, the wright shim with envelope and serve tracing, entry snapshots, expectation detectors and friction metrics, per-turn usage and context reporting, upstream-oracle grading with disagreement output, negative-overlay calibration, and a report with intervals, paired comparison, and eval diagnostics. Adds a Claude Code reference adapter. Scenarios for the new OPY requirements and the guide-tuning loop are not included.
@Teakowa
Teakowa merged commit 905ee94 into main Sep 30, 2026
15 checks passed
@Teakowa
Teakowa deleted the feat/414-benchmark-harness branch September 30, 2026 16:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants