Skip to content

Define a standardized Wright Agent Score for model comparison #467

Description

@Teakowa

Goal

Define a stable Wright Agent Score that measures how capable a model/agent system is when operating in one standardized Wright environment.

The benchmark should answer:

Given the same Wright tool + Wright skill environment, how reliably can this model complete realistic Workshop development tasks?

This is distinct from measuring whether Wright itself improves agents. Wright is fixed; the model/agent system is the variable.

Context

#414 established the product-level agent benchmark harness. Follow-up work added deterministic graders, repeated trials, oracle-backed OverPy scenarios, raw Workshop scenarios, train/test splits, usage/context metrics, and multi-agent adapters.

That makes it possible to define a model-facing benchmark score without changing the underlying verification model.

The score should behave like a conventional public benchmark headline metric while preserving the detailed diagnostic results underneath it.

Language tracks

Workshop and OverPy are different languages and must have separate canonical scores.

The benchmark defines at least:

Wright Workshop Agent Score
Wright OPY Agent Score

Each score uses only scenarios for that language, with its own test set, usable rate, confidence interval, and detailed breakdowns.

Do not average Workshop and OPY runs into a single Wright Agent Score. Cross-language presentation may show both scores side by side, but the two tracks remain independently versioned and interpretable.

Canonical benchmark condition

Each language score must use one fixed, reproducible Wright environment.

Initial contract:

Wright       = released/pinned binary
Skill        = pinned Wright skill
Knowledge    = none beyond the standardized environment
Network      = off
Split        = test

The exact Wright version, skill identity/hash, benchmark suite version/hash, grader hash, adapter, model identity, and relevant inference configuration must be recorded with every published score.

Model-specific prompt or skill tuning is not allowed in the canonical score condition.

Execution protocol

A published score must come from a fixed, reproducible execution protocol.

For each language track:

  • the task user prompt is the exact contents of the scenario's pinned prompt.md;
  • the benchmark adds no task-specific system prompt, Wright command hint, Workshop API hint, or solution note;
  • the canonical pinned Wright skill is loaded through the adapter-defined mechanism;
  • the scenario seed workspace is the only task workspace provided to the agent;
  • network is disabled for the canonical score condition;
  • each trial starts from a fresh workspace and scrubbed environment according to the benchmark harness contract;
  • the Wright binary, skill, suite, grader, adapter, model, and inference configuration are fixed for the reported score;
  • model-specific skill or prompt variants are not allowed.

The protocol must also define and record:

  • trials per scenario;
  • maximum wall-clock duration;
  • maximum turns or equivalent interaction budget when the harness exposes one;
  • model token/output limits where configurable;
  • temperature / sampling configuration where configurable;
  • reasoning effort or equivalent inference setting;
  • infrastructure retry policy.

The score evaluates an agent-system configuration, not an abstract base model. A published result therefore identifies at least:

model
+ agent harness / adapter
+ inference configuration
+ Wright version
+ Wright skill
+ benchmark suite
+ grader

A separate model-only comparison is valid only when the same canonical agent harness and execution configuration are used across models.

Grading protocol

Each completed scenario trial produces one task-level binary outcome:

usable = true  → 1
usable = false → 0

usable must be derived from the scenario's deterministic benchmark checks and existing harness contract. It is not a model-judge rating.

A run is usable only when all required scenario checks pass and no benchmark-defined blocking condition makes the result unusable. The implementation must define blocking conditions explicitly, including how error-severity lint findings, unsafe edits, unavailable required graders, and owner-layer validity failures affect usability.

Do not award partial headline credit from the fraction of individual checks passed. Splitting one requirement into more checks must not increase its benchmark weight.

Infrastructure failures and invalid runs must not silently count as model failures. The execution contract must define retry behavior and whether a run is excluded after retries are exhausted; exclusions and infrastructure failures must be published separately.

Score

Define the headline score from task-level usability:

Wright Workshop Agent Score = 100 × usable Workshop test-run rate
Wright OPY Agent Score      = 100 × usable OPY test-run rate

For a language track with N held-out scenarios and a fixed K valid trials per scenario, compute each scenario's usable rate first:

scenario_rate(s) = usable_trials(s) / valid_trials(s)

Then macro-average scenarios so each scenario has equal weight:

Wright <Language> Agent Score
= 100 × (1 / N) × Σ scenario_rate(s)

When every scenario has the same number of valid trials this is numerically equivalent to the overall usable-run rate, but the scenario-first definition prevents unequal trial counts from changing scenario weight.

Publish a 95% confidence interval using one fixed documented method. The implementation must not switch interval methods between published runs. If bootstrap is used, resampling must preserve scenario-level clustering.

The report may also include a repeated-run reliability metric such as Pass^k, but this is secondary and must not replace the headline score.

Small score differences must not be interpreted as meaningful when confidence intervals or measured trial variance do not support that conclusion.

Example:

Wright Workshop Agent Score v1
Model X: 81.2
95% CI: [74.3, 86.6]

Wright OPY Agent Score v1
Model X: 75.6
95% CI: [68.0, 81.8]

Publication format

A published result must be self-describing and machine-readable.

At minimum, publish:

Wright <Language> Agent Score <score-version>

Agent:       <agent harness / adapter>
Model:       <model id>
Inference:   <reasoning/sampling configuration>

Score:       <0-100>
95% CI:      <lower-upper>
Trials:      <K per scenario>
Scenarios:   <N held-out scenarios>

Suite:       <version/hash>
Wright:      <version>
Skill:       <version/hash>
Grader:      <hash>
Oracle:      <version/hash when applicable>

The benchmark should emit a stable machine-readable summary alongside the human report so published tables can be reproduced from stored results rather than hand-entered.

Workshop and OPY results are published as separate score cards. They may be displayed side by side but must not be averaged into a single headline score.

Detailed reporting

The headline score must not replace the multidimensional report.

Include breakdowns where applicable by:

  • scenario family;
  • language;
  • scenario;
  • failure layer.

Also report secondary efficiency metrics separately, such as:

  • tokens per run;
  • tokens per usable result;
  • peak context and context-window share;
  • turns to first valid result;
  • elapsed time;
  • tool friction;
  • Wright invocation patterns.

Efficiency metrics must not be mixed into the correctness score through arbitrary weights.

Versioning and comparability

A published score is comparable only within the same benchmark contract.

The benchmark must version or hash at least:

  • scenario suite;
  • grader;
  • Wright release;
  • Wright skill;
  • relevant oracle/reference implementation.

Adding/removing scenarios or materially changing graders creates a new score version or otherwise makes old/new scores explicitly non-comparable.

Historical scores should identify the environment they were produced under rather than being silently reinterpreted under a new suite.

Held-out evaluation

Canonical scores use the held-out test split.

Train scenarios may be used to improve Wright, its skill, or benchmark ergonomics, but model scores must not be reported from the train split.

If the Wright skill or benchmark surface is tuned, held-out test performance determines whether the change improved the standardized environment.

Model comparison

The same canonical environment should be runnable through supported adapters for different model/agent systems.

Examples may include pi, Claude Code, Devin, or future adapters, but the score contract must not depend on one vendor.

Differences in harness capabilities that materially affect the environment must be disclosed rather than treated as model differences.

Non-goals

  • Measuring Wright's lift versus docs, skills, or no-Wright conditions; that belongs to the Wright effectiveness benchmark.
  • Combining correctness and cost into one weighted composite score.
  • Tuning a different Wright skill or prompt for each model.
  • Claiming small score differences are meaningful when confidence intervals or run variance do not support that conclusion.
  • Replacing detailed failure attribution with one number.
  • Treating one stochastic run as a model score.
  • Expanding Wright semantics merely to improve the benchmark.

Acceptance criteria

  • Workshop and OverPy have separate canonical score tracks.
  • Workshop and OPY scenarios are never pooled into one headline score.
  • Each language track has its own held-out test result and confidence interval.
  • A canonical Wright + Wright-skill benchmark condition is defined and reproducible for each language track.
  • The exact task prompt contract, skill loading, workspace isolation, network policy, trial count, budgets, retry policy, and inference configuration are defined.
  • Published results identify the model + agent harness configuration rather than implying that the score belongs to an abstract base model alone.
  • usable has one explicit deterministic definition, including blocking conditions and treatment of invalid/infrastructure runs.
  • The headline score macro-averages held-out scenarios from task-level binary usable outcomes so each scenario has equal weight.
  • The confidence-interval method is fixed and documented.
  • Published scores include a confidence interval.
  • The suite, grader, Wright version, skill identity, adapter, model, and inference configuration are recorded.
  • Benchmark/suite changes have an explicit score-versioning or comparability rule.
  • Individual check count does not determine task weight.
  • Family/language/scenario breakdowns remain available beneath the headline score.
  • Token, context, latency, reliability, and tool-use metrics remain secondary metrics rather than score weights.
  • A stable human and machine-readable publication format records score identity, environment, suite, grader, Wright, skill, and model/harness configuration.
  • Multiple model/agent adapters can run the same canonical condition.
  • The benchmark prevents per-model skill/prompt customization from contaminating score comparability.

Dependencies / ownership

Wright owns the benchmark contract, canonical Wright environment, result schema, and reporting.

Language owners remain authoritative for semantic/compiler behavior used by benchmark graders.

Related work: #414 and the benchmark comparison/harness work introduced in #459-#462.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions