Skip to content

Establish a Wright effectiveness benchmark against docs and skills #466

Description

@Teakowa

Goal

Establish a controlled benchmark that measures whether Wright improves coding-agent capability and efficiency compared with prompt/document/skill-based knowledge delivery.

The benchmark should answer:

Does Wright let a general coding agent complete realistic Workshop work more reliably and with lower context/token cost than progressively disclosed documentation or a Workshop-specific skill alone?

This is a product-effectiveness benchmark. The model is a controlled variable; Wright and other knowledge/tooling interventions are the variables under test.

Context

#414 established the product-level agent benchmark harness, and the follow-up benchmark work added condition matrices, deterministic grading, usage/context metrics, paired comparison, train/test splits, and oracle-backed scenarios.

Those mechanisms are sufficient to evaluate a separate product question: whether Wright's discoverable semantic tooling is more effective than supplying equivalent domain knowledge through conventional documentation or skills.

The comparison must not collapse into a generic model leaderboard.

Scope

Workshop and OverPy are separate language tracks and must be evaluated separately. Do not pool their scenarios into one effectiveness rate or one lift number.

Each language track has its own scenario set, paired runs, rates, confidence intervals, and intervention deltas. Cross-language reporting may place the tracks side by side, but must not average them into a single headline result.

Define comparable benchmark conditions within each language over the same scenarios, models, trial identities, and graders.

At minimum, compare:

baseline
docs / progressive knowledge
skill
Wright
Wright + skill

The exact delivery mechanism for the docs condition must be pinned and reproducible. The task prompt itself must remain unchanged across paired conditions.

Measure task outcome and efficiency separately.

Primary effectiveness measures should include:

  • usable / passed rate;
  • paired Wright lift relative to baseline;
  • paired Wright lift relative to docs;
  • paired Wright lift relative to skill.

Secondary efficiency measures should include where available:

  • input/output/cache/reasoning tokens;
  • tokens per usable result;
  • peak context and context-window share;
  • turns to first valid result;
  • correction rounds;
  • elapsed time;
  • Wright/tool friction and failed tool use.

Use confidence intervals and paired comparisons rather than interpreting small raw deltas as meaningful.

Train/test separation must be preserved for any guide, skill, or documentation tuning. Improvements that only raise train performance must not be treated as product improvements.

Benchmark design

The benchmark should preserve the existing principles from #414:

  • deterministic checks remain the grading authority where possible;
  • owner-level failures remain attributable to their owning repository;
  • runtime-only claims remain explicit;
  • grader changes require regrading comparable runs;
  • stochastic results require repeated trials;
  • scenario prompts do not leak Wright commands, Workshop APIs, or solution hints.

The benchmark should make it possible to compare interventions without changing the model or task.

Conceptually:

same model
same scenario
same trial identity
same grader
        ↓
different knowledge/tooling intervention
        ↓
paired capability + cost delta

Reporting

The benchmark should report product-effectiveness deltas rather than one composite score, and report them separately for Workshop and OverPy.

Example shape:

Wright lift vs baseline     +X pp
Wright lift vs docs         +Y pp
Wright lift vs skill        +Z pp

input tokens                -A%
peak context                -B%
turns to usable             -C
tokens / usable             -D%

Do not combine correctness, token cost, context use, and tool friction into an arbitrary weighted score.

Non-goals

  • Ranking models.
  • Defining a general-purpose coding-agent benchmark.
  • Treating documentation or skills as inherently inferior without measured results.
  • Tuning benchmark graders to favor Wright-specific behavior.
  • Giving Wright credit for owner-engine capability it does not provide.
  • Replacing owner-level tests or compatibility suites.
  • Creating a parallel verification framework outside the existing benchmark harness.

Acceptance criteria

  • Workshop and OverPy are separate benchmark tracks with separate scenario sets, rates, confidence intervals, and lift results.
  • No headline effectiveness metric pools Workshop and OverPy outcomes.
  • The benchmark defines reproducible baseline, docs, skill, Wright, and Wright+skill conditions within each language track.
  • Conditions keep the task, model, grader, scenario, and paired trial identity fixed.
  • Docs and skill inputs are pinned or hashed so comparisons are reproducible.
  • Reports include usable/pass rates with confidence intervals.
  • Reports include paired lift for Wright vs baseline, docs, and skill where both runs are valid.
  • Reports include token and context efficiency metrics where the adapter provides them.
  • Reports include turns/corrections/tool-friction metrics where available.
  • Train/test separation prevents guide/docs/skill tuning from being evaluated only on tuned scenarios.
  • The result can distinguish a correctness improvement from a cost/context improvement.
  • The benchmark remains focused on Wright product effectiveness, not model ranking.

Dependencies / ownership

Wright owns the benchmark harness, product intervention surfaces, reporting, and agent/tooling integration.

Language owners remain authoritative for Workshop, OverPy, and DEL/OSTW semantics.

Related work: #414 and the benchmark comparison/harness work introduced in #459-#462.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions