feat(bench): add raw Workshop scenarios and a Tier 1 matrix example - #462
Merged
Merged
Conversation
Refs #414. Adds ana-paintball-ws, widow-headshots-ws, and impossible-request-ws: the same requirements as the OPY scenarios, written as raw Workshop script, with references and negatives compiled from the OPY ones by the pinned upstream compiler and graded by Wright and workshop-rs. Assigns train/test splits to the existing Workshop scenarios so a requirement keeps one split across languages, and adds matrix.example.json.
e54-bot
force-pushed
the
feat/414-workshop-scenarios
branch
from
September 30, 2026 16:17
c054e30 to
df76199
Compare
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #414. Builds on #461 (merged).
The benchmark now covers raw Workshop as well as OverPy, so guide tuning is not fitted to one language.
ana-paintball-ws,widow-headshots-ws,impossible-request-ws. They carry the same requirements as their OPY counterparts, with the prompt asking for raw Workshop script inmode.ws. Their references and negatives are the OPY ones compiled by the pinned upstream compiler, so the Workshop text is upstream-derived. They are graded bywright check,wright compile(workshop-rslayer), Wright's emitted text for the requirement regexes, andlint. There is no upstream oracle verdict on a raw Workshop entry, and theoraclecheck is not used for them.train/testsplits. A requirement or family keeps one split across languages (anatrain,widowandimpossibletest, and so on), so a held-out scenario is never a translation of a training one. Result: 6 train and 7 test scenarios, with every family covered in both languages except greenfield (train: ana in both languages; test: widow in both plusgreenfield-elimination-race) and refusal (test only).matrix.example.json: the Tier 1 matrix (none,bin,bin+skill, web, wiki), for whoever runs the multi-model evaluation.Verification:
agent_bench.py validatewith the oracle installed: all 13 scenarios pass calibration (reference passes, seed fails, each negative fails exactly its declared checks). Unit tests: 19 pass.Limits: 13 scenarios is still a small set for tuning, and the Workshop greenfield references are compiled from agent-adapted OPY, not verified in the game (
referenceNote).