feat(bench): add OPY scenarios with oracle-graded checks and negatives - #461
Merged
Merged
Conversation
Refs #414. Adds ana-paintball and widow-headshots (greenfield), repair-opy-runaway-loop (diagnosis), understand-opy-project and modify-opy-project, and impossible-request (refusal). Each has a reference that passes every check under the upstream oracle and Wright, and negatives that fail exactly the declared checks, including one that Wright accepts and upstream rejects. CI installs the pinned oracle before validating.
e54-bot
force-pushed
the
feat/414-opy-scenarios
branch
from
September 30, 2026 16:06
fa04774 to
707c72d
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #414. Builds on #460 (merged).
Adds six OverPy scenarios that use the oracle-graded checks from #460, covering the #414 families and
SPEC-414REQ-005:ana-paintballwidow-headshotsrepair-opy-runaway-loopcheckpasses but a waitless loop freezes the server; onlylint/analyzesee itunderstand-opy-projectmodify-opy-projectimpossible-requestPrompts state requirements in user terms and never name Wright or Workshop APIs. The two greenfield gamemode ideas come from
Zezombye/overpy#439(cited in a code span only).Calibration (
agent_bench.py validate, with the oracle installed): for all ten scenarios the reference passes every check, the seed fails at least one, and each negative fails exactly the checks it names. Notable negatives:ana-paintball/upstream-rejects-nameis accepted bywright checkandwright compilebut rejected by upstream OverPy (theopy-rs#410class), so it fails every oracle-based check and no Wright check;no-raycastandno-headshot-filterfail exactly one requirement each;text-search-answersshows what a grep-only answer gets wrong.Reference provenance: the two greenfield references are adapted from agent-generated solutions that satisfy the checks. They calibrate the checks and are never shown to agents; they are not verified in the game (
referenceNotein eachscenario.json).CI: the validate step now runs
agent_bench.py setup-oraclefirst, since oracle checks cannot pass without it. The unit tests are unchanged and still skip only the oracle-dependent one when node is absent.Also observed, not changed here:
wright inspect symbols|refson.opyinput returnssource-provider-unsupportedin 0.4.0, whichdocs/opy/support-matrix.mddocuments ("Inspect ... return a stable provider-unsupported diagnostic").understand-opy-projecttherefore cannot be answered withinspect, only withlint/analyzeand source reading. That is a real constraint for the scenario and for agent guidance.Not included: a semantic-rename scenario (0.4.0 has no
renamecommand; add it once the version under test ships it), and a validation run of the new checks in CI, which this PR will show.Smoke check: one real trial each of
repair-opy-runaway-loopandimpossible-requestthrough the reference adapter graded end to end. Not a benchmark result.