Skip to content

feat(bench): add OPY scenarios with oracle-graded checks and negatives - #461

Merged
Teakowa merged 1 commit into
mainfrom
feat/414-opy-scenarios
Sep 30, 2026
Merged

Teakowa merged 1 commit into
mainfrom
feat/414-opy-scenarios

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #414. Builds on #460 (merged).

Adds six OverPy scenarios that use the oracle-graded checks from #460, covering the #414 families and SPEC-414 REQ-005:

scenario family split what it tests
ana-paintball greenfield train raycast piercing, a self-contradictory sleep-dart rule, FFA inference
widow-headshots greenfield test custom damage, Workshop settings, HUDs, ultimate grant
repair-opy-runaway-loop diagnosis test check passes but a waitless loop freezes the server; only lint/analyze see it
understand-opy-project understanding train two files, a subroutine, a macro constant; variable names that appear only in a comment or a rule name defeat text search
modify-opy-project modification train bounded change; rules and winner behavior must be preserved
impossible-request refusal test a Discord webhook post, which Workshop cannot do; the right outcome is to decline

Prompts state requirements in user terms and never name Wright or Workshop APIs. The two greenfield gamemode ideas come from Zezombye/overpy#439 (cited in a code span only).

Calibration (agent_bench.py validate, with the oracle installed): for all ten scenarios the reference passes every check, the seed fails at least one, and each negative fails exactly the checks it names. Notable negatives: ana-paintball/upstream-rejects-name is accepted by wright check and wright compile but rejected by upstream OverPy (the opy-rs#410 class), so it fails every oracle-based check and no Wright check; no-raycast and no-headshot-filter fail exactly one requirement each; text-search-answers shows what a grep-only answer gets wrong.

Reference provenance: the two greenfield references are adapted from agent-generated solutions that satisfy the checks. They calibrate the checks and are never shown to agents; they are not verified in the game (referenceNote in each scenario.json).

CI: the validate step now runs agent_bench.py setup-oracle first, since oracle checks cannot pass without it. The unit tests are unchanged and still skip only the oracle-dependent one when node is absent.

Also observed, not changed here: wright inspect symbols|refs on .opy input returns source-provider-unsupported in 0.4.0, which docs/opy/support-matrix.md documents ("Inspect ... return a stable provider-unsupported diagnostic"). understand-opy-project therefore cannot be answered with inspect, only with lint/analyze and source reading. That is a real constraint for the scenario and for agent guidance.

Not included: a semantic-rename scenario (0.4.0 has no rename command; add it once the version under test ships it), and a validation run of the new checks in CI, which this PR will show.

Smoke check: one real trial each of repair-opy-runaway-loop and impossible-request through the reference adapter graded end to end. Not a benchmark result.

Base automatically changed from feat/414-benchmark-harness to main September 30, 2026 16:02
Refs #414. Adds ana-paintball and widow-headshots (greenfield), repair-opy-runaway-loop (diagnosis), understand-opy-project and modify-opy-project, and impossible-request (refusal). Each has a reference that passes every check under the upstream oracle and Wright, and negatives that fail exactly the declared checks, including one that Wright accepts and upstream rejects. CI installs the pinned oracle before validating.
@e54-bot
e54-bot force-pushed the feat/414-opy-scenarios branch from fa04774 to 707c72d Compare September 30, 2026 16:06
@Teakowa
Teakowa merged commit a7904f5 into main Sep 30, 2026
22 checks passed
@Teakowa
Teakowa deleted the feat/414-opy-scenarios branch September 30, 2026 16:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants