Skip to content

feat(bench): add pi and Devin adapters and an ancestor instruction-file canary - #464

Open
e54-bot wants to merge 1 commit into
mainfrom
feat/414-pi-devin-adapters
Open

e54-bot wants to merge 1 commit into
mainfrom
feat/414-pi-devin-adapters

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Refs #414. Makes the benchmark runnable with pi (GPT, Gemini) and the Devin CLI, and closes a context leak found while doing it.

Adapters (benchmarks/agent/adapters/), same BENCH_* contract as the Claude Code reference:

  • pi.py: pi -p --mode json. Disables context files, extensions, prompt templates, themes, and skill discovery, then loads the guide with --skill. Reads per-turn usage from message_end events, the context limit from pi --list-models, and loaded skills from the system message. Gemini needs BENCH_PI_EXTENSIONS (the pi-antigravity provider is an extension); BENCH_PI_WEB_EXTENSIONS (for example pi-web-access) is loaded only for knowledge web.
  • devin.py: devin -p with an isolated HOME that holds only the Devin credentials, a config with read_config_from off, MCP tools denied, web tools denied unless knowledge is web, and the guide installed as a workspace skill. Usage and transcript come from the ATIF export. Built-in skills are ignored; account-managed plugin skills are listed apart from loaded.

Context leak fixed. Devin loaded the user's global CLAUDE.md, ~/.agents/skills, and the repository's AGENTS.md, because agents discover instruction files by walking up from the workspace, and the default output directory was inside this repo. Now:

  • a run whose workspace has instruction files in an ancestor directory is invalid (--no-ancestor-check disables it);
  • the default --out is ~/.cache/wright-agent-bench;
  • results record networkEnforcement (canary-checked or declared-only), since pi and Devin shells can still reach the network;
  • .agents and .devin are ignored as unsafe edits, because skills are installed there through the agent's own mechanism.

Docs: docs/agent-benchmark.md states that the scenario prompt.md is the only text the harness gives an agent (no system prompt or hint added), lists the adapters, and explains how to set the model in --agent-cmd, absolute paths, and --env-pass HOME. matrix.example.json now lists pi (GPT, Gemini) and Devin.

Verification: 25 tests pass (5 new for adapter parsing, plus the ancestor canary), with the oracle installed. Real single trials, not results: pi with openai-codex/gpt-6-luna on understand-score-flow (usage, context share of the 272K window, and the guide were recorded); Devin swe-2-max on repair-runaway-loop with and without the guide (both valid, loaded context [wright] and [], and Wright used only with the guide); pi with antigravity/gemini-3.6-flash answered once with the extension loaded (adapter not run end to end for Gemini).

Not verified: network off enforcement (a curl from my session fails even without a sandbox, so I could not tell a real block from my own network); Devin --sandbox (available through BENCH_DEVIN_SANDBOX=1, untested on scenarios); the web cells for pi and Devin.

…le canary

Refs #414. Adds adapters for pi and the Devin CLI that report per-turn usage, transcript, and loaded context, with host configuration kept out of the run (no context files or extensions for pi; an isolated HOME, no cross-tool rules, and denied MCP tools for Devin). Runs now fail a canary when instruction files exist above the workspace, the default output directory moves outside the repository, and results record whether network off was checked. Updates the matrix example and the benchmark contract.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants