Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,9 @@ acquire diff --> parse supported lockfiles --> parse + index --> bounded evidenc
binary, rename, mode, and dependency evidence under `.postil/change-metadata`.
Git C-quoted paths are decoded to canonical identities, then reversibly C-quoted
for prompt display; model citations decode back to the same forge path.
Incremental reviews also carry an uncitable view of the complete pull-request
change, without margin numbers: the raw diff up to 24 KiB, else a per-file
summary, dropped only when it does not fit the request budget.
- `filter.rs`: grounding (uncited findings dropped; all-uncited = untrusted run),
policy suppression (ignore globs, severityThreshold, minConfidence, maxFindings),
structured retention of suppressed grounded findings, and
Expand Down
25 changes: 21 additions & 4 deletions bench/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,9 +46,11 @@ Inspect the fixture IDs and source before a live run in [`fixtures/cases.ts`](fi

## Expanded clean bank

The 25-case clean bank combines the 13 admission clean cases with 12 supplemental cases in [`fixtures/clean-screen.ts`](fixtures/clean-screen.ts). The [clean-screen entrypoint](src/clean-screen.ts) passes the exported `cleanScreenCases` to the existing `runLive` API. Default screens and admission retain the 70-case corpus and its attested evaluator inputs. The supplemental cases cover authorization, expiry, asynchronous ordering, cancellation, retry limits, pagination, integer precision, defaults, lock release, SQL parameters, array ownership, and partial updates. Behavioral tests exercise both versions of each executable module, including rejection paths and boundary values.
The 27-case clean bank combines the 13 admission clean cases with 12 single-file and 2 cross-file supplemental cases in [`fixtures/clean-screen.ts`](fixtures/clean-screen.ts). The [clean-screen entrypoint](src/clean-screen.ts) passes the exported `cleanScreenCases` to the existing `runLive` API. Default screens and admission retain the 70-case corpus and its attested evaluator inputs. The supplemental cases cover authorization, expiry, asynchronous ordering, cancellation, retry limits, pagination, integer precision, defaults, lock release, SQL parameters, array ownership, and partial updates. Behavioral tests exercise both versions of each executable module, including rejection paths and boundary values.

Live diff-file screening sends the diff. Supplemental source comments therefore contain the argument and behavior contracts; the `Object.hasOwn` case also supplies package metadata declaring Node.js 22 or later. Metadata outside the diff is not evidence available to this screen.
The cross-file cases remove a target in some files and narrow a dependent alert selector, rate-limit entry, or caller in another: an infrastructure change that deletes test edge hostnames with their Traefik routes, and an application change that deletes a disabled export feature. Each hunk looks like lost coverage in isolation; the rest of the diff shows the target is gone. [`fixtures/causality-screen.ts`](fixtures/causality-screen.ts) pairs each with a must-block contrast whose narrowing also drops a target the change keeps. Tests check that every dropped selector target is deleted by the same diff, and that each contrast drops a target that remains.

Live diff-file screening sends the diff without a pull-request title or description. Supplemental source comments therefore contain the argument and behavior contracts; the `Object.hasOwn` case also supplies package metadata declaring Node.js 22 or later. Metadata outside the diff is not evidence available to this screen.

After the build and dependency setup above, select a checked-in profile:

Expand All @@ -64,12 +66,27 @@ POSTIL_LLM_REQUEST_TIMEOUT_SECS=30 POSTIL_LLM_TOTAL_TIMEOUT_SECS=60 \
timeout 780s bun run src/clean-screen.ts
```

Each invocation retains a separate report and per-case evidence. Its `supplementalScreen` field records a separate framed SHA-256 digest of the supplemental fixture module and entrypoint; `summary.evaluatorSha256` identifies the default evaluator. The process fails if every case is unavailable; partial reports retain unavailable cases for inspection. The CLI safety cap is 180 seconds per case. Report final review silence, final findings, suppressed findings with reasons, and unavailable cases separately, with the 13 legacy and 12 supplemental cases identified. A silent final review can contain suppressed model findings.
Each invocation retains a separate report and per-case evidence. Its `supplementalScreen` field records a separate framed SHA-256 digest of the supplemental fixture module and entrypoint; `summary.evaluatorSha256` identifies the default evaluator. The process fails if every case is unavailable; partial reports retain unavailable cases for inspection. The CLI safety cap is 180 seconds per case. Report final review silence, final findings, suppressed findings with reasons, and unavailable cases separately, with the 13 legacy, 12 single-file, and 2 cross-file supplemental cases identified. A silent final review can contain suppressed model findings.

The clean-bank-v2 evidence identifies the [measured fixture and evaluator source](https://github.com/postil-dev/postil-cli/tree/a7e7c67235519fff79c6e82c44550ac29255dcdc) and its [exact invocation](https://github.com/postil-dev/postil-cli/blob/a7e7c67235519fff79c6e82c44550ac29255dcdc/bench/README.md#expanded-clean-bank). The command above uses the same 25 fixture payloads with the default evaluator source digest; its reports are separate evidence.
The clean-bank-v2 evidence identifies the [measured fixture and evaluator source](https://github.com/postil-dev/postil-cli/tree/a7e7c67235519fff79c6e82c44550ac29255dcdc) and its [exact invocation](https://github.com/postil-dev/postil-cli/blob/a7e7c67235519fff79c6e82c44550ac29255dcdc/bench/README.md#expanded-clean-bank). The command above runs those 25 fixture payloads unchanged plus the 2 cross-file cases, with the default evaluator source digest; its reports are separate evidence.

Compare models only when the selected cases, fixture hash, evaluator hash, binary hash, retry settings, and concurrency match. Provider routes remain explicit. Evidence identifies the fixture/evaluator source by immutable commit and the measured executable by SHA-256; use `POSTIL_BIN` to select that executable. A different build produces separate evidence. This authored bank is not held-out validation, and one observation per fixture does not establish a stable false-positive rate.

## Incremental screen

An incremental review cites only the pushed commits but is judged against the complete pull-request change. [`fixtures/incremental-screen.ts`](fixtures/incremental-screen.ts) holds four such cases. In the two clean cases the increment contains only a dependent cleanup: an alert selector, or a rate-limit entry and alert selectors. An earlier push in the complete change deletes their targets. The two must-block contrasts narrow a target that the complete change keeps.

The [incremental entrypoint](src/incremental-screen.ts) passes the cases to `runLive` through a generated launcher. The launcher recognizes each increment by content and adds `--since-sha` and `--pull-request-diff-file`. With `COMPLETE_CHANGE=omit` it leaves out the complete change, so the same cases also measure an increment reviewed alone. `REVIEW_SCORER_MODEL` optionally enables the scorer; the profile must then list it. From `bench/`:

```sh
REVIEW_MODEL=openai/gpt-5.6-luna REVIEW_SCORER_MODEL=openai/gpt-5.6-luna \
SCREEN_PROFILE=../provisional-models.json \
POSTIL_LLM_REQUEST_TIMEOUT_SECS=30 POSTIL_LLM_TOTAL_TIMEOUT_SECS=60 \
timeout 780s bun run src/incremental-screen.ts
```

`summary.binary` names the launcher. The `incrementalScreen` field records the measured executable and its SHA-256, the launcher digest, the complete-change mode, and a framed digest of the fixture modules and entrypoint. Launcher inputs remain in `.runs/<run-id>-inputs/`.

## Managed qualification

Managed qualification exercises an ordered generator and scorer pair through the mock forge and a real provider. It requires an exact pair, provider identity and route, three complete repeats, and a release build whose embedded profile matches the worktree. Run the manual [managed admission workflow](../.github/workflows/bench-live.yml) for the attested hosted path. `bun run verify-admission` validates checked-in admission evidence.
Expand Down
25 changes: 24 additions & 1 deletion bench/fixtures/causality-screen.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import { createHash } from "node:crypto";
import type { BenchmarkCaseInput } from "../src/harness";
import { addedLine, crossFileCase, csvExportRemoval, edgeLbHostnameRemoval, type CrossFileSpec } from "./clean-screen";

// Supplemental contrasts keep the release corpus and evaluator identity intact.
export const causalitySpecs = [
Expand Down Expand Up @@ -74,7 +75,27 @@ export function causalitySource(spec: typeof causalitySpecs[number], side: "befo
.map((line) => line.slice(1)).join("\n");
}

export const causalityScreenCases: BenchmarkCaseInput[] = causalitySpecs.map((spec, index) => {
// Cross-file contrasts: the change removes one target but also narrows a
// dependency on a target that the same change keeps.
function overreach(spec: CrossFileSpec, id: string, pullNumber: number, path: string, body: string): CrossFileSpec {
const file = spec.files.find((candidate) => candidate.path === path)!;
const added = file.lines.filter((line) => line.startsWith("+")).map((line) => addedLine(file, line.slice(1)));
const line = Math.min(...added);
return { ...spec, id, pullNumber, primaryChange: { path, line },
labels: [...spec.labels, "supplemental-causality"],
defect: { path, line, endLine: Math.max(...added), body } };
}

export const crossFileCausalityCases: BenchmarkCaseInput[] = [
overreach(edgeLbHostnameRemoval("edge|edge-legacy"), "causality-cross-file-alert-drops-kept-router", 206,
"k8s/monitoring/prometheusrule-edge-lb-traefik.yaml",
"The 429 alert also drops the portal-beta router, whose IngressRoute this change keeps. Restore portal-beta to the router selectors."),
overreach(csvExportRemoval("reports.get('/reports/:id/export.pdf', exportReportPdf);"),
"causality-cross-file-limit-dropped-from-kept-route", 207, "src/server/routes/reports.ts",
"The PDF export route loses its rate limit although only CSV export is removed. Restore rateLimit('reports.exportPdf') on the PDF route."),
].map(crossFileCase);

const singleFileCausalityCases: BenchmarkCaseInput[] = causalitySpecs.map((spec, index) => {
const before = causalitySource(spec, "before");
const after = causalitySource(spec, "after");
const expected = spec.defect === null ? [] : [{
Expand All @@ -101,3 +122,5 @@ export const causalityScreenCases: BenchmarkCaseInput[] = causalitySpecs.map((sp
expectations: { minFindings: expected.length, maxFindings: expected.length, requiredFindings: expected },
};
});

export const causalityScreenCases: BenchmarkCaseInput[] = [...singleFileCausalityCases, ...crossFileCausalityCases];
Loading
Loading