Judge incremental reviews against the complete change - #216
Merged
Merged
Conversation
Add two clean cases in which part of a change removes a target and another part narrows a dependency on it: test edge hostnames removed together with their Traefik routes and alert selectors, and a disabled export feature removed with its route, caller, flag, rate-limit entry and alert selectors. Each narrowed hunk looks like lost coverage in isolation. Add two must-block contrasts whose narrowing also drops a target that the change keeps. The cases extend only the supplemental clean and causality banks. The admission corpus and the attested evaluator sources are unchanged.
A finding anchored on a change-metadata or pull-request description line took its source role from the adjudicator's copied evidence. Unresolved results must carry empty evidence, so any fresh unresolved finding on a deleted file failed scope validation for every candidate and ended the review as invalid adjudication output. The reviewed citation now fixes the metadata role, and its evidence becomes the causal change. Confirmed results must still copy the anchored metadata evidence.
An incremental re-review sent only the commits pushed since the last review. A later commit that narrowed an alert selector, rate limit, or caller for a target an earlier commit had removed looked like lost coverage, and the review reported it as a regression. Incremental generator and scorer requests now also carry the complete pull-request change as uncitable context: the raw diff up to 24 KiB, else a per-file summary with status and line counts. The context counts toward request admission and is dropped only when it does not fit the model budget. Forge reviews fetch the complete change themselves; a failed fetch keeps the incremental review. Local reviews supply it with --pull-request-diff-file, which requires --diff-file and --since-sha. Findings still cite only the increment. A supplemental incremental screen reproduces the isolated false positive and pairs each clean case with a contrast that narrows a kept target.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An incremental re-review saw only the newly pushed commits, so a follow-up that cleaned up a reference to something an earlier commit removed (an alert selector, a rate-limit entry, a caller) read as lost coverage and was reported as a regression; incremental generator and scorer requests now also carry the complete pull-request change as uncitable context, bounded to 24 KiB with a per-file summary fallback, fetched by forge reviews and supplied locally with
--pull-request-diff-file. An unresolved adjudication of a finding anchored on change metadata, such as a deleted file, no longer fails the whole review. A new incremental screen and cross-file clean and contrast cases cover dependent cleanups; withopenai/gpt-5.6-lunaon Azure EU the incremental clean cases were silent in 6 of 6 runs with both contrasts detected, against false positives in 2 of 3 runs without the complete change. The admission corpus and evaluator sources are unchanged, and the CLI gate, exact-head canonical review and exhaustive pre-push review passed.