Skip to content

[Deferred] test: audit real-evidence candidate selection and execution quality - #759

Draft
zzylol wants to merge 71 commits into
test/resource-measurementsfrom
test/issue754-level3
Draft

zzylol wants to merge 71 commits into
test/resource-measurementsfrom
test/issue754-level3

Conversation

@zzylol

@zzylol zzylol commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Deferred follow-up. Not required for the current #728 → #742 → #775 synthetic-selection and data-plane execution milestone. Online ERP feedback and runtime replanning are explicitly out of scope for that milestone.

Distinguish measured plan-selection quality from synthetic ranking and cross-engine fixture performance.

Before this PR: “Level 3” meant a fixture benefit comparison against exact engines; that cannot show whether Backend selected the right physical candidate using real evidence.

After this PR: a separate measured-selection audit consumes the compiled candidate report, real-trace/evidence artifacts and repeated per-candidate measurements. It checks artifact integrity, workload/environment/horizon consistency, calibration/evaluation separation, evidence freshness, used cost-model/generation identity, complete candidate coverage, accurate results and latency/memory limits. It reports total-CPU selection regret and observed run ranges, separately for single-query and full-workload experiments. Historical benefit tooling/results remain explicitly labeled as a different evaluation.

Validation: eight audit contract tests pass; all five split Rust integration targets pass with strict Clippy, and nine resource exporter tests pass. Seven finite-fixture execution runs are preserved in #775 with raw reports and source provenance.

Draft: real-evidence acceptance remains unverified. Contract tests are synthetic. No matched real-trace experiment with applicable ERP/resource evidence and measurements for every admitted candidate has been completed. The audit tool does not collect production telemetry or authenticate measurement origin; missing evidence fails rather than becoming zero or a fixture pass. Historical benefit failures remain preserved.

Base: #778. Review #728 → #742 → #775 → #776 → #777 → #778 → #759. See docs/evaluation/planning-review-order.md and tools/planning-validation/README.md for responsibilities, the accepted CPU objective and reproducible audit command. Human plan approval remains outstanding.

@milindsrivastava1997 milindsrivastava1997 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@zzylol
The differential workflow is currently deterministically red for two independent reasons.

  1. ComposeLifecycle.Start starts the one-shot planner with docker compose up -d …, then separately calls docker compose wait planner. planner can successfully exit in the interval between those commands, leaving wait with no container to observe (no containers for project ...). This already prevents the existing corpora from running. Please start and await the planner atomically (for example, an up invocation that waits and propagates the planner exit code), or otherwise synchronize on the generated snapshot.

  2. The new issue-754 corpus contains quantile_over_time(0.9, data[1m]) / quantile_over_time(0.5, data[1m]), which the selected-plan gate explicitly classifies as requiring exact_fallback/. The subsequent benefit measurements therefore can never run, and the required CI job stays red. This is a valid semantic gate result, but it means the PR is not merge-ready as a level-3 benchmark: implement local execution for that query, or remove/mark it as unsupported until that work lands.

@zzylol
zzylol changed the base branch from 732-test-add-prometheus-remote-write-promql-differential-suite to issue-752 September 22, 2026 19:23
@zzylol
zzylol force-pushed the issue-752 branch 2 times, most recently from cf0395b to 59e848b Compare September 22, 2026 20:26
@zzylol
zzylol force-pushed the test/issue754-level3 branch from 4c91b8d to c338080 Compare September 24, 2026 12:55
@zzylol
zzylol changed the base branch from issue-752 to 732-test-add-prometheus-remote-write-promql-differential-suite September 24, 2026 12:57
@zzylol zzylol changed the title test: measure issue 754 level-3 benefit against exact DBs test: audit real-evidence candidate selection and execution quality Sep 28, 2026
@zzylol
zzylol changed the base branch from 732-test-add-prometheus-remote-write-promql-differential-suite to test/resource-measurements September 28, 2026 02:08
@zzylol zzylol changed the title test: audit real-evidence candidate selection and execution quality [Deferred] test: audit real-evidence candidate selection and execution quality Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants