Conversation
There was a problem hiding this comment.
@zzylol
The differential workflow is currently deterministically red for two independent reasons.
-
ComposeLifecycle.Startstarts the one-shotplannerwithdocker compose up -d …, then separately callsdocker compose wait planner.plannercan successfully exit in the interval between those commands, leavingwaitwith no container to observe (no containers for project ...). This already prevents the existing corpora from running. Please start and await the planner atomically (for example, anupinvocation that waits and propagates the planner exit code), or otherwise synchronize on the generated snapshot. -
The new issue-754 corpus contains
quantile_over_time(0.9, data[1m]) / quantile_over_time(0.5, data[1m]), which the selected-plan gate explicitly classifies as requiringexact_fallback/. The subsequent benefit measurements therefore can never run, and the required CI job stays red. This is a valid semantic gate result, but it means the PR is not merge-ready as a level-3 benchmark: implement local execution for that query, or remove/mark it as unsupported until that work lands.
cf0395b to
59e848b
Compare
4c91b8d to
c338080
Compare
# Conflicts: # docs/design_docs/planning-test-layers.md
# Conflicts: # promql-compliance/README.md # promql-compliance/runner/Makefile
Deferred follow-up. Not required for the current #728 → #742 → #775 synthetic-selection and data-plane execution milestone. Online ERP feedback and runtime replanning are explicitly out of scope for that milestone.
Distinguish measured plan-selection quality from synthetic ranking and cross-engine fixture performance.
Before this PR: “Level 3” meant a fixture benefit comparison against exact engines; that cannot show whether Backend selected the right physical candidate using real evidence.
After this PR: a separate measured-selection audit consumes the compiled candidate report, real-trace/evidence artifacts and repeated per-candidate measurements. It checks artifact integrity, workload/environment/horizon consistency, calibration/evaluation separation, evidence freshness, used cost-model/generation identity, complete candidate coverage, accurate results and latency/memory limits. It reports total-CPU selection regret and observed run ranges, separately for single-query and full-workload experiments. Historical benefit tooling/results remain explicitly labeled as a different evaluation.
Validation: eight audit contract tests pass; all five split Rust integration targets pass with strict Clippy, and nine resource exporter tests pass. Seven finite-fixture execution runs are preserved in #775 with raw reports and source provenance.
Draft: real-evidence acceptance remains unverified. Contract tests are synthetic. No matched real-trace experiment with applicable ERP/resource evidence and measurements for every admitted candidate has been completed. The audit tool does not collect production telemetry or authenticate measurement origin; missing evidence fails rather than becoming zero or a fixture pass. Historical benefit failures remain preserved.
Base: #778. Review #728 → #742 → #775 → #776 → #777 → #778 → #759. See
docs/evaluation/planning-review-order.mdandtools/planning-validation/README.mdfor responsibilities, the accepted CPU objective and reproducible audit command. Human plan approval remains outstanding.