Skip to content

feat(control-plane): scope evidence and compute ERP/analytical workload costs - #761

Open
zzylol wants to merge 59 commits into
refactor/query-plan-dag-executionfrom
issue-752
Open

zzylol wants to merge 59 commits into
refactor/query-plan-dag-executionfrom
issue-752

Conversation

@zzylol

@zzylol zzylol commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Dataset-bound SDS alignment

Workload cost manifests include the logical dataset identity from the installed ingestion contract, including raw candidates. Synthetic quotes cannot be reused across dataset scopes. This branch inherits #737 dataset binding and same-version recovery contracts, plus the whole-query revision-fence regression adapted to shared DAG execution.

Stack: main → #768 → #737 → #749 → #771 → #774 → #763 → #765 → #761 → #728 → #742 → #759. Diagnostics #756 and controls #766 branch from #761.

Before this PR

Backend selection lacked complete runtime-informed workload costs and scoped accuracy evidence. Native TopK computations and stored-state maintenance boundaries did not consistently survive deployment installation and recovery.

After this PR

Backend compares physically compiled candidates using workload demand, applicable ERP measurements, analytical resource estimates or complete provider quotes. Accuracy admission remains separate from pricing. Missing or stale proof cannot be replaced by a favorable cost.

For topk by(job)(1, rate(requests_total[1m])), Planner supplies exact ranking, query-time CMS/CountSketch heap, and fixed-window precomputed heap candidates. Backend binds counter or heap SDS and preserves the supplied operators. The precompute candidate buffers complete timestamped counter states, finalizes each series' Rate, builds a fresh heap, and stores it; query execution reads that heap. It never adds finalized rates across windows.

For sum by(job)(rate(requests_total[1m])), Planner also supplies query-time Sum and precomputed Sum. Both finalize Rate per series before grouping. Complete 60-second windows may overlap at a five-second evaluation cadence; each window is an independent result. Missing grouping labels match PromQL grouping semantics.

Spatial TopK supports native exact ranking and signed CountSketch heap over the complete current-series snapshot. Decreases, staleness and expiration replace eligibility rather than accumulating historical weights. Arbitrary signed values do not authorize CMS.

Installed query and maintenance programs retain typed boundaries. Bound SDS reads validate deployed output and semantic definition, window, generation, encoding and schema. Grouped native heap/Sum outputs publish one atomic batch per window; logical groups remain inside that batch. Recovery does not re-lower the physical operators.

Cost manifests include maintenance programs as well as query programs. Maintenance dependencies receive explicit retention requirements, and retained heap sizing includes groups inside the atomic batch.

Validation and scope

  • Real-process Remote Write → counter/heap SDS → HTTP → restart tests cover both heap families, counter resets, independent and overlapping windows, hidden/missing series labels and missing windows. Grouped Sum tests exercise 60s, 65s and 120s evaluations before and after restart.
  • Candidate tests cover accuracy rejection, exact/CMS/CountSketch cost selection, both physical graphs, recovery rejection when the maintenance program is missing, and maintenance CPU/workspace costs.
  • Operator/storage/library regression and strict Clippy results are recorded in the accompanying execution evidence.

Precomputed Rate aggregates require the finite complete-input barrier and aligned complete windows at the selected evaluation cadence. Continuous population completion is not inferred from elapsed time. Ad-hoc SDS discovery remains deferred.

Other PromQL shapes still use existing adapters. The current workload search evaluates per-root substitutions and explicitly reports that joint workload search is not exhaustive. Automated tests do not constitute manual deployment verification or human plan approval. #759's previous performance gate remains failed.

Planner #462 pin: 176c1bd565e0c400f9a35996c2bd65de32475e72.

Closes #752. See docs/design_docs/evidence-dependent-candidates.md and docs/design_docs/summary-catalog-sds-architecture.md.

@zzylol
zzylol changed the base branch from main to refactor/query-plan-dag-execution September 22, 2026 20:20
@zzylol
zzylol changed the base branch from refactor/query-plan-dag-execution to perf/issue-758 September 22, 2026 20:45
@zzylol zzylol changed the title feat(control-plane): gate evidence-dependent Planner candidates feat(control-plane): scope evidence and compute ERP/analytical workload costs Sep 23, 2026
@zzylol
zzylol changed the base branch from perf/issue-758 to refactor/query-plan-dag-execution September 24, 2026 12:56
# Conflicts:
#	Cargo.lock
#	Cargo.toml
#	control_plane/src/physical/compiler.rs
# Conflicts:
#	data_plane/src/query_engines/asap_clickhouse_query_engine/execution.rs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Integrate Planner evidence-dependent candidates into backend selection

1 participant