From b400e54dca829bbb643529fd6fab43b55d86677c Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 17:59:21 +0000 Subject: [PATCH 01/56] docs: define the Planner and deployment layering contract Add the layering proposal, its proposals-index entry, and an Output layers section in the input/output/workflow design. Every Planner layer outputs all legal candidates; the deployment keeps every summary-family candidate and selects with its own costs. Design DAG names are marked as the target API with the current main type alongside. Co-Authored-By: Claude Opus 5.5 --- .../architecture/input-output-workflow.md | 31 +++ docs/design_docs/proposals/README.md | 1 + .../proposals/planner-backend-layering.md | 244 ++++++++++++++++++ 3 files changed, 276 insertions(+) create mode 100644 docs/design_docs/proposals/planner-backend-layering.md diff --git a/docs/design_docs/architecture/input-output-workflow.md b/docs/design_docs/architecture/input-output-workflow.md index fb9a7f42..05bf0830 100644 --- a/docs/design_docs/architecture/input-output-workflow.md +++ b/docs/design_docs/architecture/input-output-workflow.md @@ -353,6 +353,37 @@ Whole-plan selection chooses compatible alternatives across those targets; construct a complete Post-ASAP DAG. Sharing each target's candidate set avoids the Cartesian-product expansion of complete DAGs and preserves shared nodes. +### Output layers + +The [Planner and deployment layering](../proposals/planner-backend-layering.md) +proposal places `PlanSpace` in a longer pipeline, from what to compute to how +to run it. Main implements layers 0 and 1; the others are in open pull +requests. The proposal's [DAG names](../proposals/planner-backend-layering.md#dag-names) +table maps each design name to its current type on main. + +| Layer | Output | Decides | +|---|---|---| +| 0. Frontends | `CandidatePreASAPDAGs` (today one `QueryExpr` root per entry) | Parse and lower PromQL, SQL, or MetricsQL; reject unsupported semantics. | +| 1. Logical Post-ASAP | `CandidatePostASAPDAGs` (today `PlanSpace`) | What to compute: summary families, rewrites, and exact candidates. No placement. | +| 2. Summary maintenance lifecycle | `CandidatePostASAPDAGs` with timing | `Ephemeral`, `Prepared`, `Shared`, or `ContinuouslyMaintained` per unique summary state; each admissible assignment sets node timing, window framework, and retention. The only source of timing. | +| 3. Physical compilation | `CandidatePhysicalDAGs` | Compile operators and kernels, then cut by timing into precompute and query DAGs with typed input contracts. | +| 4–5. Deployment | Deployment-owned | Price lifecycle assignments and summary families with its own costs; select, bind, and execute. | + +#### Candidate generation + +Every layer outputs all legal candidates: + +`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with timing → CandidatePhysicalDAGs` + +No intermediate layer selects a winner. A frontend may produce a singleton set +for an unambiguous query. Candidate sets may be compact or enumerated lazily, +and rejected candidates keep their reasons. The [selection +workflows](#workflows) below, and `PlanOutput` from `e2e_plan` or `optimize`, +are opt-in selection helpers, not stages of this pipeline. A deployment that +prices candidates itself enumerates the candidate space rather than reading a +selected result; in particular it keeps every summary-family candidate, such as +KLL and DDSketch, and compares them with its whole-plan quotes. + ## Workflows All paths start by lowering the workload and searching for candidates: diff --git a/docs/design_docs/proposals/README.md b/docs/design_docs/proposals/README.md index 2cebb3b2..dd0cdadc 100644 --- a/docs/design_docs/proposals/README.md +++ b/docs/design_docs/proposals/README.md @@ -10,3 +10,4 @@ extensions. A design document is not a promise of downstream runtime support. - [ASAP-aware mapping proposals](asap-aware-mapping/README.md) - [Operator sharing](operator-sharing.md) - [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) +- [Planner and deployment layering](planner-backend-layering.md) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md new file mode 100644 index 00000000..c3a1f702 --- /dev/null +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -0,0 +1,244 @@ +# Planner and deployment layering + +Status: proposal. Main implements layers 0 and 1 and the library lifecycle +helpers; open pull requests implement the rest (see +[Implementation status](#implementation-status)). Audience: designers of +ASAPPlanner and of deployments such as ASAPQuery-backend. + +## Goal + +A deployment selects and executes plans that ASAPPlanner compiled. ASAPPlanner +decides what can be computed and how it is computed. The deployment decides +where each piece runs, based on its own costs. It supplies data and state and +runs the compiled plans. It never re-derives the computation or keeps its own +operators. + +## Layers + +```text + PromQL / SQL / MetricsQL + │ +┌───────────────────────── ASAPPlanner ────────────────────────────┐ +│ 0. Frontends │ +│ Parse + lower -> CandidatePreASAPDAGs │ +│ Reject unsupported constructs, such as PromQL fill. │ +│ │ │ +│ 1. Logical Post-ASAP (asap-aware-mapping) │ +│ WHAT to compute; no placement. │ +│ Summary families, rewrites, exact candidates │ +│ -> CandidatePostASAPDAGs │ +│ │ │ +│ 2. Summary maintenance lifecycle │ +│ Per unique summary state / maintained population: │ +│ Ephemeral | Prepared | Shared | ContinuouslyMaintained │ +│ Each assignment -> node timing, window framework, │ +│ retention. The only source of timing. │ +│ -> CandidatePostASAPDAGs with timing │ +│ │ │ +│ 3. Physical compile │ +│ Compile and cut by timing -> CandidatePhysicalDAGs │ +│ Each candidate contains: │ +│ { precompute DAG, query DAG, typed InputContracts } │ +│ Operators + kernels implement all computation. │ +└──────────────────────┬───────────────────────────────────────────┘ + CandidatePhysicalDAGs (the boundary) +┌──────────────────────┴──────── Deployment ──────────────┐ +│ 4 Selection price lifecycle assignments; choose │ +│ 5 Execution ingest, panes, storage, readout, run │ +└─────────────────────────────────────────────────────────┘ +``` + +### Candidate generation by ASAPPlanner + +Every Planner layer outputs all legal candidates represented at that stage: + +`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with timing → CandidatePhysicalDAGs` + +No intermediate layer chooses a winning candidate. Frontend lowering can +produce a singleton candidate set for an unambiguous query; it need not invent +alternative parses. Logical planning exposes the supported computation +candidates; lifecycle enumeration exposes their admissible assignments; +physical compilation preserves their supported physical realizations and cuts. +Unsupported semantics and invalid or infeasible candidates are rejected with +reasons, rather than silently discarded by an intermediate cost selection. + +“All candidates” means the legal candidate space under the supplied semantics, +accuracy requirements, evidence, and capabilities. It can be represented +compactly or enumerated lazily; it does not require eagerly materializing the +Cartesian product. The deployment selects from this space. Library helpers +that return one selected `PlanOutput` are optional selection APIs, not stages +of this candidate-preserving pipeline. + +### DAG names + +This design distinguishes individual DAGs from the collections of candidates +passed between layers. The design names are the target API; main still uses +the earlier types, where one exists. + +| Design name | Meaning | Current main | Target API (implemented by open PRs #508, #480) | +|---|---|---|---| +| `PreASAPDAG` | Frontend-lowered query semantics before summary rewrites | `Rc` | `PreASAPDAG` | +| `CandidatePreASAPDAGs` | Frontend candidates, keyed by workload entry; deterministic frontends produce one per entry | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | +| `PostASAPDAG` | One shared logical computation graph | `Rc` tree; exported as `PostAsapDag` | `PostASAPDAG` | +| `CandidatePostASAPDAGs` | All legal logical candidates for the workload | `PlanSpace` | `CandidatePostASAPDAGs` | +| `CandidatePostASAPDAGs` with timing | The logical candidates with each admissible lifecycle assignment's timing, window framework and retention | None; lifecycle helpers return one selected `SummaryMaintenanceLifecyclePlan` | `CandidatePostASAPDAGs` with timing | +| `PhysicalDAG` | Compiled operators and kernels with typed inputs, before runtime sources are bound | None | `PhysicalDAG` | +| `CandidatePhysicalDAGs` | Physical candidates with their timing cuts, metadata and diagnostics, before deployment selection | None | `CandidatePhysicalDAGs` | +| `BoundPhysicalDAG` | The execution graph after the deployment binds runtime sources | None (deployment-owned) | `BoundPhysicalDAG` | + +A candidate collection shares graphs across its candidates; it is not a copy +of every complete DAG. Timing is attached to the shared logical graph, not +stored in a second graph representation. Compatible timing assignments share +one physical compilation, and each candidate's precompute and query DAGs are +cut from it on demand. The exported `PostAsapDag` form is an explicit +export/import format, not an additional planning layer (see +[Post-ASAP IR](../concepts/post-asap-ir.md#tree-and-exported-dag-forms)). + +### Responsibilities + +| Layer | Owns | Does not own | +|---|---|---| +| 0. Frontends | Language semantics and lowering into `CandidatePreASAPDAGs`. A construct that cannot be represented faithfully is rejected, never ignored (for example PromQL `fill`). | Summaries, placement | +| 1. Logical Post-ASAP | `CandidatePostASAPDAGs`: all legal logical candidates, including summary families, exact rewrites, compositions, and series-identity typing. | Placement | +| 2. Summary maintenance lifecycle | For each unique summary state and maintained population, the lifecycle choices (`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) and their costs under a caller-supplied cost model. Each assignment sets every node's execution timing, window framework and retention, producing `CandidatePostASAPDAGs` with timing. | The cost values themselves | +| 3. Physical compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. Compiles `CandidatePostASAPDAGs` with timing into `CandidatePhysicalDAGs`, including each candidate's timing cuts; it does not select a winner. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | +| 4. Deployment selection | Prices lifecycle assignments and, through its cost model, logical candidates. Shared state is counted once. It binds the chosen plan. | Re-lowering computation | +| 5. Deployment execution | Ingestion and routing, pane assignment and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | + +The deployment may call Planner's logical and lifecycle APIs. The rule is only +that every computation runs as a Planner-compiled physical DAG. + +## The boundary + +The deployment receives `CandidatePhysicalDAGs`, containing all supported +physical candidates, for selection. Each candidate contains: + +* a **precompute DAG**, whose inputs are raw-sample contracts (rows carrying + series labels, timestamp and value; the label set is the complete series + identity) and whose outputs are typed summary states; +* a **query DAG**, whose inputs are stored-state contracts, query-time + raw-series contracts, or both; +* the lifecycle, window framework and retention of every stored output. + +After selection, the deployment binds each input contract, stores each precompute output under +its own storage identity, and returns the query DAG's result. Semantic identity +of stored outputs is defined by the logical DAG they compute. Storage identity +and encoding belong to the deployment. + +## Timing and placement + +Timing comes only from the lifecycle layer. Logical strategies propose +computation candidates, not execution timing or placement. The rules for +applying an assignment are: + +* a retained state (`ContinuouslyMaintained`, `Shared`, `Prepared`) and every + node feeding it run at ingestion time; +* readouts, other consumers and `Ephemeral` states run at query time; +* an `Ephemeral` state that feeds a retained state runs at ingestion time, + because query-time work may not feed ingestion-time work. + +The lifecycle layer produces `CandidatePostASAPDAGs` with timing, containing +one candidate for each admissible lifecycle assignment of a logical candidate. +These are logical graphs with execution timing assigned, not a new DAG +representation or a single selected plan. The deployment prices the +assignments and selects among the candidates. Each required lowering is +compiled once, and compatible assignments share that `PhysicalDAG`; each +candidate's cut is derived from it. The timing frontier contains the +ingestion-time nodes read by query-time nodes, plus an ingestion-time root. An ingestion-time `Binary` is +the one exception. It lowers differently from a query-time one, so its timing +must match at compile time. Assignments that change this lowering need a +matching compilation; other timing cuts can reuse the compiled graph. + +## Cost and selection + +The deployment decides both placement and the summary family. + +* **Placement.** For each assignment, the deployment prices every state's + lifecycle with its own costs, then chooses the cheapest admissible + assignment for the whole workload: + * build cost; + * maintenance per update (precompute CPU); + * reads; + * retention: summary store cost, meaning state bytes × retained panes of + the installed window layout × cardinality × store price; + * retirement; + * for `Ephemeral`, processing of the raw samples read at query time. +* **Sharing.** A state shared by several queries is priced once, with all + consumers' demand. This holds only when compilation would install one shared + output: same window layout, evaluation interval and phase. +* **Family.** Every summary-family candidate (for example KLL and DDSketch + for one quantile) reaches the deployment. The deployment compares their + physical candidates with its whole-plan quotes, alongside lifecycle and + placement choices. It must not prune families early through Planner's global + selection. Those helpers remain available to callers that explicitly ask + Planner to select; they are not an intermediate pruning stage. Known gap: + ASAPQuery-backend #795 still pre-selects families through + `global_selection`; this is to be changed. +* Unknown cost stays unknown. It never becomes zero, and an alternative that + cannot be priced is not selected. + +## Query-time raw data and mixed placement + +An `Ephemeral` state is built at query time, so the deployment must provide raw +data as a query-time source (for example Prometheus raw series). If it cannot, +that alternative is not offered. + +A single query may mix `Ephemeral` and stored inputs, with bounded staleness: + +* raw inputs are read at the query evaluation time `t_q`; +* each stored input uses its latest complete revision, with watermark `t_s`; +* the mix is admitted only if `t_q − t_s ≤ max_lag` for every stored input; + `max_lag` is configurable and defaults to one slide of that output; +* the observed lag is reported with the result; +* beyond the bound, the query takes the exact fallback. A silently stale mix is + never returned. + +## Failure and compatibility + +* **Fail closed.** A query or subtree that Planner cannot compile is answered + whole by the deployment's exact engine. Nothing is approximated, ignored or + computed by deployment-owned operators. +* **Development-stage compatibility.** Persisted plans and stored formats carry + versions. A format change bumps the version and rejects old data with a clear + error. Old data is not migrated and is never misread. + +## Example: store cost changes placement + +The following uses `sum by (job) (rate(m[1m]))`, evaluated every 10 s. + +Consider one logical candidate: a per-series Rate state feeding a +grouped Sum state. Logical planning does not return placement variants. The +lifecycle layer lists choices for both states. The deployment prices them: + +* When summary storage is cheap, both states are retained. The precompute DAG + builds Rate and Sum per pane, and the query DAG only reads the Sum. +* When storage is expensive enough, Sum becomes `Ephemeral`. Precompute keeps + only Rate, and the query DAG builds Sum at query time. +* When storage is more expensive still, both states become `Ephemeral`. The + query DAG reads raw series at `t_q`. + +All three are cuts of one compilation. The deployment's decision is only which +lifecycle assignment to buy. + +## Implementation status + +Main implements frontends, logical candidates (`PlanSpace`) and the library +lifecycle and selection helpers. The remaining layers are in open, stacked pull +requests. + +* **Planner.** Split of #462: #473 → #474 → #475. Lifecycle candidates and + timing: #476, #482, #479, #485, #491. Physical layer: #483, #477. Physical + compile coverage: #484, #486–#490, #492–#495, #500–#507; the remaining gaps + are listed in physical-compile-coverage.md, added by #484. Target DAG names: + #508, #480. +* **ASAPQuery-backend.** #788 through #812: query-time raw sources, lifecycle + placement, Planner precompute and query DAGs, and codecs over Planner + kernels. + +Open items: + +* Family pruning in the backend: #795 pre-selects families through + `global_selection` instead of pricing every family candidate. +* Removing edge sampling (`sample_p`) and delta-frame ingestion; the data plane + does not ingest OTLP. +* The remaining physical-compile coverage gaps. From f1c13d4ef9f25d24f2e34e495c02f3b3138a9e5b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:07:01 +0000 Subject: [PATCH 02/56] docs: treat source binding as a PhysicalDAG execution step Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/proposals/planner-backend-layering.md | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index c3a1f702..60e57c9e 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -84,7 +84,6 @@ the earlier types, where one exists. | `CandidatePostASAPDAGs` with timing | The logical candidates with each admissible lifecycle assignment's timing, window framework and retention | None; lifecycle helpers return one selected `SummaryMaintenanceLifecyclePlan` | `CandidatePostASAPDAGs` with timing | | `PhysicalDAG` | Compiled operators and kernels with typed inputs, before runtime sources are bound | None | `PhysicalDAG` | | `CandidatePhysicalDAGs` | Physical candidates with their timing cuts, metadata and diagnostics, before deployment selection | None | `CandidatePhysicalDAGs` | -| `BoundPhysicalDAG` | The execution graph after the deployment binds runtime sources | None (deployment-owned) | `BoundPhysicalDAG` | A candidate collection shares graphs across its candidates; it is not a copy of every complete DAG. Timing is attached to the shared logical graph, not @@ -94,6 +93,11 @@ cut from it on demand. The exported `PostAsapDag` form is an explicit export/import format, not an additional planning layer (see [Post-ASAP IR](../concepts/post-asap-ir.md#tree-and-exported-dag-forms)). +Binding runtime sources is an execution step of a `PhysicalDAG`, not another +DAG. At execution the deployment supplies a source for each typed input slot, +the slots are checked against their contracts, and the graph runs; the bound +instance lives only for that execution and is not persisted or compared. + ### Responsibilities | Layer | Owns | Does not own | From 2bc772f8e662dfa2bf7379f1cc244a461a974751 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:25:00 +0000 Subject: [PATCH 03/56] docs: place summary materialization under physical planning Split physical planning into physical design (summary materialization, workload-level, like materialized-view selection) and physical implementation (per-node lowering and cutting). The annotated graph stays a Post-ASAP DAG; materialization is decided in design and realized by the cut. State node-by-node correspondence and the Fallback exception that the operator-flattening proposal removes. Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 81 ++++++++++++++----- 1 file changed, 62 insertions(+), 19 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index 60e57c9e..f63397c3 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -23,22 +23,24 @@ operators. │ Parse + lower -> CandidatePreASAPDAGs │ │ Reject unsupported constructs, such as PromQL fill. │ │ │ │ +│ Logical planning (what) │ │ 1. Logical Post-ASAP (asap-aware-mapping) │ │ WHAT to compute; no placement. │ │ Summary families, rewrites, exact candidates │ │ -> CandidatePostASAPDAGs │ │ │ │ -│ 2. Summary maintenance lifecycle │ -│ Per unique summary state / maintained population: │ -│ Ephemeral | Prepared | Shared | ContinuouslyMaintained │ -│ Each assignment -> node timing, window framework, │ -│ retention. The only source of timing. │ -│ -> CandidatePostASAPDAGs with timing │ +│ Physical planning (how) │ +│ 2. Physical design: summary materialization │ +│ Per unique summary state / maintained population, choose a │ +│ lifecycle: Ephemeral | Prepared | Shared | │ +│ ContinuouslyMaintained. Each assignment -> node timing, │ +│ window framework, retention. The only source of timing. │ +│ -> CandidatePostASAPDAGs with materialization annotations │ │ │ │ -│ 3. Physical compile │ -│ Compile and cut by timing -> CandidatePhysicalDAGs │ -│ Each candidate contains: │ -│ { precompute DAG, query DAG, typed InputContracts } │ +│ 3. Physical implementation: compilation │ +│ Lower each Post-ASAP node to operators; cut by timing │ +│ -> CandidatePhysicalDAGs │ +│ Each candidate: { precompute DAG, query DAG, InputContracts } │ │ Operators + kernels implement all computation. │ └──────────────────────┬───────────────────────────────────────────┘ CandidatePhysicalDAGs (the boundary) @@ -52,7 +54,7 @@ operators. Every Planner layer outputs all legal candidates represented at that stage: -`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with timing → CandidatePhysicalDAGs` +`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with materialization annotations → CandidatePhysicalDAGs` No intermediate layer chooses a winning candidate. Frontend lowering can produce a singleton candidate set for an unambiguous query; it need not invent @@ -81,7 +83,7 @@ the earlier types, where one exists. | `CandidatePreASAPDAGs` | Frontend candidates, keyed by workload entry; deterministic frontends produce one per entry | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | | `PostASAPDAG` | One shared logical computation graph | `Rc` tree; exported as `PostAsapDag` | `PostASAPDAG` | | `CandidatePostASAPDAGs` | All legal logical candidates for the workload | `PlanSpace` | `CandidatePostASAPDAGs` | -| `CandidatePostASAPDAGs` with timing | The logical candidates with each admissible lifecycle assignment's timing, window framework and retention | None; lifecycle helpers return one selected `SummaryMaintenanceLifecyclePlan` | `CandidatePostASAPDAGs` with timing | +| `CandidatePostASAPDAGs` with materialization annotations | The same Post-ASAP candidates, annotated with each admissible lifecycle assignment's timing, window framework and retention; still Post-ASAP DAGs, not physical ones | None; lifecycle helpers return one selected `SummaryMaintenanceLifecyclePlan` | `CandidatePostASAPDAGs` with timing | | `PhysicalDAG` | Compiled operators and kernels with typed inputs, before runtime sources are bound | None | `PhysicalDAG` | | `CandidatePhysicalDAGs` | Physical candidates with their timing cuts, metadata and diagnostics, before deployment selection | None | `CandidatePhysicalDAGs` | @@ -93,19 +95,60 @@ cut from it on demand. The exported `PostAsapDag` form is an explicit export/import format, not an additional planning layer (see [Post-ASAP IR](../concepts/post-asap-ir.md#tree-and-exported-dag-forms)). +There are only two kinds of DAG after the frontend: Post-ASAP DAGs, whose +nodes are logical operations (optionally carrying materialization annotations), +and physical DAGs, whose nodes are selected physical operators. + Binding runtime sources is an execution step of a `PhysicalDAG`, not another DAG. At execution the deployment supplies a source for each typed input slot, the slots are checked against their contracts, and the graph runs; the bound instance lives only for that execution and is not persisted or compared. +### Physical planning + +Physical planning has two parts of different character, as in databases. + +**Physical design (summary materialization).** In this document, +*materialization* means summary state kept across executions, like a +materialized view: whether a state is maintained, how it is refreshed and how +long it is retained. This is a workload-level decision, like a database's +materialized-view selection: it spans queries (shared state is kept once), it +depends on workload demand (read and update rates, horizon), and its result +persists. The Planner enumerates the admissible choices; the deployment prices +and chooses. The choice is recorded as annotations on the Post-ASAP nodes +(timing, window framework, retention), so the annotated graph is still a +Post-ASAP DAG, much as physical properties annotate logical expressions in a +database optimizer. Operator materialization (blocking operators such as sort, +aggregation or summary build) and computing a shared subexpression once are +details of compilation and the runtime, not separate layers. + +**Physical implementation (compilation).** Compilation lowers each Post-ASAP +node to physical operators and cuts the graph by timing into a precompute and a +query DAG. Materialization is *decided* by physical design and *realized* here: +the precompute DAG's outputs at the cut are the materialized states, and the +query DAG reads them through typed input slots. A physical DAG corresponds to +its Post-ASAP DAG node by node. A node may expand into several operators (for +example TopK into sort then limit, or a multi-input merge into union then +merge); helper operators are numbered from their source node, so every operator +traces back to one Post-ASAP node. One exception is a `Fallback` node, which +wraps a whole Pre-ASAP expression and compiles to many operators; the +operator-flattening proposal ([operator sharing](operator-sharing.md), from +#469) removes it by making non-ASAP operators ordinary Post-ASAP nodes. Its +export section currently groups the largest non-ASAP subtree into one +`Relational` fragment; per-node correspondence needs that export to keep one +node per operator. Today each node kind has essentially one lowering (a +`Binary` lowers differently at ingestion and query time); if a node gains +alternative physical implementations, they become further candidates in +`CandidatePhysicalDAGs`, priced and chosen by the deployment. + ### Responsibilities | Layer | Owns | Does not own | |---|---|---| | 0. Frontends | Language semantics and lowering into `CandidatePreASAPDAGs`. A construct that cannot be represented faithfully is rejected, never ignored (for example PromQL `fill`). | Summaries, placement | | 1. Logical Post-ASAP | `CandidatePostASAPDAGs`: all legal logical candidates, including summary families, exact rewrites, compositions, and series-identity typing. | Placement | -| 2. Summary maintenance lifecycle | For each unique summary state and maintained population, the lifecycle choices (`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) and their costs under a caller-supplied cost model. Each assignment sets every node's execution timing, window framework and retention, producing `CandidatePostASAPDAGs` with timing. | The cost values themselves | -| 3. Physical compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. Compiles `CandidatePostASAPDAGs` with timing into `CandidatePhysicalDAGs`, including each candidate's timing cuts; it does not select a winner. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | +| 2. Physical design: summary materialization | For each unique summary state and maintained population, the lifecycle choices (`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) and their costs under a caller-supplied cost model. Each assignment sets every node's execution timing, window framework and retention, producing `CandidatePostASAPDAGs` with materialization annotations. | The cost values themselves; operator implementation | +| 3. Physical implementation: compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. Lowers each node of `CandidatePostASAPDAGs` with materialization annotations to operators and cuts by timing into `CandidatePhysicalDAGs`; it does not select a winner. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | | 4. Deployment selection | Prices lifecycle assignments and, through its cost model, logical candidates. Shared state is counted once. It binds the chosen plan. | Re-lowering computation | | 5. Deployment execution | Ingestion and routing, pane assignment and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | @@ -131,7 +174,7 @@ and encoding belong to the deployment. ## Timing and placement -Timing comes only from the lifecycle layer. Logical strategies propose +Timing comes only from physical design, that is, from the chosen lifecycle assignment. Logical strategies propose computation candidates, not execution timing or placement. The rules for applying an assignment are: @@ -141,10 +184,10 @@ applying an assignment are: * an `Ephemeral` state that feeds a retained state runs at ingestion time, because query-time work may not feed ingestion-time work. -The lifecycle layer produces `CandidatePostASAPDAGs` with timing, containing -one candidate for each admissible lifecycle assignment of a logical candidate. -These are logical graphs with execution timing assigned, not a new DAG -representation or a single selected plan. The deployment prices the +Physical design produces `CandidatePostASAPDAGs` with materialization +annotations, containing one candidate for each admissible lifecycle assignment +of a logical candidate. These are Post-ASAP graphs with execution timing +assigned, not a new DAG representation or a single selected plan. The deployment prices the assignments and selects among the candidates. Each required lowering is compiled once, and compatible assignments share that `PhysicalDAG`; each candidate's cut is derived from it. The timing frontier contains the From 6702d3a23c6b6c7a2ef7c8b1f1e9423b736ccabc Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:27:06 +0000 Subject: [PATCH 04/56] docs: point per-operator export at #481 Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/proposals/planner-backend-layering.md | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index f63397c3..2ffb5323 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -133,10 +133,8 @@ merge); helper operators are numbered from their source node, so every operator traces back to one Post-ASAP node. One exception is a `Fallback` node, which wraps a whole Pre-ASAP expression and compiles to many operators; the operator-flattening proposal ([operator sharing](operator-sharing.md), from -#469) removes it by making non-ASAP operators ordinary Post-ASAP nodes. Its -export section currently groups the largest non-ASAP subtree into one -`Relational` fragment; per-node correspondence needs that export to keep one -node per operator. Today each node kind has essentially one lowering (a +#469) removes it by making non-ASAP operators ordinary Post-ASAP nodes, and +#481 revises its export to emit one Post-ASAP node per non-ASAP operator. Today each node kind has essentially one lowering (a `Binary` lowers differently at ingestion and query time); if a node gains alternative physical implementations, they become further candidates in `CandidatePhysicalDAGs`, priced and chosen by the deployment. From 1bb8ec5a4304606d669a5ea0cef780335c9a525a Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:32:42 +0000 Subject: [PATCH 05/56] docs: state family retention and timing-sensitive lowering plainly Describe current Planner behaviour (no family pruning) and the backend gap, and replace the Binary 'exception' with the general rule that a timing- sensitive node needs one compilation per distinct timing, with an example. Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 41 +++++++++++-------- 1 file changed, 25 insertions(+), 16 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index 2ffb5323..02ce4f95 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -185,14 +185,23 @@ applying an assignment are: Physical design produces `CandidatePostASAPDAGs` with materialization annotations, containing one candidate for each admissible lifecycle assignment of a logical candidate. These are Post-ASAP graphs with execution timing -assigned, not a new DAG representation or a single selected plan. The deployment prices the -assignments and selects among the candidates. Each required lowering is -compiled once, and compatible assignments share that `PhysicalDAG`; each -candidate's cut is derived from it. The timing frontier contains the -ingestion-time nodes read by query-time nodes, plus an ingestion-time root. An ingestion-time `Binary` is -the one exception. It lowers differently from a query-time one, so its timing -must match at compile time. Assignments that change this lowering need a -matching compilation; other timing cuts can reuse the compiled graph. +assigned, not a new DAG representation or a single selected plan. The +deployment prices the assignments and selects among the candidates. The timing +frontier contains the ingestion-time nodes read by query-time nodes, plus an +ingestion-time root. + +**Timing can change a node's implementation.** For most nodes, timing only +decides which side of the cut the node falls on, so one compilation serves every +lifecycle assignment and each assignment is a different cut of it. A node whose +physical operator depends on its timing is *timing-sensitive*. Today the only +such node is `Binary`: at ingestion time it combines aligned row streams per pane +and its result feeds a summary (`AlignedBinary`); at query time it matches +readout vectors under PromQL rules (`series_binary`). Assignments share a +compilation only when they give every timing-sensitive node the same timing; +each distinct combination is compiled once. For example, in +`quantile(0.9, sum_over_time(m[1m]) + sum_over_time(n[1m]))`, retaining the +quantile state puts `+` at ingestion time, while making that state `Ephemeral` +puts `+` at query time, so the two assignments need two compilations. ## Cost and selection @@ -211,14 +220,14 @@ The deployment decides both placement and the summary family. * **Sharing.** A state shared by several queries is priced once, with all consumers' demand. This holds only when compilation would install one shared output: same window layout, evaluation interval and phase. -* **Family.** Every summary-family candidate (for example KLL and DDSketch - for one quantile) reaches the deployment. The deployment compares their - physical candidates with its whole-plan quotes, alongside lifecycle and - placement choices. It must not prune families early through Planner's global - selection. Those helpers remain available to callers that explicitly ask - Planner to select; they are not an intermediate pruning stage. Known gap: - ASAPQuery-backend #795 still pre-selects families through - `global_selection`; this is to be changed. +* **Family.** Planner does not prune families: every summary-family candidate + (for example KLL and DDSketch for one quantile) reaches the deployment. The + deployment compares their physical candidates with its whole-plan quotes, + alongside lifecycle and placement choices. Planner's global-selection helpers + select only when a caller explicitly asks; no Planner layer calls them as an + intermediate stage. Known gap: ASAPQuery-backend #795 currently calls + `global_selection` before pricing and so pre-selects families; this is being + changed. * Unknown cost stays unknown. It never becomes zero, and an alternative that cannot be priced is not selected. From 1e84c15059ed6caed877935724951992be3b5da9 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:32:57 +0000 Subject: [PATCH 06/56] docs: drop Binary lowering detail from the layering proposal Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 17 +---------------- 1 file changed, 1 insertion(+), 16 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index 02ce4f95..9a4c6625 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -134,9 +134,7 @@ traces back to one Post-ASAP node. One exception is a `Fallback` node, which wraps a whole Pre-ASAP expression and compiles to many operators; the operator-flattening proposal ([operator sharing](operator-sharing.md), from #469) removes it by making non-ASAP operators ordinary Post-ASAP nodes, and -#481 revises its export to emit one Post-ASAP node per non-ASAP operator. Today each node kind has essentially one lowering (a -`Binary` lowers differently at ingestion and query time); if a node gains -alternative physical implementations, they become further candidates in +#481 revises its export to emit one Post-ASAP node per non-ASAP operator. If a node gains alternative physical implementations, they become further candidates in `CandidatePhysicalDAGs`, priced and chosen by the deployment. ### Responsibilities @@ -190,19 +188,6 @@ deployment prices the assignments and selects among the candidates. The timing frontier contains the ingestion-time nodes read by query-time nodes, plus an ingestion-time root. -**Timing can change a node's implementation.** For most nodes, timing only -decides which side of the cut the node falls on, so one compilation serves every -lifecycle assignment and each assignment is a different cut of it. A node whose -physical operator depends on its timing is *timing-sensitive*. Today the only -such node is `Binary`: at ingestion time it combines aligned row streams per pane -and its result feeds a summary (`AlignedBinary`); at query time it matches -readout vectors under PromQL rules (`series_binary`). Assignments share a -compilation only when they give every timing-sensitive node the same timing; -each distinct combination is compiled once. For example, in -`quantile(0.9, sum_over_time(m[1m]) + sum_over_time(n[1m]))`, retaining the -quantile state puts `+` at ingestion time, while making that state `Ephemeral` -puts `+` at query time, so the two assignments need two compilations. - ## Cost and selection The deployment decides both placement and the summary family. From cd6a18d5cfda2d978a7697bbf525714258cfd8c1 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:35:12 +0000 Subject: [PATCH 07/56] docs: drop mixed-placement and compatibility sections from the layering proposal Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 25 ------------------- 1 file changed, 25 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index 9a4c6625..7a3113a5 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -216,31 +216,6 @@ The deployment decides both placement and the summary family. * Unknown cost stays unknown. It never becomes zero, and an alternative that cannot be priced is not selected. -## Query-time raw data and mixed placement - -An `Ephemeral` state is built at query time, so the deployment must provide raw -data as a query-time source (for example Prometheus raw series). If it cannot, -that alternative is not offered. - -A single query may mix `Ephemeral` and stored inputs, with bounded staleness: - -* raw inputs are read at the query evaluation time `t_q`; -* each stored input uses its latest complete revision, with watermark `t_s`; -* the mix is admitted only if `t_q − t_s ≤ max_lag` for every stored input; - `max_lag` is configurable and defaults to one slide of that output; -* the observed lag is reported with the result; -* beyond the bound, the query takes the exact fallback. A silently stale mix is - never returned. - -## Failure and compatibility - -* **Fail closed.** A query or subtree that Planner cannot compile is answered - whole by the deployment's exact engine. Nothing is approximated, ignored or - computed by deployment-owned operators. -* **Development-stage compatibility.** Persisted plans and stored formats carry - versions. A format change bumps the version and rejects old data with a clear - error. Old data is not migrated and is never misread. - ## Example: store cost changes placement The following uses `sum by (job) (rate(m[1m]))`, evaluated every 10 s. From aa0b9594fa6eaa38dd5f0bcb31326808f548d182 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:35:56 +0000 Subject: [PATCH 08/56] docs: drop the implementation-status section from the layering proposal Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 29 ++----------------- 1 file changed, 2 insertions(+), 27 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index 7a3113a5..c58f5a0d 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -1,9 +1,7 @@ # Planner and deployment layering -Status: proposal. Main implements layers 0 and 1 and the library lifecycle -helpers; open pull requests implement the rest (see -[Implementation status](#implementation-status)). Audience: designers of -ASAPPlanner and of deployments such as ASAPQuery-backend. +Status: proposal. Audience: designers of ASAPPlanner and of deployments such +as ASAPQuery-backend. ## Goal @@ -233,26 +231,3 @@ lifecycle layer lists choices for both states. The deployment prices them: All three are cuts of one compilation. The deployment's decision is only which lifecycle assignment to buy. - -## Implementation status - -Main implements frontends, logical candidates (`PlanSpace`) and the library -lifecycle and selection helpers. The remaining layers are in open, stacked pull -requests. - -* **Planner.** Split of #462: #473 → #474 → #475. Lifecycle candidates and - timing: #476, #482, #479, #485, #491. Physical layer: #483, #477. Physical - compile coverage: #484, #486–#490, #492–#495, #500–#507; the remaining gaps - are listed in physical-compile-coverage.md, added by #484. Target DAG names: - #508, #480. -* **ASAPQuery-backend.** #788 through #812: query-time raw sources, lifecycle - placement, Planner precompute and query DAGs, and codecs over Planner - kernels. - -Open items: - -* Family pruning in the backend: #795 pre-selects families through - `global_selection` instead of pricing every family candidate. -* Removing edge sampling (`sample_p`) and delta-frame ingestion; the data plane - does not ingest OTLP. -* The remaining physical-compile coverage gaps. From 62ae3e770b57c68f3f4adaede546363731f22d3b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:36:54 +0000 Subject: [PATCH 09/56] docs: leave input-output-workflow unchanged in the layering PR Co-Authored-By: Claude Opus 5.5 --- .../architecture/input-output-workflow.md | 31 ------------------- 1 file changed, 31 deletions(-) diff --git a/docs/design_docs/architecture/input-output-workflow.md b/docs/design_docs/architecture/input-output-workflow.md index 05bf0830..fb9a7f42 100644 --- a/docs/design_docs/architecture/input-output-workflow.md +++ b/docs/design_docs/architecture/input-output-workflow.md @@ -353,37 +353,6 @@ Whole-plan selection chooses compatible alternatives across those targets; construct a complete Post-ASAP DAG. Sharing each target's candidate set avoids the Cartesian-product expansion of complete DAGs and preserves shared nodes. -### Output layers - -The [Planner and deployment layering](../proposals/planner-backend-layering.md) -proposal places `PlanSpace` in a longer pipeline, from what to compute to how -to run it. Main implements layers 0 and 1; the others are in open pull -requests. The proposal's [DAG names](../proposals/planner-backend-layering.md#dag-names) -table maps each design name to its current type on main. - -| Layer | Output | Decides | -|---|---|---| -| 0. Frontends | `CandidatePreASAPDAGs` (today one `QueryExpr` root per entry) | Parse and lower PromQL, SQL, or MetricsQL; reject unsupported semantics. | -| 1. Logical Post-ASAP | `CandidatePostASAPDAGs` (today `PlanSpace`) | What to compute: summary families, rewrites, and exact candidates. No placement. | -| 2. Summary maintenance lifecycle | `CandidatePostASAPDAGs` with timing | `Ephemeral`, `Prepared`, `Shared`, or `ContinuouslyMaintained` per unique summary state; each admissible assignment sets node timing, window framework, and retention. The only source of timing. | -| 3. Physical compilation | `CandidatePhysicalDAGs` | Compile operators and kernels, then cut by timing into precompute and query DAGs with typed input contracts. | -| 4–5. Deployment | Deployment-owned | Price lifecycle assignments and summary families with its own costs; select, bind, and execute. | - -#### Candidate generation - -Every layer outputs all legal candidates: - -`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with timing → CandidatePhysicalDAGs` - -No intermediate layer selects a winner. A frontend may produce a singleton set -for an unambiguous query. Candidate sets may be compact or enumerated lazily, -and rejected candidates keep their reasons. The [selection -workflows](#workflows) below, and `PlanOutput` from `e2e_plan` or `optimize`, -are opt-in selection helpers, not stages of this pipeline. A deployment that -prices candidates itself enumerates the candidate space rather than reading a -selected result; in particular it keeps every summary-family candidate, such as -KLL and DDSketch, and compares them with its whole-plan quotes. - ## Workflows All paths start by lowering the workload and searching for candidates: From 3607270b8f31eeb7908d6e394a878296640e34f7 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 19:11:23 +0000 Subject: [PATCH 10/56] docs: move selection into ASAPPlanner; deployment supplies prices Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 98 +++++++++++-------- 1 file changed, 56 insertions(+), 42 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index c58f5a0d..a22e5868 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -5,11 +5,13 @@ as ASAPQuery-backend. ## Goal -A deployment selects and executes plans that ASAPPlanner compiled. ASAPPlanner -decides what can be computed and how it is computed. The deployment decides -where each piece runs, based on its own costs. It supplies data and state and -runs the compiled plans. It never re-derives the computation or keeps its own -operators. +ASAPPlanner decides what can be computed, how it is computed and which plan is +best, and hands the deployment one optimal physical plan to execute. The +deployment prices the candidates with its own costs, but it does not rank or +select: it passes its prices to ASAPPlanner's selection function and receives +the optimal plan. It then supplies data and state and runs that plan. It never +re-derives the computation, keeps its own operators, or implements its own +ranking or selection. ## Layers @@ -40,21 +42,27 @@ operators. │ -> CandidatePhysicalDAGs │ │ Each candidate: { precompute DAG, query DAG, InputContracts } │ │ Operators + kernels implement all computation. │ +│ │ │ +│ 4. Selection (library function) │ +│ Input: deployment's prices, accuracy requirements and │ +│ capabilities. Rank and combine candidates for the workload. │ +│ -> one optimal PhysicalDAG │ └──────────────────────┬───────────────────────────────────────────┘ - CandidatePhysicalDAGs (the boundary) -┌──────────────────────┴──────── Deployment ──────────────┐ -│ 4 Selection price lifecycle assignments; choose │ -│ 5 Execution ingest, panes, storage, readout, run │ -└─────────────────────────────────────────────────────────┘ + optimal PhysicalDAG (the boundary) +┌──────────────────────┴──────── Deployment ───────────────────────┐ +│ Pricing: computes the costs and quotes passed to selection │ +│ 5. Execution: ingest, panes, storage, readout, run │ +└──────────────────────────────────────────────────────────────────┘ ``` ### Candidate generation by ASAPPlanner -Every Planner layer outputs all legal candidates represented at that stage: +Layers 0 to 3 output all legal candidates represented at their stage, and +selection is the only step that chooses: -`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with materialization annotations → CandidatePhysicalDAGs` +`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with materialization annotations → CandidatePhysicalDAGs → optimal PhysicalDAG` -No intermediate layer chooses a winning candidate. Frontend lowering can +No layer before selection chooses a winning candidate. Frontend lowering can produce a singleton candidate set for an unambiguous query; it need not invent alternative parses. Logical planning exposes the supported computation candidates; lifecycle enumeration exposes their admissible assignments; @@ -65,9 +73,9 @@ reasons, rather than silently discarded by an intermediate cost selection. “All candidates” means the legal candidate space under the supplied semantics, accuracy requirements, evidence, and capabilities. It can be represented compactly or enumerated lazily; it does not require eagerly materializing the -Cartesian product. The deployment selects from this space. Library helpers -that return one selected `PlanOutput` are optional selection APIs, not stages -of this candidate-preserving pipeline. +Cartesian product. Selection chooses from this space using the deployment's +prices. Library helpers that return one selected `PlanOutput` without deployment +prices are optional convenience APIs, not stages of this pipeline. ### DAG names @@ -83,7 +91,8 @@ the earlier types, where one exists. | `CandidatePostASAPDAGs` | All legal logical candidates for the workload | `PlanSpace` | `CandidatePostASAPDAGs` | | `CandidatePostASAPDAGs` with materialization annotations | The same Post-ASAP candidates, annotated with each admissible lifecycle assignment's timing, window framework and retention; still Post-ASAP DAGs, not physical ones | None; lifecycle helpers return one selected `SummaryMaintenanceLifecyclePlan` | `CandidatePostASAPDAGs` with timing | | `PhysicalDAG` | Compiled operators and kernels with typed inputs, before runtime sources are bound | None | `PhysicalDAG` | -| `CandidatePhysicalDAGs` | Physical candidates with their timing cuts, metadata and diagnostics, before deployment selection | None | `CandidatePhysicalDAGs` | +| `CandidatePhysicalDAGs` | Physical candidates with their timing cuts, metadata and diagnostics, before selection | None | `CandidatePhysicalDAGs` | +| optimal `PhysicalDAG` | The selected candidate: its precompute and query `PhysicalDAG`s with every stored output's lifecycle, window framework and retention | None | one materialized element of `CandidatePhysicalDAGs` | A candidate collection shares graphs across its candidates; it is not a copy of every complete DAG. Timing is attached to the shared logical graph, not @@ -113,8 +122,8 @@ long it is retained. This is a workload-level decision, like a database's materialized-view selection: it spans queries (shared state is kept once), it depends on workload demand (read and update rates, horizon), and its result persists. The Planner enumerates the admissible choices; the deployment prices -and chooses. The choice is recorded as annotations on the Post-ASAP nodes -(timing, window framework, retention), so the annotated graph is still a +them, and selection chooses. The choice is recorded as annotations on the +Post-ASAP nodes (timing, window framework, retention), so the annotated graph is still a Post-ASAP DAG, much as physical properties annotate logical expressions in a database optimizer. Operator materialization (blocking operators such as sort, aggregation or summary build) and computing a shared subexpression once are @@ -133,7 +142,7 @@ wraps a whole Pre-ASAP expression and compiles to many operators; the operator-flattening proposal ([operator sharing](operator-sharing.md), from #469) removes it by making non-ASAP operators ordinary Post-ASAP nodes, and #481 revises its export to emit one Post-ASAP node per non-ASAP operator. If a node gains alternative physical implementations, they become further candidates in -`CandidatePhysicalDAGs`, priced and chosen by the deployment. +`CandidatePhysicalDAGs`, priced by the deployment and chosen by selection. ### Responsibilities @@ -143,16 +152,19 @@ operator-flattening proposal ([operator sharing](operator-sharing.md), from | 1. Logical Post-ASAP | `CandidatePostASAPDAGs`: all legal logical candidates, including summary families, exact rewrites, compositions, and series-identity typing. | Placement | | 2. Physical design: summary materialization | For each unique summary state and maintained population, the lifecycle choices (`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) and their costs under a caller-supplied cost model. Each assignment sets every node's execution timing, window framework and retention, producing `CandidatePostASAPDAGs` with materialization annotations. | The cost values themselves; operator implementation | | 3. Physical implementation: compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. Lowers each node of `CandidatePostASAPDAGs` with materialization annotations to operators and cuts by timing into `CandidatePhysicalDAGs`; it does not select a winner. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | -| 4. Deployment selection | Prices lifecycle assignments and, through its cost model, logical candidates. Shared state is counted once. It binds the chosen plan. | Re-lowering computation | +| 4. Selection (Planner library function) | Ranking and combining all candidates, including every summary family and lifecycle assignment, across the workload with the deployment's prices, counting shared state once; returning one optimal `PhysicalDAG`. Candidates that cannot be priced are not selected. | The prices themselves | +| Deployment pricing | Computing the prices passed to selection: unit costs, whole-plan quotes, store price; supplying accuracy requirements and capabilities (for example whether query-time raw data is available). | Ranking, sorting or selection | | 5. Deployment execution | Ingestion and routing, pane assignment and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | -The deployment may call Planner's logical and lifecycle APIs. The rule is only -that every computation runs as a Planner-compiled physical DAG. +The deployment may call Planner's earlier-layer APIs, for example to inspect +candidates while pricing them. The rules are that every computation runs as a +Planner-compiled physical DAG and every choice among candidates is made by +Planner's selection function. ## The boundary -The deployment receives `CandidatePhysicalDAGs`, containing all supported -physical candidates, for selection. Each candidate contains: +The deployment receives the optimal `PhysicalDAG` returned by selection. It +contains: * a **precompute DAG**, whose inputs are raw-sample contracts (rows carrying series labels, timestamp and value; the label set is the complete series @@ -161,7 +173,7 @@ physical candidates, for selection. Each candidate contains: raw-series contracts, or both; * the lifecycle, window framework and retention of every stored output. -After selection, the deployment binds each input contract, stores each precompute output under +The deployment binds each input contract, stores each precompute output under its own storage identity, and returns the query DAG's result. Semantic identity of stored outputs is defined by the logical DAG they compute. Storage identity and encoding belong to the deployment. @@ -182,17 +194,17 @@ Physical design produces `CandidatePostASAPDAGs` with materialization annotations, containing one candidate for each admissible lifecycle assignment of a logical candidate. These are Post-ASAP graphs with execution timing assigned, not a new DAG representation or a single selected plan. The -deployment prices the assignments and selects among the candidates. The timing +deployment prices the assignments and selection chooses among them. The timing frontier contains the ingestion-time nodes read by query-time nodes, plus an ingestion-time root. ## Cost and selection -The deployment decides both placement and the summary family. +The deployment prices; ASAPPlanner selects. Selection decides both placement and +the summary family, using only the deployment's prices. -* **Placement.** For each assignment, the deployment prices every state's - lifecycle with its own costs, then chooses the cheapest admissible - assignment for the whole workload: +* **Prices supplied by the deployment.** For each lifecycle assignment, the price + of every state's lifecycle: * build cost; * maintenance per update (precompute CPU); * reads; @@ -200,17 +212,18 @@ The deployment decides both placement and the summary family. the installed window layout × cardinality × store price; * retirement; * for `Ephemeral`, processing of the raw samples read at query time. + + The deployment may also supply whole-plan quotes for physical candidates. +* **Selection.** Selection ranks and combines the candidates and returns the + cheapest admissible plan for the whole workload that meets the accuracy + requirements. * **Sharing.** A state shared by several queries is priced once, with all consumers' demand. This holds only when compilation would install one shared output: same window layout, evaluation interval and phase. -* **Family.** Planner does not prune families: every summary-family candidate - (for example KLL and DDSketch for one quantile) reaches the deployment. The - deployment compares their physical candidates with its whole-plan quotes, - alongside lifecycle and placement choices. Planner's global-selection helpers - select only when a caller explicitly asks; no Planner layer calls them as an - intermediate stage. Known gap: ASAPQuery-backend #795 currently calls - `global_selection` before pricing and so pre-selects families; this is being - changed. +* **Family.** Planner does not prune families before selection: every + summary-family candidate (for example KLL and DDSketch for one quantile) + reaches selection, which compares them with the deployment's prices and + quotes. * Unknown cost stays unknown. It never becomes zero, and an alternative that cannot be priced is not selected. @@ -220,7 +233,8 @@ The following uses `sum by (job) (rate(m[1m]))`, evaluated every 10 s. Consider one logical candidate: a per-series Rate state feeding a grouped Sum state. Logical planning does not return placement variants. The -lifecycle layer lists choices for both states. The deployment prices them: +lifecycle layer lists choices for both states. The deployment prices them, and +selection chooses: * When summary storage is cheap, both states are retained. The precompute DAG builds Rate and Sum per pane, and the query DAG only reads the Sum. @@ -229,5 +243,5 @@ lifecycle layer lists choices for both states. The deployment prices them: * When storage is more expensive still, both states become `Ephemeral`. The query DAG reads raw series at `t_q`. -All three are cuts of one compilation. The deployment's decision is only which -lifecycle assignment to buy. +All three are cuts of one compilation. Only the deployment's store price +changes; selection picks the lifecycle assignment that is cheapest under it. From dc275fd65a66c9808d095c4482b82ec8690ddbcf Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 20:06:43 +0000 Subject: [PATCH 11/56] docs: address review of layering proposal Planner takes the query and data workloads plus the deployment's cost model, accuracy requirements and capabilities, and returns one optimal PhysicalDAG; candidate sets stay internal. Name the annotated stage MaterializedPostASAPDAG, tabulate what each DAG encodes, fold timing and selection into the layer descriptions, rename physical design to summary materialization, and make the example trace one query through every DAG. Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 304 +++++++----------- 1 file changed, 111 insertions(+), 193 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index a22e5868..9805eb22 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -5,166 +5,135 @@ as ASAPQuery-backend. ## Goal -ASAPPlanner decides what can be computed, how it is computed and which plan is -best, and hands the deployment one optimal physical plan to execute. The -deployment prices the candidates with its own costs, but it does not rank or -select: it passes its prices to ASAPPlanner's selection function and receives -the optimal plan. It then supplies data and state and runs that plan. It never -re-derives the computation, keeps its own operators, or implements its own -ranking or selection. +ASAPPlanner takes a query workload, a data workload and the deployment's +inputs, and returns one optimal physical plan. It decides what is computed, how +it is computed, and which plan is best. The deployment only supplies inputs and +executes the plan: it supplies its own cost model but never ranks or selects, +never re-derives the computation, and keeps no operators of its own. ## Layers ```text - PromQL / SQL / MetricsQL + Query workload (PromQL / SQL / MetricsQL, query repeating pattern, etc.) + + data workload + + deployment inputs: cost model, accuracy requirements, capabilities │ -┌───────────────────────── ASAPPlanner ────────────────────────────┐ +┌───────────────────────┴───────── ASAPPlanner ────────────────────┐ │ 0. Frontends │ │ Parse + lower -> CandidatePreASAPDAGs │ │ Reject unsupported constructs, such as PromQL fill. │ │ │ │ │ Logical planning (what) │ -│ 1. Logical Post-ASAP (asap-aware-mapping) │ -│ WHAT to compute; no placement. │ +│ 1. Logical optimization │ │ Summary families, rewrites, exact candidates │ │ -> CandidatePostASAPDAGs │ │ │ │ │ Physical planning (how) │ -│ 2. Physical design: summary materialization │ -│ Per unique summary state / maintained population, choose a │ -│ lifecycle: Ephemeral | Prepared | Shared | │ -│ ContinuouslyMaintained. Each assignment -> node timing, │ -│ window framework, retention. The only source of timing. │ -│ -> CandidatePostASAPDAGs with materialization annotations │ +│ 2. Summary materialization │ +│ Per summary state: Ephemeral | Prepared | Shared | │ +│ ContinuouslyMaintained -> node timing, window framework, │ +│ retention -> CandidateMaterializedPostASAPDAGs │ │ │ │ -│ 3. Physical implementation: compilation │ -│ Lower each Post-ASAP node to operators; cut by timing │ +│ 3. Compilation │ +│ Lower each node to operators; cut by timing │ │ -> CandidatePhysicalDAGs │ -│ Each candidate: { precompute DAG, query DAG, InputContracts } │ -│ Operators + kernels implement all computation. │ │ │ │ -│ 4. Selection (library function) │ -│ Input: deployment's prices, accuracy requirements and │ -│ capabilities. Rank and combine candidates for the workload. │ -│ -> one optimal PhysicalDAG │ +│ 4. Selection │ +│ Cost every candidate with the deployment's cost model; │ +│ choose the cheapest admissible one for the workload. │ └──────────────────────┬───────────────────────────────────────────┘ - optimal PhysicalDAG (the boundary) + one optimal PhysicalDAG (the boundary) ┌──────────────────────┴──────── Deployment ───────────────────────┐ -│ Pricing: computes the costs and quotes passed to selection │ │ 5. Execution: ingest, panes, storage, readout, run │ └──────────────────────────────────────────────────────────────────┘ ``` -### Candidate generation by ASAPPlanner - -Layers 0 to 3 output all legal candidates represented at their stage, and -selection is the only step that chooses: - -`CandidatePreASAPDAGs → CandidatePostASAPDAGs → CandidatePostASAPDAGs with materialization annotations → CandidatePhysicalDAGs → optimal PhysicalDAG` - -No layer before selection chooses a winning candidate. Frontend lowering can -produce a singleton candidate set for an unambiguous query; it need not invent -alternative parses. Logical planning exposes the supported computation -candidates; lifecycle enumeration exposes their admissible assignments; -physical compilation preserves their supported physical realizations and cuts. -Unsupported semantics and invalid or infeasible candidates are rejected with -reasons, rather than silently discarded by an intermediate cost selection. - -“All candidates” means the legal candidate space under the supplied semantics, -accuracy requirements, evidence, and capabilities. It can be represented -compactly or enumerated lazily; it does not require eagerly materializing the -Cartesian product. Selection chooses from this space using the deployment's -prices. Library helpers that return one selected `PlanOutput` without deployment -prices are optional convenience APIs, not stages of this pipeline. - -### DAG names - -This design distinguishes individual DAGs from the collections of candidates -passed between layers. The design names are the target API; main still uses -the earlier types, where one exists. - -| Design name | Meaning | Current main | Target API (implemented by open PRs #508, #480) | -|---|---|---|---| -| `PreASAPDAG` | Frontend-lowered query semantics before summary rewrites | `Rc` | `PreASAPDAG` | -| `CandidatePreASAPDAGs` | Frontend candidates, keyed by workload entry; deterministic frontends produce one per entry | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | -| `PostASAPDAG` | One shared logical computation graph | `Rc` tree; exported as `PostAsapDag` | `PostASAPDAG` | -| `CandidatePostASAPDAGs` | All legal logical candidates for the workload | `PlanSpace` | `CandidatePostASAPDAGs` | -| `CandidatePostASAPDAGs` with materialization annotations | The same Post-ASAP candidates, annotated with each admissible lifecycle assignment's timing, window framework and retention; still Post-ASAP DAGs, not physical ones | None; lifecycle helpers return one selected `SummaryMaintenanceLifecyclePlan` | `CandidatePostASAPDAGs` with timing | -| `PhysicalDAG` | Compiled operators and kernels with typed inputs, before runtime sources are bound | None | `PhysicalDAG` | -| `CandidatePhysicalDAGs` | Physical candidates with their timing cuts, metadata and diagnostics, before selection | None | `CandidatePhysicalDAGs` | -| optimal `PhysicalDAG` | The selected candidate: its precompute and query `PhysicalDAG`s with every stored output's lifecycle, window framework and retention | None | one materialized element of `CandidatePhysicalDAGs` | - -A candidate collection shares graphs across its candidates; it is not a copy -of every complete DAG. Timing is attached to the shared logical graph, not -stored in a second graph representation. Compatible timing assignments share -one physical compilation, and each candidate's precompute and query DAGs are -cut from it on demand. The exported `PostAsapDag` form is an explicit -export/import format, not an additional planning layer (see -[Post-ASAP IR](../concepts/post-asap-ir.md#tree-and-exported-dag-forms)). - -There are only two kinds of DAG after the frontend: Post-ASAP DAGs, whose -nodes are logical operations (optionally carrying materialization annotations), -and physical DAGs, whose nodes are selected physical operators. +Layers 0 to 3 each output a candidate set, `Candidates`, holding every +legal candidate of their stage (for example KLL and DDSketch for one quantile, +or each lifecycle assignment), and selection is the only step that chooses. Candidate sets are internal to ASAPPlanner: they may +be shared or enumerated lazily, and the deployment never sees them. Unsupported +or infeasible candidates are rejected with reasons, not silently dropped. + +## DAGs and what each encodes + +Each stage adds decisions to the DAG it receives. The table shows which +decisions each DAG carries. + +| | `PreASAPDAG` | `PostASAPDAG` | `MaterializedPostASAPDAG` | `PhysicalDAG` | +|---|---|---|---|---| +| Produced by | 0. Frontends | 1. Logical optimization | 2. Summary materialization | 3. Compilation; 4. selects one | +| Node | Query operation | Logical operation, including summary operations | Same, plus annotations | Physical operator | +| Logical optimization (summary family, rewrites) | No | Yes | Yes | Yes | +| Materialization decided (which summary states persist) | No | No | Yes | Yes | +| Data lifecycle (how each state is maintained: `Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) | No | No | Yes | Yes | +| Retention (how long each state lives) and window framework | No | No | Yes | Yes | +| Execution time (ingestion or query) | No | No | Yes, per node | Yes, as the precompute / query split | +| Physical optimization (operator choice, e.g. TopK as sort + limit) | No | No | No | Yes | +| Seen by the deployment | No | No | No | Only the selected one | + +Name mapping to code: + +| Design name | Current main | Target API (open PRs #508, #480) | +|---|---|---| +| `PreASAPDAG` | `Rc` | `PreASAPDAG` | +| `PostASAPDAG` | `Rc` tree; exported as `PostAsapDag` | `PostASAPDAG` | +| `MaterializedPostASAPDAG` | `SummaryMaintenanceLifecyclePlan`, one selected assignment beside the DAG | `MaterializedPostASAPDAG`; `SummaryMaintenanceLifecyclePlan` is merged into it | +| `PhysicalDAG` | None | `PhysicalDAG` | +| `CandidatePreASAPDAGs` | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | +| `CandidatePostASAPDAGs` | `PlanSpace` | `CandidatePostASAPDAGs` | +| `CandidateMaterializedPostASAPDAGs` | None | `CandidateMaterializedPostASAPDAGs` (#480 currently names it `CandidatePostASAPDAGsWithTiming`) | +| `CandidatePhysicalDAGs` | None | `CandidatePhysicalDAGs` | + +**Post-ASAP DAG to materialized DAG.** In this document, *materialization* +means summary state kept across executions, like a materialized view. Choosing +it is a workload-level decision, like a database's materialized-view +selection: it spans queries (a shared state is kept once) and depends on +workload demand (read and update rates, horizon). A lifecycle assignment +annotates every node with execution timing, and every stored state with window +framework and retention: + +* a retained state (`ContinuouslyMaintained`, `Shared`, `Prepared`) and every + node feeding it run at ingestion time; +* readouts, other consumers and `Ephemeral` states run at query time; +* an `Ephemeral` state that feeds a retained state runs at ingestion time, + because query-time work may not feed ingestion-time work. + +The result is still a Post-ASAP DAG, much as physical properties annotate +logical expressions in a database optimizer. This is the only source of +timing; logical optimization proposes computations, never timing or placement. + +**Materialized DAG to physical DAG.** Compilation lowers each node to physical +operators and cuts the graph at the timing frontier (ingestion-time nodes read +by query-time nodes, plus an ingestion-time root) into a precompute and a query +DAG. Materialization is decided in layer 2 and realized here: the precompute +DAG's outputs at the cut are the materialized states, and the query DAG reads +them through typed input slots. A physical DAG corresponds to its +Post-ASAP DAG node by node; a node may expand into several operators, whose +helper operators are numbered from their source node. The one exception, a +`Fallback` node wrapping a whole Pre-ASAP expression, is removed by the +operator-flattening proposal ([operator sharing](operator-sharing.md), #469, +#481). Operator materialization (sort, aggregation, summary build) and +computing a shared subexpression once are compilation and runtime details. Binding runtime sources is an execution step of a `PhysicalDAG`, not another -DAG. At execution the deployment supplies a source for each typed input slot, -the slots are checked against their contracts, and the graph runs; the bound -instance lives only for that execution and is not persisted or compared. - -### Physical planning - -Physical planning has two parts of different character, as in databases. - -**Physical design (summary materialization).** In this document, -*materialization* means summary state kept across executions, like a -materialized view: whether a state is maintained, how it is refreshed and how -long it is retained. This is a workload-level decision, like a database's -materialized-view selection: it spans queries (shared state is kept once), it -depends on workload demand (read and update rates, horizon), and its result -persists. The Planner enumerates the admissible choices; the deployment prices -them, and selection chooses. The choice is recorded as annotations on the -Post-ASAP nodes (timing, window framework, retention), so the annotated graph is still a -Post-ASAP DAG, much as physical properties annotate logical expressions in a -database optimizer. Operator materialization (blocking operators such as sort, -aggregation or summary build) and computing a shared subexpression once are -details of compilation and the runtime, not separate layers. - -**Physical implementation (compilation).** Compilation lowers each Post-ASAP -node to physical operators and cuts the graph by timing into a precompute and a -query DAG. Materialization is *decided* by physical design and *realized* here: -the precompute DAG's outputs at the cut are the materialized states, and the -query DAG reads them through typed input slots. A physical DAG corresponds to -its Post-ASAP DAG node by node. A node may expand into several operators (for -example TopK into sort then limit, or a multi-input merge into union then -merge); helper operators are numbered from their source node, so every operator -traces back to one Post-ASAP node. One exception is a `Fallback` node, which -wraps a whole Pre-ASAP expression and compiles to many operators; the -operator-flattening proposal ([operator sharing](operator-sharing.md), from -#469) removes it by making non-ASAP operators ordinary Post-ASAP nodes, and -#481 revises its export to emit one Post-ASAP node per non-ASAP operator. If a node gains alternative physical implementations, they become further candidates in -`CandidatePhysicalDAGs`, priced by the deployment and chosen by selection. - -### Responsibilities +DAG: the deployment supplies a source for each typed input slot, the slots are +checked against their contracts, and the graph runs. + +## Responsibilities | Layer | Owns | Does not own | |---|---|---| -| 0. Frontends | Language semantics and lowering into `CandidatePreASAPDAGs`. A construct that cannot be represented faithfully is rejected, never ignored (for example PromQL `fill`). | Summaries, placement | -| 1. Logical Post-ASAP | `CandidatePostASAPDAGs`: all legal logical candidates, including summary families, exact rewrites, compositions, and series-identity typing. | Placement | -| 2. Physical design: summary materialization | For each unique summary state and maintained population, the lifecycle choices (`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) and their costs under a caller-supplied cost model. Each assignment sets every node's execution timing, window framework and retention, producing `CandidatePostASAPDAGs` with materialization annotations. | The cost values themselves; operator implementation | -| 3. Physical implementation: compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. Lowers each node of `CandidatePostASAPDAGs` with materialization annotations to operators and cuts by timing into `CandidatePhysicalDAGs`; it does not select a winner. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | -| 4. Selection (Planner library function) | Ranking and combining all candidates, including every summary family and lifecycle assignment, across the workload with the deployment's prices, counting shared state once; returning one optimal `PhysicalDAG`. Candidates that cannot be priced are not selected. | The prices themselves | -| Deployment pricing | Computing the prices passed to selection: unit costs, whole-plan quotes, store price; supplying accuracy requirements and capabilities (for example whether query-time raw data is available). | Ranking, sorting or selection | -| 5. Deployment execution | Ingestion and routing, pane assignment and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | - -The deployment may call Planner's earlier-layer APIs, for example to inspect -candidates while pricing them. The rules are that every computation runs as a -Planner-compiled physical DAG and every choice among candidates is made by -Planner's selection function. +| 0. Frontends | Language semantics and lowering. A construct that cannot be represented faithfully is rejected, never ignored (for example PromQL `fill`). | Summaries, placement | +| 1. Logical optimization | All legal logical candidates: summary families, exact rewrites, compositions, series-identity typing. | Placement, timing | +| 2. Summary materialization | For each unique summary state and maintained population, the admissible lifecycle assignments and their timing, window framework and retention. | Cost values; operator implementation | +| 3. Compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | +| 4. Selection | Costing every candidate with the deployment's cost model and returning the cheapest admissible `PhysicalDAG` for the whole workload that meets the accuracy requirements. A state shared by several queries is costed once with all consumers' demand (only when compilation installs one shared output: same window layout, evaluation interval and phase). Unknown cost stays unknown and such a candidate is not selected. | The cost values | +| 5. Deployment | Inputs: the cost model (build, per-update maintenance, read, store price per byte-second, retirement, query-time raw processing; optionally whole-plan quotes), accuracy requirements and capabilities (for example whether query-time raw data is available). Execution: ingestion and routing, panes and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Ranking or selection; any computation algorithm | ## The boundary -The deployment receives the optimal `PhysicalDAG` returned by selection. It -contains: +The deployment passes its inputs to ASAPPlanner and receives one optimal +`PhysicalDAG`, which contains: * a **precompute DAG**, whose inputs are raw-sample contracts (rows carrying series labels, timestamp and value; the label set is the complete series @@ -178,70 +147,19 @@ its own storage identity, and returns the query DAG's result. Semantic identity of stored outputs is defined by the logical DAG they compute. Storage identity and encoding belong to the deployment. -## Timing and placement +## Example -Timing comes only from physical design, that is, from the chosen lifecycle assignment. Logical strategies propose -computation candidates, not execution timing or placement. The rules for -applying an assignment are: +This traces `sum by (job) (rate(m[1m]))`, evaluated every 10 s, through the +four DAGs, and shows that only the deployment's store price changes the plan. -* a retained state (`ContinuouslyMaintained`, `Shared`, `Prepared`) and every - node feeding it run at ingestion time; -* readouts, other consumers and `Ephemeral` states run at query time; -* an `Ephemeral` state that feeds a retained state runs at ingestion time, - because query-time work may not feed ingestion-time work. +* `PreASAPDAG`: `sum by (job)` over `rate` over the range selector `m[1m]`. +* `PostASAPDAG`: a per-series Rate state feeding a grouped Sum state. +* `MaterializedPostASAPDAG`: one per lifecycle assignment, for example + (a) both retained, (b) Rate retained and Sum `Ephemeral`, (c) both + `Ephemeral`. +* `PhysicalDAG`: one compilation, cut three ways. (a) Precompute builds Rate + and Sum per pane; the query only reads Sum. (b) Precompute keeps Rate; the + query builds Sum. (c) No precompute; the query reads raw series at `t_q`. -Physical design produces `CandidatePostASAPDAGs` with materialization -annotations, containing one candidate for each admissible lifecycle assignment -of a logical candidate. These are Post-ASAP graphs with execution timing -assigned, not a new DAG representation or a single selected plan. The -deployment prices the assignments and selection chooses among them. The timing -frontier contains the ingestion-time nodes read by query-time nodes, plus an -ingestion-time root. - -## Cost and selection - -The deployment prices; ASAPPlanner selects. Selection decides both placement and -the summary family, using only the deployment's prices. - -* **Prices supplied by the deployment.** For each lifecycle assignment, the price - of every state's lifecycle: - * build cost; - * maintenance per update (precompute CPU); - * reads; - * retention: summary store cost, meaning state bytes × retained panes of - the installed window layout × cardinality × store price; - * retirement; - * for `Ephemeral`, processing of the raw samples read at query time. - - The deployment may also supply whole-plan quotes for physical candidates. -* **Selection.** Selection ranks and combines the candidates and returns the - cheapest admissible plan for the whole workload that meets the accuracy - requirements. -* **Sharing.** A state shared by several queries is priced once, with all - consumers' demand. This holds only when compilation would install one shared - output: same window layout, evaluation interval and phase. -* **Family.** Planner does not prune families before selection: every - summary-family candidate (for example KLL and DDSketch for one quantile) - reaches selection, which compares them with the deployment's prices and - quotes. -* Unknown cost stays unknown. It never becomes zero, and an alternative that - cannot be priced is not selected. - -## Example: store cost changes placement - -The following uses `sum by (job) (rate(m[1m]))`, evaluated every 10 s. - -Consider one logical candidate: a per-series Rate state feeding a -grouped Sum state. Logical planning does not return placement variants. The -lifecycle layer lists choices for both states. The deployment prices them, and -selection chooses: - -* When summary storage is cheap, both states are retained. The precompute DAG - builds Rate and Sum per pane, and the query DAG only reads the Sum. -* When storage is expensive enough, Sum becomes `Ephemeral`. Precompute keeps - only Rate, and the query DAG builds Sum at query time. -* When storage is more expensive still, both states become `Ephemeral`. The - query DAG reads raw series at `t_q`. - -All three are cuts of one compilation. Only the deployment's store price -changes; selection picks the lifecycle assignment that is cheapest under it. +Selection returns (a) when storage is cheap, (b) when it is expensive, and (c) +when it is more expensive still. The deployment only changed its store price. From a095ce510b7ed23abda65fdf99baffde0d7bc1f5 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 20:22:50 +0000 Subject: [PATCH 12/56] docs: name layer 2 summary lifecycle planning; drop deployment ranking row A lifecycle fixes materialization, timing, maintenance, retention and window framework together, so the stage and its DAG are named after the lifecycle (LifecyclePostASAPDAG) rather than one of those aspects. The deployment row no longer mentions ranking or selection. Co-Authored-By: Claude Opus 5.5 --- .../proposals/planner-backend-layering.md | 38 ++++++++++--------- 1 file changed, 21 insertions(+), 17 deletions(-) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-backend-layering.md index 9805eb22..9642453d 100644 --- a/docs/design_docs/proposals/planner-backend-layering.md +++ b/docs/design_docs/proposals/planner-backend-layering.md @@ -29,10 +29,10 @@ never re-derives the computation, and keeps no operators of its own. │ -> CandidatePostASAPDAGs │ │ │ │ │ Physical planning (how) │ -│ 2. Summary materialization │ +│ 2. Summary lifecycle planning │ │ Per summary state: Ephemeral | Prepared | Shared | │ │ ContinuouslyMaintained -> node timing, window framework, │ -│ retention -> CandidateMaterializedPostASAPDAGs │ +│ retention -> CandidateLifecyclePostASAPDAGs │ │ │ │ │ 3. Compilation │ │ Lower each node to operators; cut by timing │ @@ -59,9 +59,9 @@ or infeasible candidates are rejected with reasons, not silently dropped. Each stage adds decisions to the DAG it receives. The table shows which decisions each DAG carries. -| | `PreASAPDAG` | `PostASAPDAG` | `MaterializedPostASAPDAG` | `PhysicalDAG` | +| | `PreASAPDAG` | `PostASAPDAG` | `LifecyclePostASAPDAG` | `PhysicalDAG` | |---|---|---|---|---| -| Produced by | 0. Frontends | 1. Logical optimization | 2. Summary materialization | 3. Compilation; 4. selects one | +| Produced by | 0. Frontends | 1. Logical optimization | 2. Summary lifecycle planning | 3. Compilation; 4. selects one | | Node | Query operation | Logical operation, including summary operations | Same, plus annotations | Physical operator | | Logical optimization (summary family, rewrites) | No | Yes | Yes | Yes | | Materialization decided (which summary states persist) | No | No | Yes | Yes | @@ -77,20 +77,24 @@ Name mapping to code: |---|---|---| | `PreASAPDAG` | `Rc` | `PreASAPDAG` | | `PostASAPDAG` | `Rc` tree; exported as `PostAsapDag` | `PostASAPDAG` | -| `MaterializedPostASAPDAG` | `SummaryMaintenanceLifecyclePlan`, one selected assignment beside the DAG | `MaterializedPostASAPDAG`; `SummaryMaintenanceLifecyclePlan` is merged into it | +| `LifecyclePostASAPDAG` | `SummaryMaintenanceLifecyclePlan`, one selected assignment beside the DAG | `LifecyclePostASAPDAG`; `SummaryMaintenanceLifecyclePlan` is merged into it | | `PhysicalDAG` | None | `PhysicalDAG` | | `CandidatePreASAPDAGs` | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | | `CandidatePostASAPDAGs` | `PlanSpace` | `CandidatePostASAPDAGs` | -| `CandidateMaterializedPostASAPDAGs` | None | `CandidateMaterializedPostASAPDAGs` (#480 currently names it `CandidatePostASAPDAGsWithTiming`) | +| `CandidateLifecyclePostASAPDAGs` | None | `CandidateLifecyclePostASAPDAGs` | | `CandidatePhysicalDAGs` | None | `CandidatePhysicalDAGs` | -**Post-ASAP DAG to materialized DAG.** In this document, *materialization* -means summary state kept across executions, like a materialized view. Choosing -it is a workload-level decision, like a database's materialized-view -selection: it spans queries (a shared state is kept once) and depends on -workload demand (read and update rates, horizon). A lifecycle assignment -annotates every node with execution timing, and every stored state with window -framework and retention: +**Post-ASAP DAG to lifecycle DAG.** A summary state's *lifecycle* +(`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) fixes several +separate aspects together: whether the state is materialized (kept across +executions, like a materialized view), when it is computed, how it is +maintained, how long it is retained, and its window framework. The table above +lists these aspects separately; the lifecycle is the one choice that sets them. +Choosing lifecycles is a workload-level decision, like a database's +materialized-view selection: it spans queries (a shared state is kept once) and +depends on workload demand (read and update rates, horizon). A lifecycle +assignment annotates every node with execution timing, and every stored state +with window framework and retention: * a retained state (`ContinuouslyMaintained`, `Shared`, `Prepared`) and every node feeding it run at ingestion time; @@ -102,7 +106,7 @@ The result is still a Post-ASAP DAG, much as physical properties annotate logical expressions in a database optimizer. This is the only source of timing; logical optimization proposes computations, never timing or placement. -**Materialized DAG to physical DAG.** Compilation lowers each node to physical +**Lifecycle DAG to physical DAG.** Compilation lowers each node to physical operators and cuts the graph at the timing frontier (ingestion-time nodes read by query-time nodes, plus an ingestion-time root) into a precompute and a query DAG. Materialization is decided in layer 2 and realized here: the precompute @@ -125,10 +129,10 @@ checked against their contracts, and the graph runs. |---|---|---| | 0. Frontends | Language semantics and lowering. A construct that cannot be represented faithfully is rejected, never ignored (for example PromQL `fill`). | Summaries, placement | | 1. Logical optimization | All legal logical candidates: summary families, exact rewrites, compositions, series-identity typing. | Placement, timing | -| 2. Summary materialization | For each unique summary state and maintained population, the admissible lifecycle assignments and their timing, window framework and retention. | Cost values; operator implementation | +| 2. Summary lifecycle planning | For each unique summary state and maintained population, the admissible lifecycle assignments and their timing, window framework and retention. | Cost values; operator implementation | | 3. Compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | | 4. Selection | Costing every candidate with the deployment's cost model and returning the cheapest admissible `PhysicalDAG` for the whole workload that meets the accuracy requirements. A state shared by several queries is costed once with all consumers' demand (only when compilation installs one shared output: same window layout, evaluation interval and phase). Unknown cost stays unknown and such a candidate is not selected. | The cost values | -| 5. Deployment | Inputs: the cost model (build, per-update maintenance, read, store price per byte-second, retirement, query-time raw processing; optionally whole-plan quotes), accuracy requirements and capabilities (for example whether query-time raw data is available). Execution: ingestion and routing, panes and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Ranking or selection; any computation algorithm | +| 5. Deployment | Inputs: the cost model (build, per-update maintenance, read, store price per byte-second, retirement, query-time raw processing; optionally whole-plan quotes), accuracy requirements and capabilities (for example whether query-time raw data is available). Execution: ingestion and routing, panes and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | ## The boundary @@ -154,7 +158,7 @@ four DAGs, and shows that only the deployment's store price changes the plan. * `PreASAPDAG`: `sum by (job)` over `rate` over the range selector `m[1m]`. * `PostASAPDAG`: a per-series Rate state feeding a grouped Sum state. -* `MaterializedPostASAPDAG`: one per lifecycle assignment, for example +* `LifecyclePostASAPDAG`: one per lifecycle assignment, for example (a) both retained, (b) Rate retained and Sum `Ephemeral`, (c) both `Ephemeral`. * `PhysicalDAG`: one compilation, cut three ways. (a) Precompute builds Rate From a3c6dbcd064a8039b0795f49838c6c647a16fc92 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 21:54:12 +0000 Subject: [PATCH 13/56] docs: rename layering proposal to planner-layering.md Co-Authored-By: Claude Opus 5.5 --- docs/design_docs/proposals/README.md | 2 +- .../{planner-backend-layering.md => planner-layering.md} | 0 2 files changed, 1 insertion(+), 1 deletion(-) rename docs/design_docs/proposals/{planner-backend-layering.md => planner-layering.md} (100%) diff --git a/docs/design_docs/proposals/README.md b/docs/design_docs/proposals/README.md index dd0cdadc..5a96c6a8 100644 --- a/docs/design_docs/proposals/README.md +++ b/docs/design_docs/proposals/README.md @@ -10,4 +10,4 @@ extensions. A design document is not a promise of downstream runtime support. - [ASAP-aware mapping proposals](asap-aware-mapping/README.md) - [Operator sharing](operator-sharing.md) - [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) -- [Planner and deployment layering](planner-backend-layering.md) +- [Planner and deployment layering](planner-layering.md) diff --git a/docs/design_docs/proposals/planner-backend-layering.md b/docs/design_docs/proposals/planner-layering.md similarity index 100% rename from docs/design_docs/proposals/planner-backend-layering.md rename to docs/design_docs/proposals/planner-layering.md From 41076ca602cd0d8ec6e21130348a608ba2bca8c6 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 21:55:35 +0000 Subject: [PATCH 14/56] docs: name the DAG stages LogicalPostASAPDAG and PhysicalPostASAPDAG Every stage after the frontend is a Post-ASAP DAG; the prefix says how far planning has gone: Logical, Lifecycle, Physical. Co-Authored-By: Claude Opus 5.5 --- .../design_docs/proposals/planner-layering.md | 34 +++++++++---------- 1 file changed, 17 insertions(+), 17 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 9642453d..f6355166 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -26,7 +26,7 @@ never re-derives the computation, and keeps no operators of its own. │ Logical planning (what) │ │ 1. Logical optimization │ │ Summary families, rewrites, exact candidates │ -│ -> CandidatePostASAPDAGs │ +│ -> CandidateLogicalPostASAPDAGs │ │ │ │ │ Physical planning (how) │ │ 2. Summary lifecycle planning │ @@ -36,13 +36,13 @@ never re-derives the computation, and keeps no operators of its own. │ │ │ │ 3. Compilation │ │ Lower each node to operators; cut by timing │ -│ -> CandidatePhysicalDAGs │ +│ -> CandidatePhysicalPostASAPDAGs │ │ │ │ │ 4. Selection │ │ Cost every candidate with the deployment's cost model; │ │ choose the cheapest admissible one for the workload. │ └──────────────────────┬───────────────────────────────────────────┘ - one optimal PhysicalDAG (the boundary) + one optimal PhysicalPostASAPDAG (the boundary) ┌──────────────────────┴──────── Deployment ───────────────────────┐ │ 5. Execution: ingest, panes, storage, readout, run │ └──────────────────────────────────────────────────────────────────┘ @@ -59,7 +59,7 @@ or infeasible candidates are rejected with reasons, not silently dropped. Each stage adds decisions to the DAG it receives. The table shows which decisions each DAG carries. -| | `PreASAPDAG` | `PostASAPDAG` | `LifecyclePostASAPDAG` | `PhysicalDAG` | +| | `PreASAPDAG` | `LogicalPostASAPDAG` | `LifecyclePostASAPDAG` | `PhysicalPostASAPDAG` | |---|---|---|---|---| | Produced by | 0. Frontends | 1. Logical optimization | 2. Summary lifecycle planning | 3. Compilation; 4. selects one | | Node | Query operation | Logical operation, including summary operations | Same, plus annotations | Physical operator | @@ -76,13 +76,13 @@ Name mapping to code: | Design name | Current main | Target API (open PRs #508, #480) | |---|---|---| | `PreASAPDAG` | `Rc` | `PreASAPDAG` | -| `PostASAPDAG` | `Rc` tree; exported as `PostAsapDag` | `PostASAPDAG` | +| `LogicalPostASAPDAG` | `Rc` tree; exported as `PostAsapDag` | `LogicalPostASAPDAG` | | `LifecyclePostASAPDAG` | `SummaryMaintenanceLifecyclePlan`, one selected assignment beside the DAG | `LifecyclePostASAPDAG`; `SummaryMaintenanceLifecyclePlan` is merged into it | -| `PhysicalDAG` | None | `PhysicalDAG` | +| `PhysicalPostASAPDAG` | None | `PhysicalPostASAPDAG` | | `CandidatePreASAPDAGs` | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | -| `CandidatePostASAPDAGs` | `PlanSpace` | `CandidatePostASAPDAGs` | +| `CandidateLogicalPostASAPDAGs` | `PlanSpace` | `CandidateLogicalPostASAPDAGs` | | `CandidateLifecyclePostASAPDAGs` | None | `CandidateLifecyclePostASAPDAGs` | -| `CandidatePhysicalDAGs` | None | `CandidatePhysicalDAGs` | +| `CandidatePhysicalPostASAPDAGs` | None | `CandidatePhysicalPostASAPDAGs` | **Post-ASAP DAG to lifecycle DAG.** A summary state's *lifecycle* (`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) fixes several @@ -119,9 +119,9 @@ operator-flattening proposal ([operator sharing](operator-sharing.md), #469, #481). Operator materialization (sort, aggregation, summary build) and computing a shared subexpression once are compilation and runtime details. -Binding runtime sources is an execution step of a `PhysicalDAG`, not another -DAG: the deployment supplies a source for each typed input slot, the slots are -checked against their contracts, and the graph runs. +Binding runtime sources is an execution step of a `PhysicalPostASAPDAG`, not +another DAG: the deployment supplies a source for each typed input slot, the +slots are checked against their contracts, and the graph runs. ## Responsibilities @@ -131,13 +131,13 @@ checked against their contracts, and the graph runs. | 1. Logical optimization | All legal logical candidates: summary families, exact rewrites, compositions, series-identity typing. | Placement, timing | | 2. Summary lifecycle planning | For each unique summary state and maintained population, the admissible lifecycle assignments and their timing, window framework and retention. | Cost values; operator implementation | | 3. Compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | -| 4. Selection | Costing every candidate with the deployment's cost model and returning the cheapest admissible `PhysicalDAG` for the whole workload that meets the accuracy requirements. A state shared by several queries is costed once with all consumers' demand (only when compilation installs one shared output: same window layout, evaluation interval and phase). Unknown cost stays unknown and such a candidate is not selected. | The cost values | +| 4. Selection | Costing every candidate with the deployment's cost model and returning the cheapest admissible `PhysicalPostASAPDAG` for the whole workload that meets the accuracy requirements. A state shared by several queries is costed once with all consumers' demand (only when compilation installs one shared output: same window layout, evaluation interval and phase). Unknown cost stays unknown and such a candidate is not selected. | The cost values | | 5. Deployment | Inputs: the cost model (build, per-update maintenance, read, store price per byte-second, retirement, query-time raw processing; optionally whole-plan quotes), accuracy requirements and capabilities (for example whether query-time raw data is available). Execution: ingestion and routing, panes and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | ## The boundary The deployment passes its inputs to ASAPPlanner and receives one optimal -`PhysicalDAG`, which contains: +`PhysicalPostASAPDAG`, which contains: * a **precompute DAG**, whose inputs are raw-sample contracts (rows carrying series labels, timestamp and value; the label set is the complete series @@ -157,13 +157,13 @@ This traces `sum by (job) (rate(m[1m]))`, evaluated every 10 s, through the four DAGs, and shows that only the deployment's store price changes the plan. * `PreASAPDAG`: `sum by (job)` over `rate` over the range selector `m[1m]`. -* `PostASAPDAG`: a per-series Rate state feeding a grouped Sum state. +* `LogicalPostASAPDAG`: a per-series Rate state feeding a grouped Sum state. * `LifecyclePostASAPDAG`: one per lifecycle assignment, for example (a) both retained, (b) Rate retained and Sum `Ephemeral`, (c) both `Ephemeral`. -* `PhysicalDAG`: one compilation, cut three ways. (a) Precompute builds Rate - and Sum per pane; the query only reads Sum. (b) Precompute keeps Rate; the - query builds Sum. (c) No precompute; the query reads raw series at `t_q`. +* `PhysicalPostASAPDAG`: one compilation, cut three ways. (a) Precompute builds + Rate and Sum per pane; the query only reads Sum. (b) Precompute keeps Rate; + the query builds Sum. (c) No precompute; the query reads raw series at `t_q`. Selection returns (a) when storage is cheap, (b) when it is expensive, and (c) when it is more expensive still. The deployment only changed its store price. From a467e1d20b321a16d17ec4cb062ce1578b7ee713 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:08:09 -0400 Subject: [PATCH 15/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index f6355166..57acdfca 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -14,9 +14,9 @@ never re-derives the computation, and keeps no operators of its own. ## Layers ```text - Query workload (PromQL / SQL / MetricsQL, query repeating pattern, etc.) - + data workload - + deployment inputs: cost model, accuracy requirements, capabilities + Query workload (PromQL / SQL / MetricsQL, query repeating pattern, accuracy requirements, query latency requirement) + + data workload (data arrival pattern: streaming data vs data at rest, data distribution, cardinality) + + deployment inputs: empirical cost estimation, empirical accuracy estimation, capabilities │ ┌───────────────────────┴───────── ASAPPlanner ────────────────────┐ │ 0. Frontends │ From 3b2ce8e887a033ce407a5157bd3da0b613137aa8 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Wed, 30 Sep 2026 18:09:31 -0400 Subject: [PATCH 16/56] Update README.md --- docs/design_docs/proposals/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/README.md b/docs/design_docs/proposals/README.md index 5a96c6a8..cb44fadb 100644 --- a/docs/design_docs/proposals/README.md +++ b/docs/design_docs/proposals/README.md @@ -10,4 +10,4 @@ extensions. A design document is not a promise of downstream runtime support. - [ASAP-aware mapping proposals](asap-aware-mapping/README.md) - [Operator sharing](operator-sharing.md) - [Decoupling operators from scalar expressions](decoupling_op_and_expr.md) -- [Planner and deployment layering](planner-layering.md) +- [ASAPPlanner layering](planner-layering.md) From 0c88b6de46a052f79eec950242da945947ab0c9f Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 09:00:05 -0400 Subject: [PATCH 17/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 11 +++++------ 1 file changed, 5 insertions(+), 6 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 57acdfca..3e014360 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -1,6 +1,6 @@ -# Planner and deployment layering +# ASAPPlanner Layering Design -Status: proposal. Audience: designers of ASAPPlanner and of deployments such +Status: proposal. Audience: designers and developers of ASAPPlanner and of deployments such as ASAPQuery-backend. ## Goal @@ -8,15 +8,14 @@ as ASAPQuery-backend. ASAPPlanner takes a query workload, a data workload and the deployment's inputs, and returns one optimal physical plan. It decides what is computed, how it is computed, and which plan is best. The deployment only supplies inputs and -executes the plan: it supplies its own cost model but never ranks or selects, -never re-derives the computation, and keeps no operators of its own. +executes the plan: it supplies its own empirical cost estimation, empirical accuracy estimation and capabilities of deployment but never does the query planning or plan selection. ## Layers ```text Query workload (PromQL / SQL / MetricsQL, query repeating pattern, accuracy requirements, query latency requirement) - + data workload (data arrival pattern: streaming data vs data at rest, data distribution, cardinality) - + deployment inputs: empirical cost estimation, empirical accuracy estimation, capabilities + + Data workload (data arrival pattern: streaming data vs data at rest, data distribution, cardinality) + + Deployment inputs: empirical cost estimation, empirical accuracy estimation, capabilities │ ┌───────────────────────┴───────── ASAPPlanner ────────────────────┐ │ 0. Frontends │ From 037f5be760e86cbd1f3c96dc4215cd9d5ee1dee9 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 09:02:22 -0400 Subject: [PATCH 18/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 3e014360..e494dea6 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -5,15 +5,15 @@ as ASAPQuery-backend. ## Goal -ASAPPlanner takes a query workload, a data workload and the deployment's -inputs, and returns one optimal physical plan. It decides what is computed, how +ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the [deployment's +inputs], and returns one optimal physical plan. It decides what is computed, how it is computed, and which plan is best. The deployment only supplies inputs and executes the plan: it supplies its own empirical cost estimation, empirical accuracy estimation and capabilities of deployment but never does the query planning or plan selection. ## Layers ```text - Query workload (PromQL / SQL / MetricsQL, query repeating pattern, accuracy requirements, query latency requirement) + Query workload (PromQL/SQL/MetricsQL, query repeating pattern, accuracy requirements, query latency requirement) + Data workload (data arrival pattern: streaming data vs data at rest, data distribution, cardinality) + Deployment inputs: empirical cost estimation, empirical accuracy estimation, capabilities │ From d90e113dbef8fa1ff36c5a98dcf793478a03d575 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 09:05:04 -0400 Subject: [PATCH 19/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index e494dea6..d91fc018 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -6,7 +6,7 @@ as ASAPQuery-backend. ## Goal ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the [deployment's -inputs], and returns one optimal physical plan. It decides what is computed, how +inputs](TODO: A data structure should be explicitly defined in PR #510), and returns one optimal physical plan. It decides what is computed, how it is computed, and which plan is best. The deployment only supplies inputs and executes the plan: it supplies its own empirical cost estimation, empirical accuracy estimation and capabilities of deployment but never does the query planning or plan selection. From 0e0301c2c6777d3f02f6983009dde211b85cafee Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 09:19:25 -0400 Subject: [PATCH 20/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 34 ++++++++----------- 1 file changed, 15 insertions(+), 19 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index d91fc018..391d2960 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -19,31 +19,27 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu │ ┌───────────────────────┴───────── ASAPPlanner ────────────────────┐ │ 0. Frontends │ -│ Parse + lower -> CandidatePreASAPDAGs │ -│ Reject unsupported constructs, such as PromQL fill. │ +│ Parse + lower + -> Output: LogicalDAG, CandidateLogicalDAGs │ +│ Reject unsupported query expressions. │ │ │ │ -│ Logical planning (what) │ -│ 1. Logical optimization │ -│ Summary families, rewrites, exact candidates │ -│ -> CandidateLogicalPostASAPDAGs │ +│ Logical planning (what to compute) │ +│ 1. Logical ASAP-aware optimization │ +│ ASAPPrimitives: Summary families x exact candidates x query rewrites x supporting multiple computation nodes with one summary node │ +│ -> Output: LogicalDAGswithASAPPrimitive, CandidateLogicalDAGswithASAPPrimitive │ │ │ │ -│ Physical planning (how) │ -│ 2. Summary lifecycle planning │ -│ Per summary state: Ephemeral | Prepared | Shared | │ -│ ContinuouslyMaintained -> node timing, window framework, │ -│ retention -> CandidateLifecyclePostASAPDAGs │ +│ Physical planning (how to compute) │ │ +│ 2. Physical ASAP-aware optimization | + Materialization or not for a subDAG x Physical operator implementation selection x Parallelism & Partitioning x resource management | + -> Output: PhysicalDAGwithASAPPrimitives, CandidatePhysicalDAGswithASAPPrimitives │ │ │ │ -│ 3. Compilation │ -│ Lower each node to operators; cut by timing │ -│ -> CandidatePhysicalPostASAPDAGs │ -│ │ │ -│ 4. Selection │ -│ Cost every candidate with the deployment's cost model; │ +│ 3. Plan Selection │ +│ Cost every candidate with the deployment's cost model and accuracy model; │ │ choose the cheapest admissible one for the workload. │ └──────────────────────┬───────────────────────────────────────────┘ - one optimal PhysicalPostASAPDAG (the boundary) + one optimal PhysicalDAGwithASAPPrimitives (the boundary) ┌──────────────────────┴──────── Deployment ───────────────────────┐ -│ 5. Execution: ingest, panes, storage, readout, run │ +│ 5. Execution: ingest, storage, query │ └──────────────────────────────────────────────────────────────────┘ ``` From fcd5b459ea6f841002d830e527260540dea9a715 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 09:25:01 -0400 Subject: [PATCH 21/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 119 +++++++++++++----- 1 file changed, 91 insertions(+), 28 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 391d2960..ee0fd338 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -13,36 +13,99 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu ## Layers ```text - Query workload (PromQL/SQL/MetricsQL, query repeating pattern, accuracy requirements, query latency requirement) - + Data workload (data arrival pattern: streaming data vs data at rest, data distribution, cardinality) - + Deployment inputs: empirical cost estimation, empirical accuracy estimation, capabilities - │ -┌───────────────────────┴───────── ASAPPlanner ────────────────────┐ -│ 0. Frontends │ -│ Parse + lower - -> Output: LogicalDAG, CandidateLogicalDAGs │ -│ Reject unsupported query expressions. │ -│ │ │ -│ Logical planning (what to compute) │ -│ 1. Logical ASAP-aware optimization │ -│ ASAPPrimitives: Summary families x exact candidates x query rewrites x supporting multiple computation nodes with one summary node │ -│ -> Output: LogicalDAGswithASAPPrimitive, CandidateLogicalDAGswithASAPPrimitive │ -│ │ │ -│ Physical planning (how to compute) │ │ -│ 2. Physical ASAP-aware optimization | - Materialization or not for a subDAG x Physical operator implementation selection x Parallelism & Partitioning x resource management | - -> Output: PhysicalDAGwithASAPPrimitives, CandidatePhysicalDAGswithASAPPrimitives │ -│ │ │ -│ 3. Plan Selection │ -│ Cost every candidate with the deployment's cost model and accuracy model; │ -│ choose the cheapest admissible one for the workload. │ -└──────────────────────┬───────────────────────────────────────────┘ - one optimal PhysicalDAGwithASAPPrimitives (the boundary) -┌──────────────────────┴──────── Deployment ───────────────────────┐ -│ 5. Execution: ingest, storage, query │ -└──────────────────────────────────────────────────────────────────┘ + Query workload + (PromQL / SQL / MetricsQL, + query recurrence, + accuracy requirements, + latency requirements) + + + Data workload + (streaming vs. data at rest, + data distribution, + cardinality) + + + Deployment inputs + (empirical cost model, + empirical accuracy model, + deployment capabilities) + │ + ▼ +┌────────────────────────────── ASAPPlanner ──────────────────────────────┐ +│ │ +│ 0. Frontends │ +│ Parse and lower source-language queries into a common logical │ +│ representation. Reject unsupported query expressions. │ +│ │ +│ Output: CandidateLogicalQueryDAGs │ +│ │ +│ │ │ +│ ▼ │ +│ Logical planning — what to compute │ +│ │ +│ 1. Logical ASAP-aware optimization │ +│ Explore semantically valid logical candidates: │ +│ │ +│ summary families │ +│ × exact candidates │ +│ × query rewrites │ +│ × sharing one summary across multiple computations │ +│ │ +│ Output: CandidateLogicalASAPDAGs │ +│ │ +│ │ │ +│ ▼ │ +│ Physical planning — how to compute │ +│ │ +│ 2. Physical ASAP-aware optimization │ +│ Explore executable implementations of each logical candidate: │ +│ │ +│ materialization decisions │ +│ × physical operator implementations │ +│ × parallelism and partitioning │ +│ × resource management │ +│ │ +│ Output: CandidatePhysicalASAPDAGs │ +│ │ +│ │ │ +│ ▼ │ +│ 3. Plan selection │ +│ Evaluate complete physical candidates using the deployment's │ +│ empirical cost and accuracy models. Reject candidates that violate │ +│ accuracy, latency, or capability constraints. │ +│ │ +│ Choose the cheapest admissible plan for the whole workload. │ +│ │ +└────────────────────────────────┬───────────────────────────────────────┘ + │ + ▼ + one selected PhysicalASAPDAG + │ + ▼ +┌────────────────────────────── Deployment ───────────────────────────────┐ +│ │ +│ 4. Execution │ +│ Bind inputs and execute ingestion, storage, precomputation, and │ +│ query-time computation described by the selected plan. │ +│ │ +└────────────────────────────────────────────────────────────────────────┘ ``` +Each optimization stage produces a candidate plan space. Candidate sets are +internal to ASAPPlanner and may be represented explicitly or enumerated lazily. +The deployment never performs query planning or plan selection. + +Each stage adds a different class of decisions: + +- **Frontend:** represents the original query semantics. +- **Logical ASAP-aware optimization:** decides **what computation** can satisfy + those semantics, including ASAP primitives, exact alternatives, rewrites, + and sharing. +- **Physical ASAP-aware optimization:** decides **how that computation runs**, + including materialization, physical operators, partitioning, parallelism, + and resource allocation. +- **Plan selection:** compares complete physical candidates and returns one + selected plan. + Layers 0 to 3 each output a candidate set, `Candidates`, holding every legal candidate of their stage (for example KLL and DDSketch for one quantile, or each lifecycle assignment), and selection is the only step that chooses. Candidate sets are internal to ASAPPlanner: they may From dee00dc22afbc4e6fa172182afa1f0dffb0c8edb Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 09:26:32 -0400 Subject: [PATCH 22/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index ee0fd338..c2e5a27c 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -6,7 +6,7 @@ as ASAPQuery-backend. ## Goal ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the [deployment's -inputs](TODO: A data structure should be explicitly defined in PR #510), and returns one optimal physical plan. It decides what is computed, how +inputs](TODO: A data structure should be explicitly defined in another PR), and returns one optimal physical plan. It decides what is computed, how it is computed, and which plan is best. The deployment only supplies inputs and executes the plan: it supplies its own empirical cost estimation, empirical accuracy estimation and capabilities of deployment but never does the query planning or plan selection. From e0282f3a995b2c8419f6b0d26dc8a367f299d67d Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 10:20:22 -0400 Subject: [PATCH 23/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 192 +++++------------- 1 file changed, 49 insertions(+), 143 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index c2e5a27c..0e8d376a 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -32,21 +32,20 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu ▼ ┌────────────────────────────── ASAPPlanner ──────────────────────────────┐ │ │ -│ 0. Frontends │ -│ Parse and lower source-language queries into a common logical │ +│ 0. Language-specific frontends │ +│ Parse and convert source-language queries into a common logical │ │ representation. Reject unsupported query expressions. │ │ │ -│ Output: CandidateLogicalQueryDAGs │ +│ Output: CandidateLogicalDAGs │ │ │ │ │ │ │ ▼ │ │ Logical planning — what to compute │ │ │ │ 1. Logical ASAP-aware optimization │ -│ Explore semantically valid logical candidates: │ +│ Explore semantically equivalent and legal logical candidates: │ │ │ │ summary families │ -│ × exact candidates │ │ × query rewrites │ │ × sharing one summary across multiple computations │ │ │ @@ -73,7 +72,7 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu │ empirical cost and accuracy models. Reject candidates that violate │ │ accuracy, latency, or capability constraints. │ │ │ -│ Choose the cheapest admissible plan for the whole workload. │ +│ Choose the cheapest valid plan for the whole workload. │ │ │ └────────────────────────────────┬───────────────────────────────────────┘ │ @@ -84,144 +83,51 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu ┌────────────────────────────── Deployment ───────────────────────────────┐ │ │ │ 4. Execution │ -│ Bind inputs and execute ingestion, storage, precomputation, and │ -│ query-time computation described by the selected plan. │ +│ deployment executes the selected DAG (plan). │ │ │ └────────────────────────────────────────────────────────────────────────┘ ``` -Each optimization stage produces a candidate plan space. Candidate sets are -internal to ASAPPlanner and may be represented explicitly or enumerated lazily. -The deployment never performs query planning or plan selection. - -Each stage adds a different class of decisions: - -- **Frontend:** represents the original query semantics. -- **Logical ASAP-aware optimization:** decides **what computation** can satisfy - those semantics, including ASAP primitives, exact alternatives, rewrites, - and sharing. -- **Physical ASAP-aware optimization:** decides **how that computation runs**, - including materialization, physical operators, partitioning, parallelism, - and resource allocation. -- **Plan selection:** compares complete physical candidates and returns one - selected plan. - -Layers 0 to 3 each output a candidate set, `Candidates`, holding every -legal candidate of their stage (for example KLL and DDSketch for one quantile, -or each lifecycle assignment), and selection is the only step that chooses. Candidate sets are internal to ASAPPlanner: they may -be shared or enumerated lazily, and the deployment never sees them. Unsupported -or infeasible candidates are rejected with reasons, not silently dropped. - -## DAGs and what each encodes - -Each stage adds decisions to the DAG it receives. The table shows which -decisions each DAG carries. - -| | `PreASAPDAG` | `LogicalPostASAPDAG` | `LifecyclePostASAPDAG` | `PhysicalPostASAPDAG` | -|---|---|---|---|---| -| Produced by | 0. Frontends | 1. Logical optimization | 2. Summary lifecycle planning | 3. Compilation; 4. selects one | -| Node | Query operation | Logical operation, including summary operations | Same, plus annotations | Physical operator | -| Logical optimization (summary family, rewrites) | No | Yes | Yes | Yes | -| Materialization decided (which summary states persist) | No | No | Yes | Yes | -| Data lifecycle (how each state is maintained: `Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) | No | No | Yes | Yes | -| Retention (how long each state lives) and window framework | No | No | Yes | Yes | -| Execution time (ingestion or query) | No | No | Yes, per node | Yes, as the precompute / query split | -| Physical optimization (operator choice, e.g. TopK as sort + limit) | No | No | No | Yes | -| Seen by the deployment | No | No | No | Only the selected one | - -Name mapping to code: - -| Design name | Current main | Target API (open PRs #508, #480) | -|---|---|---| -| `PreASAPDAG` | `Rc` | `PreASAPDAG` | -| `LogicalPostASAPDAG` | `Rc` tree; exported as `PostAsapDag` | `LogicalPostASAPDAG` | -| `LifecyclePostASAPDAG` | `SummaryMaintenanceLifecyclePlan`, one selected assignment beside the DAG | `LifecyclePostASAPDAG`; `SummaryMaintenanceLifecyclePlan` is merged into it | -| `PhysicalPostASAPDAG` | None | `PhysicalPostASAPDAG` | -| `CandidatePreASAPDAGs` | None; one `QueryExpr` root per entry | `CandidatePreASAPDAGs` | -| `CandidateLogicalPostASAPDAGs` | `PlanSpace` | `CandidateLogicalPostASAPDAGs` | -| `CandidateLifecyclePostASAPDAGs` | None | `CandidateLifecyclePostASAPDAGs` | -| `CandidatePhysicalPostASAPDAGs` | None | `CandidatePhysicalPostASAPDAGs` | - -**Post-ASAP DAG to lifecycle DAG.** A summary state's *lifecycle* -(`Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) fixes several -separate aspects together: whether the state is materialized (kept across -executions, like a materialized view), when it is computed, how it is -maintained, how long it is retained, and its window framework. The table above -lists these aspects separately; the lifecycle is the one choice that sets them. -Choosing lifecycles is a workload-level decision, like a database's -materialized-view selection: it spans queries (a shared state is kept once) and -depends on workload demand (read and update rates, horizon). A lifecycle -assignment annotates every node with execution timing, and every stored state -with window framework and retention: - -* a retained state (`ContinuouslyMaintained`, `Shared`, `Prepared`) and every - node feeding it run at ingestion time; -* readouts, other consumers and `Ephemeral` states run at query time; -* an `Ephemeral` state that feeds a retained state runs at ingestion time, - because query-time work may not feed ingestion-time work. - -The result is still a Post-ASAP DAG, much as physical properties annotate -logical expressions in a database optimizer. This is the only source of -timing; logical optimization proposes computations, never timing or placement. - -**Lifecycle DAG to physical DAG.** Compilation lowers each node to physical -operators and cuts the graph at the timing frontier (ingestion-time nodes read -by query-time nodes, plus an ingestion-time root) into a precompute and a query -DAG. Materialization is decided in layer 2 and realized here: the precompute -DAG's outputs at the cut are the materialized states, and the query DAG reads -them through typed input slots. A physical DAG corresponds to its -Post-ASAP DAG node by node; a node may expand into several operators, whose -helper operators are numbered from their source node. The one exception, a -`Fallback` node wrapping a whole Pre-ASAP expression, is removed by the -operator-flattening proposal ([operator sharing](operator-sharing.md), #469, -#481). Operator materialization (sort, aggregation, summary build) and -computing a shared subexpression once are compilation and runtime details. - -Binding runtime sources is an execution step of a `PhysicalPostASAPDAG`, not -another DAG: the deployment supplies a source for each typed input slot, the -slots are checked against their contracts, and the graph runs. - -## Responsibilities - -| Layer | Owns | Does not own | -|---|---|---| -| 0. Frontends | Language semantics and lowering. A construct that cannot be represented faithfully is rejected, never ignored (for example PromQL `fill`). | Summaries, placement | -| 1. Logical optimization | All legal logical candidates: summary families, exact rewrites, compositions, series-identity typing. | Placement, timing | -| 2. Summary lifecycle planning | For each unique summary state and maintained population, the admissible lifecycle assignments and their timing, window framework and retention. | Cost values; operator implementation | -| 3. Compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | -| 4. Selection | Costing every candidate with the deployment's cost model and returning the cheapest admissible `PhysicalPostASAPDAG` for the whole workload that meets the accuracy requirements. A state shared by several queries is costed once with all consumers' demand (only when compilation installs one shared output: same window layout, evaluation interval and phase). Unknown cost stays unknown and such a candidate is not selected. | The cost values | -| 5. Deployment | Inputs: the cost model (build, per-update maintenance, read, store price per byte-second, retirement, query-time raw processing; optionally whole-plan quotes), accuracy requirements and capabilities (for example whether query-time raw data is available). Execution: ingestion and routing, panes and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | - -## The boundary - -The deployment passes its inputs to ASAPPlanner and receives one optimal -`PhysicalPostASAPDAG`, which contains: - -* a **precompute DAG**, whose inputs are raw-sample contracts (rows carrying - series labels, timestamp and value; the label set is the complete series - identity) and whose outputs are typed summary states; -* a **query DAG**, whose inputs are stored-state contracts, query-time - raw-series contracts, or both; -* the lifecycle, window framework and retention of every stored output. - -The deployment binds each input contract, stores each precompute output under -its own storage identity, and returns the query DAG's result. Semantic identity -of stored outputs is defined by the logical DAG they compute. Storage identity -and encoding belong to the deployment. - -## Example - -This traces `sum by (job) (rate(m[1m]))`, evaluated every 10 s, through the -four DAGs, and shows that only the deployment's store price changes the plan. - -* `PreASAPDAG`: `sum by (job)` over `rate` over the range selector `m[1m]`. -* `LogicalPostASAPDAG`: a per-series Rate state feeding a grouped Sum state. -* `LifecyclePostASAPDAG`: one per lifecycle assignment, for example - (a) both retained, (b) Rate retained and Sum `Ephemeral`, (c) both - `Ephemeral`. -* `PhysicalPostASAPDAG`: one compilation, cut three ways. (a) Precompute builds - Rate and Sum per pane; the query only reads Sum. (b) Precompute keeps Rate; - the query builds Sum. (c) No precompute; the query reads raw series at `t_q`. - -Selection returns (a) when storage is cheap, (b) when it is expensive, and (c) -when it is more expensive still. The deployment only changed its store price. + +## Layers/DAGs and what decisions each layer makes + + +Stages 0 to 2 each output a candidate set holding every semantically equivalent and legal candidate DAG of that stage; stage 3 is the only step that chooses one candidate DAG as output. Candidate sets are internal to ASAPPlanner and may be shared or enumerated lazily. A stage may prune a candidate early only when it is provably inadmissible (for example, a summary family that cannot meet the query's accuracy target), and every rejected candidate carries a reason. + +| Stage | Input | Decides | Output | +|---|---|---|---| +| 0. Frontends | Each `QueryWorkloadEntry.query` and `QueryWorkload.language` | Language semantics and converting source-language queries into a common logical representation. Rejects constructs it cannot represent faithfully. | `CandidateLogicalDAGs`; nodes are logical operations with no summary operations | +| 1. Logical ASAP-aware optimization | Logical query DAGs, each query's accuracy requirements across the workload | **Pass 1, per subDAG:** replace subDAGs with summary families, apply query rewriting rules and generate candidates. **Pass 2, across subDAGs:** detect Common Subexpression with ASAP awareness, meaning CSE with exactly the same subexpression, one summary type can support multiple computation nodes and replace the per-node summary with one shared summary (e.g., UnivMon supporting distinct counting, entropy, L2 norm node), window-overlap patterns (sliding windows, sub-interval windows) and replace the per-node window with one shared window primitive (sliding window, tumbling window, Exponential Histogram). | `CandidateLogicalASAPDAGs`; nodes include summary and window-primitive operations | +| 2. Physical ASAP-aware optimization | Logical ASAP DAGs, each entry's `recurrence` and `predictability`, the `DataWorkload` | **Materialization:** for subDAGs, whether it is materialized and therefore computed at ingestion time or query time, and how long it is retained. **Lowering:** physical operators for every node, and the cut into a precompute DAG and a query DAG. **Parallelism, partitioning, resources:** TODO. | `CandidatePhysicalASAPDAGs` | +| 3. Plan selection | Physical ASAP DAG candidates, `requirements`, the deployment's cost model, accuracy model and capabilities | Rejects candidates that miss an accuracy target, a latency bound or a capability; picks the cheapest admissible plan for the whole workload. A shared state is costed once with all its consumers' demand. Unknown cost is never selected. | One `PhysicalASAPDAG` | +| 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Binds raw samples and stored states to the plan's inputs, runs ingestion, storage, precomputation and query-time computation. | Query results | + +## End-to-End Example for Layers and Decision Space for Logical and Physical ASAP-aware Optimization +Every example below is a PlanningWorkload in the serialized form of [workload.rs](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). Times and durations are milliseconds. + +TODO: write concrete query expressions, workload pattern based on the workload data structure, data workload based on the data structure definition per subsection. +### Aggregation over dimensions and motivation for summary replacement decisions in logical planning + +sum by (job) (rate(data[1m])) as an example + +logical asap-aware planning can represent sum with exact aggregation SUM. + +topk by (job) (sum_over_time(data[1m])), logical planning asap-aware can represent this as Count-Min Sketch with heap per job row, or Hydra group-by for the entire job column. + + + +### Aggregation over windows and motivation for summary replacement and sharing one summary for multiple computation nodes in logical planning + +query can be quantile aggregation of past 5 year data, past 1 year data, past year between 1 - 2 years, 2 to 3, 2 to 5 etc. And during logical asap-aware planning stage, these aggregation over time windows can be mapped to KLL for quantile over these time windows. query can also be, querying data over every 5min windows, and repeating the query every 1min, as the repeatedness of the query workload. + +In the logical asap-aware planning stage, a second pass will detect that the window overlapping patten, e.g., it appears as sliding windows, or sub-interval window queries, and then the logical asap-aware planning stage, can map these batch queries' window overlapping pattern to a shared ASAP window primitive node, e.g., sliding window, Exponential Histogram for sub-interval window queries. + +### Aggregation over windows and motivation for materialization decisions in physical planning + +query can be quantile aggregation of past 5 year data, past 1 year data, past year between 1 - 2 years, 2 to 3, 2 to 5 etc. And during logical asap-aware planning stage, these aggregation over time windows can be mapped to KLL for quantile over these time windows. query can also be, querying data over every 5min windows, and repeating the query every 1min, as the repeatedness of the query workload. + +In the logical asap-aware planning stage, a second pass will detect that the window overlapping patten, e.g., it appears as sliding windows, or sub-interval window queries, and then the logical asap-aware planning stage, can map these batch queries' window overlapping pattern to a shared ASAP window primitive node, e.g., sliding window, Exponential Histogram for sub-interval window queries. + +Note that the ASAP primitives with windows (e.g., sliding window, tumbling window, Exponential Histogram window frameworks) replacement to a logical query subDAG happens in the logical asap-aware planning stage, the decision of whether materializing the window summary nodes or not happens in the physical planning stage and thus decoupled from the window primitive candidate enumeration. + + From 22f927932e7872f2983736e26f4666e51a5b1bc0 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 10:22:07 -0400 Subject: [PATCH 24/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 0e8d376a..bad615d1 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -25,15 +25,15 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu cardinality) + Deployment inputs - (empirical cost model, - empirical accuracy model, + (cost model, + accuracy model, deployment capabilities) │ ▼ ┌────────────────────────────── ASAPPlanner ──────────────────────────────┐ │ │ -│ 0. Language-specific frontends │ -│ Parse and convert source-language queries into a common logical │ +│ 0. Language-specific frontends │ +│ Parse and convert source-language queries into a common logical │ │ representation. Reject unsupported query expressions. │ │ │ │ Output: CandidateLogicalDAGs │ From 127310cddf8f51136805e126cf64e7b9d0e52e90 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 11:31:04 -0400 Subject: [PATCH 25/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 545 +++++++++++++++++- 1 file changed, 523 insertions(+), 22 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index bad615d1..0f20e551 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -72,7 +72,7 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu │ empirical cost and accuracy models. Reject candidates that violate │ │ accuracy, latency, or capability constraints. │ │ │ -│ Choose the cheapest valid plan for the whole workload. │ +│ Choose the cheapest valid plan for the whole workload. │ │ │ └────────────────────────────────┬───────────────────────────────────────┘ │ @@ -88,46 +88,547 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu └────────────────────────────────────────────────────────────────────────┘ ``` +The planner receives three groups of inputs: -## Layers/DAGs and what decisions each layer makes +| Input | Contents | +|---|---| +| Query workload | Source-language expressions, recurrence, predictability, time selection, accuracy requirements, and latency requirements | +| Data workload | Data arrival pattern, sampling cadence, ingestion volume and rate, cardinality, and distribution | +| Deployment inputs | Empirical cost model, empirical accuracy model, and execution capabilities | + +Logical planning decides **what** to compute: summary families, query rewrites, +and sharing across computations and windows. Physical planning decides **how** +to compute it: materialization, execution placement, physical operators, +partitioning, parallelism, and resource allocation. Selection evaluates +complete physical candidates and chooses one plan for the whole workload. +## Layers/DAGs and what decisions each layer makes -Stages 0 to 2 each output a candidate set holding every semantically equivalent and legal candidate DAG of that stage; stage 3 is the only step that chooses one candidate DAG as output. Candidate sets are internal to ASAPPlanner and may be shared or enumerated lazily. A stage may prune a candidate early only when it is provably inadmissible (for example, a summary family that cannot meet the query's accuracy target), and every rejected candidate carries a reason. +Stages 0 to 2 each output a candidate set holding every semantically equivalent +and legal candidate DAG of that stage; stage 3 is the only step that chooses +one candidate DAG as output. Candidate sets are internal to ASAPPlanner and may +be shared or enumerated lazily. A stage may prune a candidate early only when +it is provably invalid (for example, a summary family that cannot meet the +query's accuracy target), and every rejected candidate carries a reason. | Stage | Input | Decides | Output | |---|---|---|---| -| 0. Frontends | Each `QueryWorkloadEntry.query` and `QueryWorkload.language` | Language semantics and converting source-language queries into a common logical representation. Rejects constructs it cannot represent faithfully. | `CandidateLogicalDAGs`; nodes are logical operations with no summary operations | -| 1. Logical ASAP-aware optimization | Logical query DAGs, each query's accuracy requirements across the workload | **Pass 1, per subDAG:** replace subDAGs with summary families, apply query rewriting rules and generate candidates. **Pass 2, across subDAGs:** detect Common Subexpression with ASAP awareness, meaning CSE with exactly the same subexpression, one summary type can support multiple computation nodes and replace the per-node summary with one shared summary (e.g., UnivMon supporting distinct counting, entropy, L2 norm node), window-overlap patterns (sliding windows, sub-interval windows) and replace the per-node window with one shared window primitive (sliding window, tumbling window, Exponential Histogram). | `CandidateLogicalASAPDAGs`; nodes include summary and window-primitive operations | -| 2. Physical ASAP-aware optimization | Logical ASAP DAGs, each entry's `recurrence` and `predictability`, the `DataWorkload` | **Materialization:** for subDAGs, whether it is materialized and therefore computed at ingestion time or query time, and how long it is retained. **Lowering:** physical operators for every node, and the cut into a precompute DAG and a query DAG. **Parallelism, partitioning, resources:** TODO. | `CandidatePhysicalASAPDAGs` | -| 3. Plan selection | Physical ASAP DAG candidates, `requirements`, the deployment's cost model, accuracy model and capabilities | Rejects candidates that miss an accuracy target, a latency bound or a capability; picks the cheapest admissible plan for the whole workload. A shared state is costed once with all its consumers' demand. Unknown cost is never selected. | One `PhysicalASAPDAG` | -| 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Binds raw samples and stored states to the plan's inputs, runs ingestion, storage, precomputation and query-time computation. | Query results | +| 0. Frontends | Each `QueryWorkloadEntry.query` and `QueryWorkload.language` | Language semantics, and converting source-language queries into a common logical representation. Rejects constructs it cannot represent faithfully. | `CandidateLogicalDAGs`; nodes are logical operations with no summary operations | +| 1. Logical ASAP-aware optimization | Logical DAGs, and each query's accuracy requirements across the workload | **Pass 1 — Summary replacement and query rewriting:** apply query rewriting rules to each eligible sub-DAG and generate exact and summary-based candidates that satisfy its semantics and accuracy requirements. **Pass 2 — ASAP-aware CSE:** apply traditional CSE and summary-specific CSE rules across sub-DAGs and queries to generate shared computation candidates, while preserving independent candidates. | `CandidateLogicalASAPDAGs`; nodes include summary and window-primitive operations | +| 2. Physical ASAP-aware optimization | Logical ASAP DAGs, each entry's `recurrence` and `predictability`, the `DataWorkload` | **Materialization:** for each sub-DAG, whether its output is materialized and therefore computed at ingestion time or query time, and how long it is retained. **Physical operator implementation:** physical operators for every node. **Parallelism, partitioning, resources:** TODO. | `CandidatePhysicalASAPDAGs` | +| 3. Plan selection | Physical ASAP DAG candidates, `requirements`, the deployment's cost model, accuracy model and capabilities | Rejects candidates that miss an accuracy target, a latency bound or a capability; picks the cheapest valid plan for the whole workload. A shared state is costed once with all its consumers' demand. | One `PhysicalASAPDAG` | +| 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Executes ingestion, storage, precomputation and query-time computation. | Query results | + +### 0. Language-specific frontends + +The frontend converts each query into a `LogicalDAG`. Nodes represent logical +query operations, including selectors, transformations, aggregations, grouping +and window semantics. They contain no ASAP summary choices. + +The frontend preserves source-language behavior, including series identity, +evaluation timing and missing-data semantics. A construct that cannot be +represented faithfully is rejected. + +### 1. Logical ASAP-aware optimization + +Logical optimization runs in two passes. Pass 1 generates candidates for each +computation on its own; Pass 2 finds candidates that share computation across +sub-DAGs and queries. Neither pass decides materialization or execution +placement, and neither picks one summary per computation: every candidate is +kept for selection. + +#### Pass 1: Local candidate generation + +For each eligible sub-DAG, Pass 1 identifies its computation semantics, applies +rewrite rules, and generates every candidate that can meet its accuracy +requirement. + +| Original computation | Local candidates | +|---|---| +| `Sum(x) by (g)` | Exact grouped sum | +| `TopK(k, x) by (g)` | Exact sort and limit per group, Count-Min Sketch with a top-*k* heap per group, Hydra over all groups | +| `Distinct(x)` | Exact distinct, a specialized distinct summary, UnivMon | +| `Entropy(x)` | Exact entropy, a specialized entropy summary, UnivMon | +| `L2(x)` | Exact L2 norm, a specialized norm summary, UnivMon | +| `Quantile(x, window)` | Exact quantile, KLL over the requested window | + +Each candidate records its input expression, filter, grouping, window, +supported readouts and accuracy requirement. Pass 2 uses these to decide +whether candidates can share a producer. + +A summary-based candidate has two kinds of nodes. The **summary node** builds +and maintains the summary from the input data, for example a KLL sketch over +`latency_ms`. A **readout node** computes a query's answer from that summary, +for example the p99 estimate from the KLL, or the entropy estimate from a +UnivMon. One summary node can feed several readout nodes, which is what Pass 2 +exploits. + +#### Pass 2: ASAP-aware common-subexpression elimination + +ASAP-aware CSE extends traditional CSE with summary-specific sharing rules. +Computations can share a producer when they use identical expressions, when +one summary supports multiple readouts, or when summary composition supports +their overlapping windows. + +The rules compare computations by their **summary input data**: what a summary for +that computation would ingest, namely the data source, the filters, and the key +or value being summarized together with its grouping. The summary input data does +not include the window; the window-composition rule compares windows +separately. + +| ASAP-aware CSE rule | Sharing condition | Shared computation | +|---|---|---| +| Identical-expression rule | The input and computation semantics are identical. | One common computation node serving multiple consumers. | +| Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary node connected to multiple readout nodes; for example, one UnivMon over `src_ip` from `flows` in the last minute serves three queries refreshed every 10 s: `COUNT(DISTINCT src_ip)`, the entropy of the `src_ip` distribution, and the L2 norm of per-`src_ip` counts. Each flow record updates the one UnivMon node once; a distinct-count readout node, an entropy readout node and an L2 readout node each read their statistic from it. The UnivMon is sized for the strictest of the three accuracy requirements (Example 2). | +| Window-composition rule | The computations have the same summary input data, and the summary's merge and window capabilities can reconstruct the requested windows within their accuracy requirements. | Shared window-summary nodes connected to window-specific reconstruction and readout nodes; for example, (1) one sliding window of 1-min KLL panes serves every evaluation of `quantile_over_time(0.99, latency_ms[5m])` repeated every minute: each evaluation's readout node merges the latest 5 panes, so consecutive evaluations share 4 of their 5 panes (Example 3, Pattern B); (2) one Exponential Histogram of KLL buckets over the last 5 years serves the p99 queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and `[3y] offset 2y`: each query's reconstruction node merges the buckets covering its interval, and its readout node reads p99 from the merged KLL (Example 3, Pattern A). The same panes or buckets also serve other quantiles, since one KLL answers every quantile: adding `quantile_over_time(0.5, latency_ms[5m])` to the dashboard in (1) adds only a p50 readout node next to the p99 one, both reading the same 5 merged panes, with no new summary. | + +Rules are defined by each summary family's capabilities and semantic +requirements. A shared summary must meet the strictest accuracy requirement +among its consumers. Applying a rule adds a shared candidate and keeps the +independent candidates, so selection can compare both. + +### 2. Physical ASAP-aware optimization + +Physical optimization turns each logical candidate into executable candidates. +It makes two ASAP-specific decisions, described below. Parallelism, partitioning +and resource management are TODO. + +#### Materialization + +Materialization decides, for each sub-DAG, whether its output is kept across +(batch) query executions, and if so, when it is computed and how long it is +stored. Materialization does not imply ingestion time; a sub-DAG has three +options: -## End-to-End Example for Layers and Decision Space for Logical and Physical ASAP-aware Optimization -Every example below is a PlanningWorkload in the serialized form of [workload.rs](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). Times and durations are milliseconds. +* **Materialized at ingestion time:** the sub-DAG runs as data arrives, and its + output is stored before any query asks for it. For example, the 1-min KLL + panes in Example 4, Pattern B. +* **Materialized at query time:** the sub-DAG runs when a query first needs + it, and its output is stored so that later executions, or other queries in + the same batch, reuse it instead of recomputing it. For example, an + Exponential Histogram built when the batch in Example 4, Pattern A runs and + read by all five of its queries. +* **Not materialized:** the sub-DAG runs at query time for each execution, and + its output is discarded afterward. -TODO: write concrete query expressions, workload pattern based on the workload data structure, data workload based on the data structure definition per subsection. -### Aggregation over dimensions and motivation for summary replacement decisions in logical planning +Whichever option is chosen for each sub-DAG, the plan must also satisfy these +constraints: -sum by (job) (rate(data[1m])) as an example +* A materialized output is stored for as long as any of its consumers still + needs it. +* Query-time work cannot feed ingestion-time work, so every node feeding an + ingestion-time sub-DAG also runs at ingestion time. +* Readout nodes, which compute a query's answer from a summary, always run at + query time. -logical asap-aware planning can represent sum with exact aggregation SUM. +The decision depends on the workload's `recurrence` and `predictability` and on +the `DataWorkload`, none of which logical planning reads. Typical outcomes: -topk by (job) (sum_over_time(data[1m])), logical planning asap-aware can represent this as Count-Min Sketch with heap per job row, or Hydra group-by for the entire job column. +* Read by repeated queries while data keeps arriving: materialize at ingestion + time. +* Read by several queries in one batch, or over data at rest: materialize at + query time. +* Read once by an ad hoc query: do not materialize. +A shared summary is materialized once for all its consumers. See Example 4. +#### Physical operator implementation -### Aggregation over windows and motivation for summary replacement and sharing one summary for multiple computation nodes in logical planning +Physical operator implementation lowers every node to physical operators, for +example TopK as a sort followed by a limit, or a KLL node as summary build, +merge and quantile readout operators. -query can be quantile aggregation of past 5 year data, past 1 year data, past year between 1 - 2 years, 2 to 3, 2 to 5 etc. And during logical asap-aware planning stage, these aggregation over time windows can be mapped to KLL for quantile over these time windows. query can also be, querying data over every 5min windows, and repeating the query every 1min, as the repeatedness of the query workload. +### 3. Plan selection + +Selection rejects every physical candidate that the deployment's accuracy model +estimates will miss an accuracy target, that misses a latency bound, or that +needs a capability the deployment lacks. Among the remaining candidates, it +uses the deployment's cost model to choose the cheapest plan for the whole +workload. A shared state is costed once, with the demand of all its consumers. + +### 4. Execution + +The deployment executes the selected `PhysicalASAPDAG`: ingestion, storage, +precomputation and query-time computation. It makes no planning decisions. + +## End-to-end examples + +Every example below is a `PlanningWorkload` in the serialized form of +[`workload.rs`](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). +Times and durations are milliseconds. + +| Example | Shows | +|---|---| +| 1. Aggregation over dimensions | Pass 1: summary replacement | +| 2. One summary for several computations | Pass 2: summary-capability rule | +| 3. Aggregation over windows | Pass 1 and Pass 2: window-composition rule | +| 4. Materialization of window summaries | Stage 2: materialization, decoupled from stage 1 | + +In the diagrams below, grey cylinders are input data, white boxes are exact +operations, blue boxes are summary nodes, and green rounded boxes are readout +nodes. + +### Shared data workload + +Unless an example says otherwise, the queried metric is scraped every 15 s +from about one million series with Zipf-distributed keys, and is still +arriving: + +```json +"data_workload": { + "arrival": "continuously_ingesting", + "data_ingestion_interval": {"value": 15000, "source": "declared", "observed_at_ms": null, "valid_for_ms": null}, + "ingestion_volume": {"value": null, "source": "unknown", "observed_at_ms": null, "valid_for_ms": null}, + "ingestion_rate": {"value": 66667.0, "source": "declared", "observed_at_ms": null, "valid_for_ms": null}, + "input_cardinality": {"value": 1000000, "source": "declared", "observed_at_ms": null, "valid_for_ms": null}, + "distribution": {"value": "zipf", "source": "declared", "observed_at_ms": null, "valid_for_ms": null} +} +``` + +### Example 1: Aggregation over dimensions — summary replacement in Pass 1 + +**Query workload.** Two dashboard panels refresh every 10 s over the last +minute. The first needs an exact total; the second tolerates error. + +```json +"query_workload": { + "language": "promql", + "query_batch": null, + "repeating_queries": [ + { + "query": "sum by (job) (rate(http_requests_total[1m]))", + "demand": {"fixed_interval": 10000}, + "requirements": {"accuracy": "implicit_exact", "response_latency": "unspecified"}, + "predictability": {"predictable": {"known_at": null}}, + "time_selection": {"scope": "real_time", "lookback": 60000, "as_of": null} + }, + { + "query": "topk by (job) (10, sum_over_time(http_requests_total[1m]))", + "demand": {"fixed_interval": 10000}, + "requirements": { + "accuracy": {"explicit": {"EpsilonDelta": {"epsilon": 0.01, "delta": 0.001}}}, + "response_latency": {"explicit_max_ms": 100.0} + }, + "predictability": {"predictable": {"known_at": null}}, + "time_selection": {"scope": "real_time", "lookback": 60000, "as_of": null} + } + ] +} +``` + +**Stage 0.** The two `LogicalDAG`s are +`range m[1m] → rate → sum by (job)` and +`range m[1m] → sum_over_time → topk by (job) (10)`. + +**Stage 1, Pass 1.** + +* `sum by (job) (rate(...))`: the accuracy requirement is `implicit_exact`, and + an exact grouped sum already keeps one value per `job`. The only candidate is + a per-series Rate feeding an exact per-`job` Sum. No summary helps here. +* `topk by (job) (10, ...)`: the exact candidate keeps a per-series sum and + sorts within each `job`, which is costly at one million Zipf-distributed + series. The `EpsilonDelta` target admits two summary candidates: + 1. **Count-Min Sketch with a top-*k* heap per `job`.** One sketch per group; + each answers its own top 10. + 2. **Hydra over the whole `job` column.** One sketch covers every (`job`, + series) key and answers the top 10 for any `job`. + + Pass 1 keeps all three candidates. The better summary depends on the number + of jobs and on costs that only the deployment knows. + +```mermaid +flowchart LR + subgraph L["Stage 0 · LogicalDAG"] + direction LR + a1[("http_requests_total
last 1m")]:::data --> a2["sum_over_time"]:::exact --> a3["topk by (job) (10)"]:::exact + end + subgraph C["Stage 1, Pass 1 · CandidateLogicalASAPDAGs"] + direction TB + subgraph E["Exact"] + direction LR + e1[("input")]:::data --> e2["sum_over_time
per series"]:::exact --> e3["sort + limit 10
per job"]:::exact + end + subgraph CM["Count-Min + heap per job"] + direction LR + c1[("input")]:::data --> c2["Count-Min Sketch +
top-10 heap, one per job"]:::summary --> c3(["top 10
per job"]):::readout + end + subgraph H["Hydra"] + direction LR + h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::readout + end + end + S{{"Stage 3 · cheapest valid candidate
many small jobs → Hydra
few large jobs → Count-Min per job"}} + L --> C --> S + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef readout fill:#e6f4ea,stroke:#188038,color:#000; +``` -In the logical asap-aware planning stage, a second pass will detect that the window overlapping patten, e.g., it appears as sliding windows, or sub-interval window queries, and then the logical asap-aware planning stage, can map these batch queries' window overlapping pattern to a shared ASAP window primitive node, e.g., sliding window, Exponential Histogram for sub-interval window queries. +**Stage 3.** The deployment's accuracy model checks that each summary candidate +meets ε = 0.01, δ = 0.001, and its cost model compares per-`job` sketches with +one Hydra sketch. With many small jobs, one shared Hydra sketch is typically +cheaper; with a few large jobs, per-`job` Count-Min sketches may win. -### Aggregation over windows and motivation for materialization decisions in physical planning +### Example 2: One summary for several computations — the summary-capability rule in Pass 2 -query can be quantile aggregation of past 5 year data, past 1 year data, past year between 1 - 2 years, 2 to 3, 2 to 5 etc. And during logical asap-aware planning stage, these aggregation over time windows can be mapped to KLL for quantile over these time windows. query can also be, querying data over every 5min windows, and repeating the query every 1min, as the repeatedness of the query workload. +**Query workload.** A network-monitoring dashboard computes three statistics of +source IPs over the last minute, every 10 s. Each SQL query below is one +`RepeatingEntry` with `"demand": {"fixed_interval": 10000}` and +`"time_selection": {"scope": "real_time", "lookback": 60000, "as_of": null}`, +in a workload with `"language": {"sql": "datafusion_sql"}`. -In the logical asap-aware planning stage, a second pass will detect that the window overlapping patten, e.g., it appears as sliding windows, or sub-interval window queries, and then the logical asap-aware planning stage, can map these batch queries' window overlapping pattern to a shared ASAP window primitive node, e.g., sliding window, Exponential Histogram for sub-interval window queries. +| Query | Computation | Accuracy | +|---|---|---| +| `SELECT COUNT(DISTINCT src_ip) FROM flows WHERE ts >= now() - INTERVAL '1 minute'` | `Distinct(src_ip)` | ε = 0.02, δ = 0.01 | +| `SELECT -SUM(p * LN(p)) FROM (SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `Entropy(src_ip)` | ε = 0.05, δ = 0.01 | +| `SELECT SQRT(SUM(c * c)) FROM (SELECT src_ip, COUNT(*) AS c FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `L2(src_ip)` | ε = 0.01, δ = 0.01 | + +The data workload is as above, except that `input_cardinality` is the number of +distinct source IPs (say 10 million) and `data_ingestion_interval` is not +needed for SQL. + +**Pass 1.** Rewrite rules recognize the three computations, and each gets its +local candidates from the Pass 1 table: exact, a specialized summary, or +UnivMon. + +**Pass 2.** All three computations have the same summary input data (`src_ip` +from `flows`, no other filter) and the same 1-min window. UnivMon supports all three +readouts, so the summary-capability rule adds a shared candidate: **one +UnivMon node feeding three readout nodes**. It must be sized for the strictest +requirement, ε = 0.01. The independent candidates are kept as well. + +```mermaid +flowchart LR + subgraph IND["Independent candidates (Pass 1, summary option shown)"] + direction LR + i1[("flows.src_ip
last 1m")]:::data + i1 --> d1["distinct
summary"]:::summary --> rd1(["distinct count"]):::readout + i1 --> en1["entropy
summary"]:::summary --> re1(["entropy"]):::readout + i1 --> l1["norm
summary"]:::summary --> rl1(["L2 norm"]):::readout + end + subgraph SH["Shared candidate (Pass 2, summary-capability rule)"] + direction LR + i2[("flows.src_ip
last 1m")]:::data --> u["UnivMon
sized for ε = 0.01"]:::summary + u --> rd2(["distinct count"]):::readout + u --> re2(["entropy"]):::readout + u --> rl2(["L2 norm"]):::readout + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef readout fill:#e6f4ea,stroke:#188038,color:#000; +``` + +Exact candidates for each computation are also kept but omitted from the +diagram. + +**Stage 3.** Selection compares one UnivMon sized for ε = 0.01 against three +separate summaries, each sized for its own requirement. The shared candidate +usually wins because each flow record updates one summary instead of three. + +### Example 3: Aggregation over windows — the window-composition rule in Pass 2 + +This example has two workload patterns that both lead to a shared window +summary. + +**Pattern A: a batch of sub-interval queries over historical data.** An analyst +submits a batch of p99 latency reports over different historical intervals, +executed together at T = `1790000000000`: + +| Query | `lookback` | `as_of` | +|---|---|---| +| `quantile_over_time(0.99, latency_ms[5y])` | 5 y | T | +| `quantile_over_time(0.99, latency_ms[1y])` | 1 y | T | +| `quantile_over_time(0.99, latency_ms[1y] offset 1y)` | 1 y | T − 1 y | +| `quantile_over_time(0.99, latency_ms[1y] offset 2y)` | 1 y | T − 2 y | +| `quantile_over_time(0.99, latency_ms[3y] offset 2y)` | 3 y | T − 2 y | + +Each row is a `BatchEntry` such as: + +```json +{ + "query": "quantile_over_time(0.99, latency_ms[1y] offset 1y)", + "requirements": {"accuracy": {"explicit": {"EpsilonDelta": {"epsilon": 0.005, "delta": 0.01}}}, "response_latency": "unspecified"}, + "predictability": "ad_hoc", + "invocations": 1, + "execute_at": 1790000000000, + "time_selection": {"scope": "longitudinal", "lookback": 31536000000, "as_of": 1758464000000} +} +``` + +The data workload has `"arrival": "mixed"`: five years at rest plus data still +arriving. + +* **Pass 1.** Each `quantile_over_time` gets an exact candidate and a KLL over + its own interval: five independent KLL candidates over overlapping data. +* **Pass 2.** Every interval is a sub-interval of [T − 5 y, T], and KLL is + mergeable. The window-composition rule adds a shared candidate: **one + Exponential Histogram of KLL buckets over [T − 5 y, T]**, with one + reconstruction and readout node per query that merges the buckets covering + its interval. The five independent candidates are kept. + +The five query intervals overlap, and all lie inside the last five years: + +```mermaid +gantt + title Pattern A · query intervals (T = batch execution time) + dateFormat YYYY + axisFormat %Y + section Queries + q1 · [5y] :q1, 2021, 2026 + q2 · [1y] :q2, 2025, 2026 + q3 · [1y] offset 1y :q3, 2024, 2025 + q4 · [1y] offset 2y :q4, 2023, 2024 + q5 · [3y] offset 2y :q5, 2021, 2024 +``` + +The shared candidate replaces five KLL sketches with one Exponential Histogram +and a reconstruction and readout node per query: + +```mermaid +flowchart LR + in[("latency_ms
T − 5y to T")]:::data --> eh["Exponential Histogram
of KLL buckets"]:::summary + eh --> m1["merge buckets
T−5y … T"]:::exact --> o1(["q1 p99"]):::readout + eh --> m2["merge buckets
T−1y … T"]:::exact --> o2(["q2 p99"]):::readout + eh --> m3["merge buckets
T−2y … T−1y"]:::exact --> o3(["q3 p99"]):::readout + eh --> m4["merge buckets
T−3y … T−2y"]:::exact --> o4(["q4 p99"]):::readout + eh --> m5["merge buckets
T−5y … T−2y"]:::exact --> o5(["q5 p99"]):::readout + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef readout fill:#e6f4ea,stroke:#188038,color:#000; +``` + +**Pattern B: one repeating query over a sliding window.** A real-time p99 panel +over the last 5 min, refreshed every minute: + +```json +{ + "query": "quantile_over_time(0.99, latency_ms[5m])", + "demand": {"fixed_interval": 60000}, + "requirements": {"accuracy": {"explicit": {"EpsilonDelta": {"epsilon": 0.01, "delta": 0.01}}}, "response_latency": {"explicit_max_ms": 200.0}}, + "predictability": {"predictable": {"known_at": null}}, + "time_selection": {"scope": "real_time", "lookback": 300000, "as_of": null} +} +``` + +* **Pass 1.** One KLL over 5 min for each evaluation. +* **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The + window-composition rule adds a shared candidate: **a sliding window of + 1-min KLL panes**, where each evaluation merges the latest 5 panes. + +Each evaluation reads five 1-min panes, and consecutive evaluations share +four of them: + +```mermaid +gantt + title Pattern B · 1-min KLL panes and 5-min evaluations + dateFormat HH:mm + axisFormat %H:%M + section KLL panes + pane 1 :p1, 00:00, 1m + pane 2 :p2, 00:01, 1m + pane 3 :p3, 00:02, 1m + pane 4 :p4, 00:03, 1m + pane 5 :p5, 00:04, 1m + pane 6 :p6, 00:05, 1m + pane 7 :p7, 00:06, 1m + section Evaluations + eval at 00:05 (panes 1–5) :e1, 00:00, 5m + eval at 00:06 (panes 2–6) :e2, 00:01, 5m + eval at 00:07 (panes 3–7) :e3, 00:02, 5m +``` + +In both patterns, stage 1 decides only **which** window summary computes the +answer. It says nothing about when panes or buckets are built or whether they +are stored. + +### Example 4: Materialization of window summaries in physical planning + +Stage 2 takes the shared window summaries from Example 3 and decides whether to +materialize them. That choice is driven by the workload's `recurrence`, +`predictability` and `data_workload.arrival`, none of which stage 1 reads. + +**Pattern B (sliding window, repeating).** + +| Candidate | Materialized | At ingestion time | At query time | +|---|---|---|---| +| B1 | 1-min KLL panes, retained 5 min | Build one KLL pane per minute | Merge the latest 5 panes, read p99 | +| B2 | Nothing | Nothing | Read 5 min of raw samples, build one KLL, read p99 | + +```mermaid +flowchart LR + subgraph B1["B1 · panes materialized at ingestion time"] + direction LR + subgraph B1I["Ingestion time"] + s1[("samples")]:::data --> p1["1-min KLL pane
stored 5 min"]:::summary + end + subgraph B1Q["Query time, every 1 min"] + g1["merge latest
5 panes"]:::exact --> r1(["p99"]):::readout + end + p1 --> g1 + end + subgraph B2["B2 · not materialized"] + direction LR + subgraph B2Q["Query time, every 1 min"] + s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::readout + end + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef readout fill:#e6f4ea,stroke:#188038,color:#000; +``` + +The query repeats every minute and the data is continuously ingesting, so B1 +builds each pane once and reuses it in five evaluations, while B2 rescans raw +data every time. Selection usually picks B1. B2 wins only if storage is +expensive and raw data is available at query time. + +**Pattern A (sub-interval batch).** + +| Candidate | Materialized | When the Exponential Histogram is built | +|---|---|---| +| A1 | The Exponential Histogram, at query time | At query time, when the batch runs at T; read by all five queries, then discarded | +| A2 | The Exponential Histogram, at ingestion time | At ingestion time, with each new sample; old data backfilled once | + +```mermaid +flowchart LR + subgraph A1["A1 · materialized at query time"] + direction LR + subgraph A1Q["Query time, once at T"] + s3[("5 years of
stored samples")]:::data --> h3["build Exponential
Histogram once"]:::summary + h3 --> r3(["q1 … q5
readouts"]):::readout + end + end + subgraph A2["A2 · materialized at ingestion time"] + direction LR + subgraph A2I["Ingestion time, continuously"] + s4[("each new sample
+ one-time backfill")]:::data --> h4["maintain Exponential
Histogram"]:::summary + end + subgraph A2Q["Query time, at T"] + r4(["q1 … q5
readouts"]):::readout + end + h4 --> r4 + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef readout fill:#e6f4ea,stroke:#188038,color:#000; +``` -Note that the ASAP primitives with windows (e.g., sliding window, tumbling window, Exponential Histogram window frameworks) replacement to a logical query subDAG happens in the logical asap-aware planning stage, the decision of whether materializing the window summary nodes or not happens in the physical planning stage and thus decoupled from the window primitive candidate enumeration. +Which candidate wins depends on the workload: +* **As given** (`invocations: 1`, `ad_hoc`): selection picks A1. A2 would + maintain the histogram for years only to serve one batch. +* **Repeated monthly and `Predictable { known_at }`:** A2 can win, because its + maintenance cost is shared by many batches. +* **Data `"at_rest"`:** A2 is not generated, because there is no ingestion to + maintain the histogram. +**What this shows.** The same logical candidate (a sliding window of KLL panes, +or one shared Exponential Histogram) yields different physical plans depending +only on recurrence, predictability and data arrival. This is why window-summary +replacement happens in logical planning, while materialization is decided +separately in physical planning. From 843c20331c49e7a8ae53cf4a76257e0128aeb4b4 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 11:36:39 -0400 Subject: [PATCH 26/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 0f20e551..ace7a619 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -143,6 +143,7 @@ For each eligible sub-DAG, Pass 1 identifies its computation semantics, applies rewrite rules, and generates every candidate that can meet its accuracy requirement. +Example for summary candidates: | Original computation | Local candidates | |---|---| | `Sum(x) by (g)` | Exact grouped sum | From f0aa068d916a3b8f1ba1a9bbdd0b3f1b68427fab Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 11:39:34 -0400 Subject: [PATCH 27/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index ace7a619..b1be98bc 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -12,6 +12,8 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu ## Layers +x represents Cartesian product for enumerating and combining different optimization angles in planning. + ```text Query workload (PromQL / SQL / MetricsQL, From 276598d1a07a72dd5bc51e75d1406a3ebfea8c6a Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 11:39:54 -0400 Subject: [PATCH 28/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index b1be98bc..deb42929 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -12,7 +12,7 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu ## Layers -x represents Cartesian product for enumerating and combining different optimization angles in planning. +`x` represents Cartesian product for enumerating and combining different optimization angles in planning. ```text Query workload From 05e1b7248162ccaa5a670a31898d3d6ad1672098 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 11:46:34 -0400 Subject: [PATCH 29/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 59 +++++++++++++++---- 1 file changed, 48 insertions(+), 11 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index deb42929..5c49d67d 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -403,25 +403,62 @@ requirement, ε = 0.01. The independent candidates are kept as well. ```mermaid flowchart LR - subgraph IND["Independent candidates (Pass 1, summary option shown)"] - direction LR - i1[("flows.src_ip
last 1m")]:::data - i1 --> d1["distinct
summary"]:::summary --> rd1(["distinct count"]):::readout - i1 --> en1["entropy
summary"]:::summary --> re1(["entropy"]):::readout - i1 --> l1["norm
summary"]:::summary --> rl1(["L2 norm"]):::readout + in[("flows.src_ip
last 1m")]:::data + + subgraph P0["Stage 0 · LogicalDAGs"] + direction TB + q1["Distinct(src_ip)"]:::exact + q2["Entropy(src_ip)"]:::exact + q3["L2(src_ip)"]:::exact + end + + subgraph P1["Stage 1, Pass 1 · local candidates per computation"] + direction TB + subgraph D["Distinct"] + direction LR + d0["exact distinct"]:::exact + d1["distinct summary"]:::summary + d2["UnivMon"]:::summary + end + subgraph E["Entropy"] + direction LR + e0["exact entropy"]:::exact + e1["entropy summary"]:::summary + e2["UnivMon"]:::summary + end + subgraph L["L2"] + direction LR + l0["exact L2"]:::exact + l1["norm summary"]:::summary + l2["UnivMon"]:::summary + end end - subgraph SH["Shared candidate (Pass 2, summary-capability rule)"] + + subgraph P2["Stage 1, Pass 2 · summary-capability rule adds a shared candidate"] direction LR - i2[("flows.src_ip
last 1m")]:::data --> u["UnivMon
sized for ε = 0.01"]:::summary - u --> rd2(["distinct count"]):::readout - u --> re2(["entropy"]):::readout - u --> rl2(["L2 norm"]):::readout + u["one UnivMon
sized for ε = 0.01"]:::summary + u --> rd(["distinct count"]):::readout + u --> re(["entropy"]):::readout + u --> rl(["L2 norm"]):::readout end + + in --> P0 + q1 --> D + q2 --> E + q3 --> L + d2 -. "same summary input data
and window" .-> u + e2 -.-> u + l2 -.-> u + + OUT[["CandidateLogicalASAPDAGs:
all Pass 1 candidates + the shared candidate"]] + P1 --> OUT + P2 --> OUT classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef readout fill:#e6f4ea,stroke:#188038,color:#000; ``` + Exact candidates for each computation are also kept but omitted from the diagram. From 0a5df1b08b22eec9164237ea27dd7316257f1109 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 11:51:50 -0400 Subject: [PATCH 30/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 179 +++++++++--------- 1 file changed, 87 insertions(+), 92 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 5c49d67d..a97622b9 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -146,6 +146,7 @@ rewrite rules, and generates every candidate that can meet its accuracy requirement. Example for summary candidates: + | Original computation | Local candidates | |---|---| | `Sum(x) by (g)` | Exact grouped sum | @@ -256,9 +257,11 @@ precomputation and query-time computation. It makes no planning decisions. ## End-to-end examples -Every example below is a `PlanningWorkload` in the serialized form of +Each example's workload is shown as tables. Field names in code font are the +fields of [`workload.rs`](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). -Times and durations are milliseconds. +An `as_of` of "evaluation time" means `as_of: None`: the window ends when the +query runs. | Example | Shows | |---|---| @@ -273,51 +276,34 @@ nodes. ### Shared data workload -Unless an example says otherwise, the queried metric is scraped every 15 s -from about one million series with Zipf-distributed keys, and is still -arriving: - -```json -"data_workload": { - "arrival": "continuously_ingesting", - "data_ingestion_interval": {"value": 15000, "source": "declared", "observed_at_ms": null, "valid_for_ms": null}, - "ingestion_volume": {"value": null, "source": "unknown", "observed_at_ms": null, "valid_for_ms": null}, - "ingestion_rate": {"value": 66667.0, "source": "declared", "observed_at_ms": null, "valid_for_ms": null}, - "input_cardinality": {"value": 1000000, "source": "declared", "observed_at_ms": null, "valid_for_ms": null}, - "distribution": {"value": "zipf", "source": "declared", "observed_at_ms": null, "valid_for_ms": null} -} -``` +Unless an example says otherwise, every example uses this data workload: + +| `DataWorkload` field | Value | +|---|---| +| `arrival` | `continuously_ingesting` | +| `data_ingestion_interval` | 15 s (declared) | +| `ingestion_volume` | unknown | +| `ingestion_rate` | about 66,667 samples/s (declared) | +| `input_cardinality` | 1,000,000 series (declared) | +| `distribution` | `zipf` (declared) | ### Example 1: Aggregation over dimensions — summary replacement in Pass 1 -**Query workload.** Two dashboard panels refresh every 10 s over the last -minute. The first needs an exact total; the second tolerates error. - -```json -"query_workload": { - "language": "promql", - "query_batch": null, - "repeating_queries": [ - { - "query": "sum by (job) (rate(http_requests_total[1m]))", - "demand": {"fixed_interval": 10000}, - "requirements": {"accuracy": "implicit_exact", "response_latency": "unspecified"}, - "predictability": {"predictable": {"known_at": null}}, - "time_selection": {"scope": "real_time", "lookback": 60000, "as_of": null} - }, - { - "query": "topk by (job) (10, sum_over_time(http_requests_total[1m]))", - "demand": {"fixed_interval": 10000}, - "requirements": { - "accuracy": {"explicit": {"EpsilonDelta": {"epsilon": 0.01, "delta": 0.001}}}, - "response_latency": {"explicit_max_ms": 100.0} - }, - "predictability": {"predictable": {"known_at": null}}, - "time_selection": {"scope": "real_time", "lookback": 60000, "as_of": null} - } - ] -} -``` +**Query workload.** Two PromQL dashboard panels over the last minute. The +first needs an exact total; the second tolerates error. + +| Workload field | Value (both queries) | +|---|---| +| `language` | `promql` | +| Entry type | `repeating_queries` | +| `demand` | every 10 s (`fixed_interval`) | +| `predictability` | `predictable` | +| `time_selection.scope` | `real_time` | + +| Query | `lookback` | `as_of` | Accuracy | Latency | +|---|---|---|---|---| +| `sum by (job) (rate(http_requests_total[1m]))` | 1 m | evaluation time | exact (`implicit_exact`) | unspecified | +| `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | **Stage 0.** The two `LogicalDAG`s are `range m[1m] → rate → sum by (job)` and @@ -376,20 +362,29 @@ cheaper; with a few large jobs, per-`job` Count-Min sketches may win. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 **Query workload.** A network-monitoring dashboard computes three statistics of -source IPs over the last minute, every 10 s. Each SQL query below is one -`RepeatingEntry` with `"demand": {"fixed_interval": 10000}` and -`"time_selection": {"scope": "real_time", "lookback": 60000, "as_of": null}`, -in a workload with `"language": {"sql": "datafusion_sql"}`. - -| Query | Computation | Accuracy | -|---|---|---| -| `SELECT COUNT(DISTINCT src_ip) FROM flows WHERE ts >= now() - INTERVAL '1 minute'` | `Distinct(src_ip)` | ε = 0.02, δ = 0.01 | -| `SELECT -SUM(p * LN(p)) FROM (SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `Entropy(src_ip)` | ε = 0.05, δ = 0.01 | -| `SELECT SQRT(SUM(c * c)) FROM (SELECT src_ip, COUNT(*) AS c FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `L2(src_ip)` | ε = 0.01, δ = 0.01 | +source IPs over the last minute. -The data workload is as above, except that `input_cardinality` is the number of -distinct source IPs (say 10 million) and `data_ingestion_interval` is not -needed for SQL. +| Workload field | Value (all three queries) | +|---|---| +| `language` | `sql` (`datafusion_sql`) | +| Entry type | `repeating_queries` | +| `demand` | every 10 s (`fixed_interval`) | +| `predictability` | `predictable` | +| `time_selection.scope` | `real_time` | +| Latency | unspecified | + +| Query | Computation | `lookback` | `as_of` | Accuracy | +|---|---|---|---|---| +| `SELECT COUNT(DISTINCT src_ip) FROM flows WHERE ts >= now() - INTERVAL '1 minute'` | `Distinct(src_ip)` | 1 m | evaluation time | ε = 0.02, δ = 0.01 | +| `SELECT -SUM(p * LN(p)) FROM (SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `Entropy(src_ip)` | 1 m | evaluation time | ε = 0.05, δ = 0.01 | +| `SELECT SQRT(SUM(c * c)) FROM (SELECT src_ip, COUNT(*) AS c FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `L2(src_ip)` | 1 m | evaluation time | ε = 0.01, δ = 0.01 | + +The data workload differs from the shared one in two fields: + +| `DataWorkload` field | Value | +|---|---| +| `input_cardinality` | 10,000,000 distinct source IPs (declared) | +| `data_ingestion_interval` | not needed for SQL | **Pass 1.** Rewrite rules recognize the three computations, and each gets its local candidates from the Pass 1 table: exact, a specialized summary, or @@ -404,14 +399,14 @@ requirement, ε = 0.01. The independent candidates are kept as well. ```mermaid flowchart LR in[("flows.src_ip
last 1m")]:::data - + subgraph P0["Stage 0 · LogicalDAGs"] direction TB q1["Distinct(src_ip)"]:::exact q2["Entropy(src_ip)"]:::exact q3["L2(src_ip)"]:::exact end - + subgraph P1["Stage 1, Pass 1 · local candidates per computation"] direction TB subgraph D["Distinct"] @@ -433,7 +428,7 @@ flowchart LR l2["UnivMon"]:::summary end end - + subgraph P2["Stage 1, Pass 2 · summary-capability rule adds a shared candidate"] direction LR u["one UnivMon
sized for ε = 0.01"]:::summary @@ -441,7 +436,7 @@ flowchart LR u --> re(["entropy"]):::readout u --> rl(["L2 norm"]):::readout end - + in --> P0 q1 --> D q2 --> E @@ -449,7 +444,7 @@ flowchart LR d2 -. "same summary input data
and window" .-> u e2 -.-> u l2 -.-> u - + OUT[["CandidateLogicalASAPDAGs:
all Pass 1 candidates + the shared candidate"]] P1 --> OUT P2 --> OUT @@ -458,10 +453,10 @@ flowchart LR classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef readout fill:#e6f4ea,stroke:#188038,color:#000; ``` - -Exact candidates for each computation are also kept but omitted from the -diagram. +The three UnivMon options from Pass 1 (dashed arrows) are merged by Pass 2 into +one shared UnivMon. The independent candidates are kept, so stage 1 outputs +both. **Stage 3.** Selection compares one UnivMon sized for ε = 0.01 against three separate summaries, each sized for its own requirement. The shared candidate @@ -474,7 +469,18 @@ summary. **Pattern A: a batch of sub-interval queries over historical data.** An analyst submits a batch of p99 latency reports over different historical intervals, -executed together at T = `1790000000000`: +all executed together at time T. + +| Workload field | Value (all five queries) | +|---|---| +| `language` | `promql` | +| Entry type | `query_batch` | +| `invocations` | 1 | +| `execute_at` | T | +| `predictability` | `ad_hoc` | +| `time_selection.scope` | `longitudinal` | +| Accuracy | ε = 0.005, δ = 0.01 | +| Latency | unspecified | | Query | `lookback` | `as_of` | |---|---|---| @@ -484,21 +490,8 @@ executed together at T = `1790000000000`: | `quantile_over_time(0.99, latency_ms[1y] offset 2y)` | 1 y | T − 2 y | | `quantile_over_time(0.99, latency_ms[3y] offset 2y)` | 3 y | T − 2 y | -Each row is a `BatchEntry` such as: - -```json -{ - "query": "quantile_over_time(0.99, latency_ms[1y] offset 1y)", - "requirements": {"accuracy": {"explicit": {"EpsilonDelta": {"epsilon": 0.005, "delta": 0.01}}}, "response_latency": "unspecified"}, - "predictability": "ad_hoc", - "invocations": 1, - "execute_at": 1790000000000, - "time_selection": {"scope": "longitudinal", "lookback": 31536000000, "as_of": 1758464000000} -} -``` - -The data workload has `"arrival": "mixed"`: five years at rest plus data still -arriving. +The data workload is the shared one, except `arrival` is `mixed`: five years +of data at rest, plus data still arriving. * **Pass 1.** Each `quantile_over_time` gets an exact candidate and a KLL over its own interval: five independent KLL candidates over overlapping data. @@ -541,17 +534,19 @@ flowchart LR ``` **Pattern B: one repeating query over a sliding window.** A real-time p99 panel -over the last 5 min, refreshed every minute: - -```json -{ - "query": "quantile_over_time(0.99, latency_ms[5m])", - "demand": {"fixed_interval": 60000}, - "requirements": {"accuracy": {"explicit": {"EpsilonDelta": {"epsilon": 0.01, "delta": 0.01}}}, "response_latency": {"explicit_max_ms": 200.0}}, - "predictability": {"predictable": {"known_at": null}}, - "time_selection": {"scope": "real_time", "lookback": 300000, "as_of": null} -} -``` +over the last 5 min, refreshed every minute. + +| Workload field | Value | +|---|---| +| `language` | `promql` | +| Entry type | `repeating_queries` | +| `demand` | every 1 min (`fixed_interval`) | +| `predictability` | `predictable` | +| `time_selection.scope` | `real_time` | + +| Query | `lookback` | `as_of` | Accuracy | Latency | +|---|---|---|---|---| +| `quantile_over_time(0.99, latency_ms[5m])` | 5 m | evaluation time | ε = 0.01, δ = 0.01 | ≤ 200 ms | * **Pass 1.** One KLL over 5 min for each evaluation. * **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The From e4379bdf310a930cbf636424db7186e33c784cdc Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 12:32:39 -0400 Subject: [PATCH 31/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index a97622b9..96024713 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -180,6 +180,20 @@ or value being summarized together with its grouping. The summary input data doe not include the window; the window-composition rule compares windows separately. +The window-composition rule shares a summary across windows by splitting time +into pieces that each carry their own summary: + +* A **pane** is a fixed-length, non-overlapping slice of time, for example one + minute, with one summary of the data that arrived in that slice. A query + window is answered by merging the summaries of the panes it covers. The pane + length is chosen so that every requested window is an exact union of panes: + for a 5-min window evaluated every 1 min, 1-min panes work, since each window + is exactly 5 consecutive panes. +* A **bucket** of an Exponential Histogram plays the same role, but bucket + lengths grow with age: recent data sits in short buckets and older data in + longer ones. This keeps few buckets over a long history, at the cost that old + window boundaries may fall inside a bucket and are then approximate. + | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| | Identical-expression rule | The input and computation semantics are identical. | One common computation node serving multiple consumers. | From 560edc3b9fe41c9a281366b606a4323f96703041 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 12:46:41 -0400 Subject: [PATCH 32/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 208 +++++++++++------- 1 file changed, 129 insertions(+), 79 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 96024713..79ee18b3 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -1,16 +1,19 @@ -# ASAPPlanner Layering Design +# ASAPPlanner Planning Stages Design Status: proposal. Audience: designers and developers of ASAPPlanner and of deployments such as ASAPQuery-backend. +Read Stages for the overview, the stage sections for the rules, and the +examples for why the rules are needed. + ## Goal -ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the [deployment's -inputs](TODO: A data structure should be explicitly defined in another PR), and returns one optimal physical plan. It decides what is computed, how +ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the deployment's +inputs (TODO: define this data structure in a follow-up PR), and returns one optimal physical plan. It decides what is computed, how it is computed, and which plan is best. The deployment only supplies inputs and executes the plan: it supplies its own empirical cost estimation, empirical accuracy estimation and capabilities of deployment but never does the query planning or plan selection. -## Layers +## Stages `x` represents Cartesian product for enumerating and combining different optimization angles in planning. @@ -98,13 +101,7 @@ The planner receives three groups of inputs: | Data workload | Data arrival pattern, sampling cadence, ingestion volume and rate, cardinality, and distribution | | Deployment inputs | Empirical cost model, empirical accuracy model, and execution capabilities | -Logical planning decides **what** to compute: summary families, query rewrites, -and sharing across computations and windows. Physical planning decides **how** -to compute it: materialization, execution placement, physical operators, -partitioning, parallelism, and resource allocation. Selection evaluates -complete physical candidates and chooses one plan for the whole workload. - -## Layers/DAGs and what decisions each layer makes +## Stages and their decisions Stages 0 to 2 each output a candidate set holding every semantically equivalent and legal candidate DAG of that stage; stage 3 is the only step that chooses @@ -116,8 +113,8 @@ query's accuracy target), and every rejected candidate carries a reason. | Stage | Input | Decides | Output | |---|---|---|---| | 0. Frontends | Each `QueryWorkloadEntry.query` and `QueryWorkload.language` | Language semantics, and converting source-language queries into a common logical representation. Rejects constructs it cannot represent faithfully. | `CandidateLogicalDAGs`; nodes are logical operations with no summary operations | -| 1. Logical ASAP-aware optimization | Logical DAGs, and each query's accuracy requirements across the workload | **Pass 1 — Summary replacement and query rewriting:** apply query rewriting rules to each eligible sub-DAG and generate exact and summary-based candidates that satisfy its semantics and accuracy requirements. **Pass 2 — ASAP-aware CSE:** apply traditional CSE and summary-specific CSE rules across sub-DAGs and queries to generate shared computation candidates, while preserving independent candidates. | `CandidateLogicalASAPDAGs`; nodes include summary and window-primitive operations | -| 2. Physical ASAP-aware optimization | Logical ASAP DAGs, each entry's `recurrence` and `predictability`, the `DataWorkload` | **Materialization:** for each sub-DAG, whether its output is materialized and therefore computed at ingestion time or query time, and how long it is retained. **Physical operator implementation:** physical operators for every node. **Parallelism, partitioning, resources:** TODO. | `CandidatePhysicalASAPDAGs` | +| 1. Logical ASAP-aware optimization | Logical DAGs, and each query's accuracy requirement, `time_selection` and repetition interval (from `recurrence`) across the workload | **Pass 1 — Summary replacement and query rewriting:** apply query rewriting rules to each eligible sub-DAG and generate exact and summary-based candidates that satisfy its semantics and accuracy requirements. **Pass 2 — ASAP-aware CSE:** apply traditional CSE and summary-specific CSE rules across sub-DAGs and queries to generate shared computation candidates, while preserving independent candidates. | `CandidateLogicalASAPDAGs`; nodes include summary and window-summary operations | +| 2. Physical ASAP-aware optimization | Logical ASAP DAGs, each entry's `recurrence` and `predictability`, the `DataWorkload` | **Materialization:** for each sub-DAG, whether its output is materialized, when it is computed (ingestion time or query time), and how long it is retained. **Physical operator implementation:** physical operators for every node. **Parallelism, partitioning, resources:** TODO. | `CandidatePhysicalASAPDAGs` | | 3. Plan selection | Physical ASAP DAG candidates, `requirements`, the deployment's cost model, accuracy model and capabilities | Rejects candidates that miss an accuracy target, a latency bound or a capability; picks the cheapest valid plan for the whole workload. A shared state is costed once with all its consumers' demand. | One `PhysicalASAPDAG` | | 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Executes ingestion, storage, precomputation and query-time computation. | Query results | @@ -128,8 +125,7 @@ query operations, including selectors, transformations, aggregations, grouping and window semantics. They contain no ASAP summary choices. The frontend preserves source-language behavior, including series identity, -evaluation timing and missing-data semantics. A construct that cannot be -represented faithfully is rejected. +evaluation timing and missing-data semantics. ### 1. Logical ASAP-aware optimization @@ -137,7 +133,9 @@ Logical optimization runs in two passes. Pass 1 generates candidates for each computation on its own; Pass 2 finds candidates that share computation across sub-DAGs and queries. Neither pass decides materialization or execution placement, and neither picks one summary per computation: every candidate is -kept for selection. +kept for selection. Stage 1 reads the repetition interval only to detect +windows that overlap across evaluations; deciding when anything is computed is +left to stage 2. #### Pass 1: Local candidate generation @@ -158,7 +156,7 @@ Example for summary candidates: Each candidate records its input expression, filter, grouping, window, supported readouts and accuracy requirement. Pass 2 uses these to decide -whether candidates can share a producer. +whether candidates can share a summary node. A summary-based candidate has two kinds of nodes. The **summary node** builds and maintains the summary from the input data, for example a KLL sketch over @@ -170,8 +168,8 @@ exploits. #### Pass 2: ASAP-aware common-subexpression elimination ASAP-aware CSE extends traditional CSE with summary-specific sharing rules. -Computations can share a producer when they use identical expressions, when -one summary supports multiple readouts, or when summary composition supports +Computations can share work when they use identical expressions, when one +summary node supports several readouts, or when one window summary can answer their overlapping windows. The rules compare computations by their **summary input data**: what a summary for @@ -193,12 +191,39 @@ into pieces that each carry their own summary: lengths grow with age: recent data sits in short buckets and older data in longer ones. This keeps few buckets over a long history, at the cost that old window boundaries may fall inside a bucket and are then approximate. +* A **window summary** is a summary organized as panes or buckets so that it + can answer many windows, for example a sliding window of panes, a tumbling + window, or an Exponential Histogram. | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| | Identical-expression rule | The input and computation semantics are identical. | One common computation node serving multiple consumers. | -| Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary node connected to multiple readout nodes; for example, one UnivMon over `src_ip` from `flows` in the last minute serves three queries refreshed every 10 s: `COUNT(DISTINCT src_ip)`, the entropy of the `src_ip` distribution, and the L2 norm of per-`src_ip` counts. Each flow record updates the one UnivMon node once; a distinct-count readout node, an entropy readout node and an L2 readout node each read their statistic from it. The UnivMon is sized for the strictest of the three accuracy requirements (Example 2). | -| Window-composition rule | The computations have the same summary input data, and the summary's merge and window capabilities can reconstruct the requested windows within their accuracy requirements. | Shared window-summary nodes connected to window-specific reconstruction and readout nodes; for example, (1) one sliding window of 1-min KLL panes serves every evaluation of `quantile_over_time(0.99, latency_ms[5m])` repeated every minute: each evaluation's readout node merges the latest 5 panes, so consecutive evaluations share 4 of their 5 panes (Example 3, Pattern B); (2) one Exponential Histogram of KLL buckets over the last 5 years serves the p99 queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and `[3y] offset 2y`: each query's reconstruction node merges the buckets covering its interval, and its readout node reads p99 from the merged KLL (Example 3, Pattern A). The same panes or buckets also serve other quantiles, since one KLL answers every quantile: adding `quantile_over_time(0.5, latency_ms[5m])` to the dashboard in (1) adds only a p50 readout node next to the p99 one, both reading the same 5 merged panes, with no new summary. | +| Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary node feeding several readout nodes, e.g. UnivMon → distinct count, entropy, L2 norm. | +| Window-composition rule | The computations have the same summary input data, and one window summary can reconstruct the requested windows within their accuracy requirements. | One window summary feeding per-window reconstruction and readout nodes, e.g. KLL panes in a sliding window, or KLL buckets in an Exponential Histogram. | + +The examples behind these rules: + +* **Summary-capability rule (Example 2).** One UnivMon over `src_ip` from + `flows` in the last minute serves three queries refreshed every 10 s: + `COUNT(DISTINCT src_ip)`, the entropy of the `src_ip` distribution, and the + L2 norm of per-`src_ip` counts. Each flow record updates the UnivMon once; a + distinct-count, an entropy and an L2 readout node each read their statistic + from it. The UnivMon is sized for the strictest of the three accuracy + requirements. +* **Window-composition rule, sliding window (Example 3, Pattern B).** One + sliding window of 1-min KLL panes serves every evaluation of + `quantile_over_time(0.99, latency_ms[5m])`, repeated every minute. Each + evaluation's readout node merges the latest 5 panes, so consecutive + evaluations share 4 of their 5 panes. +* **Window-composition rule, Exponential Histogram (Example 3, Pattern A).** + One Exponential Histogram of KLL buckets over the last 5 years serves the p99 + queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and + `[3y] offset 2y`. Each query's reconstruction node merges the buckets + covering its interval, and its readout node reads p99 from the merged KLL. +* **Other quantiles share for free.** One KLL answers every quantile, so adding + `quantile_over_time(0.5, latency_ms[5m])` to the sliding-window dashboard + adds only a p50 readout node next to the p99 one, reading the same 5 merged + panes, with no new summary. Rules are defined by each summary family's capabilities and semantic requirements. A shared summary must meet the strictest accuracy requirement @@ -240,7 +265,7 @@ constraints: query time. The decision depends on the workload's `recurrence` and `predictability` and on -the `DataWorkload`, none of which logical planning reads. Typical outcomes: +the `DataWorkload`. Typical outcomes: * Read by repeated queries while data keeps arriving: materialize at ingestion time. @@ -258,16 +283,16 @@ merge and quantile readout operators. ### 3. Plan selection -Selection rejects every physical candidate that the deployment's accuracy model -estimates will miss an accuracy target, that misses a latency bound, or that -needs a capability the deployment lacks. Among the remaining candidates, it -uses the deployment's cost model to choose the cheapest plan for the whole -workload. A shared state is costed once, with the demand of all its consumers. +Selection is the only stage that uses the deployment's cost and accuracy +models, and the only stage that discards valid candidates. Accuracy is +estimated by the deployment's accuracy model, not assumed from a summary's +nominal bound. Cost is evaluated for the whole workload rather than per query, +which is what lets one shared summary beat several cheaper independent ones. ### 4. Execution -The deployment executes the selected `PhysicalASAPDAG`: ingestion, storage, -precomputation and query-time computation. It makes no planning decisions. +Execution runs outside ASAPPlanner. The deployment runs the selected plan as +given: it does not choose among summaries or decide what to materialize. ## End-to-end examples @@ -275,7 +300,8 @@ Each example's workload is shown as tables. Field names in code font are the fields of [`workload.rs`](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). An `as_of` of "evaluation time" means `as_of: None`: the window ends when the -query runs. +query runs. Approximate accuracy targets are `EpsilonDelta`: the answer's error +is at most ε with probability at least 1 − δ. | Example | Shows | |---|---| @@ -320,8 +346,8 @@ first needs an exact total; the second tolerates error. | `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | **Stage 0.** The two `LogicalDAG`s are -`range m[1m] → rate → sum by (job)` and -`range m[1m] → sum_over_time → topk by (job) (10)`. +`range http_requests_total[1m] → rate → sum by (job)` and +`range http_requests_total[1m] → sum_over_time → topk by (job) (10)`. **Stage 1, Pass 1.** @@ -385,13 +411,39 @@ source IPs over the last minute. | `demand` | every 10 s (`fixed_interval`) | | `predictability` | `predictable` | | `time_selection.scope` | `real_time` | +| `as_of` | evaluation time | | Latency | unspecified | -| Query | Computation | `lookback` | `as_of` | Accuracy | -|---|---|---|---|---| -| `SELECT COUNT(DISTINCT src_ip) FROM flows WHERE ts >= now() - INTERVAL '1 minute'` | `Distinct(src_ip)` | 1 m | evaluation time | ε = 0.02, δ = 0.01 | -| `SELECT -SUM(p * LN(p)) FROM (SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `Entropy(src_ip)` | 1 m | evaluation time | ε = 0.05, δ = 0.01 | -| `SELECT SQRT(SUM(c * c)) FROM (SELECT src_ip, COUNT(*) AS c FROM flows WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip)` | `L2(src_ip)` | 1 m | evaluation time | ε = 0.01, δ = 0.01 | +```sql +-- Q1: Distinct(src_ip) +SELECT COUNT(DISTINCT src_ip) +FROM flows +WHERE ts >= now() - INTERVAL '1 minute'; + +-- Q2: Entropy(src_ip) +SELECT -SUM(p * LN(p)) +FROM ( + SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p + FROM flows + WHERE ts >= now() - INTERVAL '1 minute' + GROUP BY src_ip +); + +-- Q3: L2(src_ip) +SELECT SQRT(SUM(c * c)) +FROM ( + SELECT src_ip, COUNT(*) AS c + FROM flows + WHERE ts >= now() - INTERVAL '1 minute' + GROUP BY src_ip +); +``` + +| Query | Computation | `lookback` | Accuracy | +|---|---|---|---| +| Q1 | `Distinct(src_ip)` | 1 m | ε = 0.02, δ = 0.01 | +| Q2 | `Entropy(src_ip)` | 1 m | ε = 0.05, δ = 0.01 | +| Q3 | `L2(src_ip)` | 1 m | ε = 0.01, δ = 0.01 | The data workload differs from the shared one in two fields: @@ -589,51 +641,13 @@ gantt eval at 00:07 (panes 3–7) :e3, 00:02, 5m ``` -In both patterns, stage 1 decides only **which** window summary computes the -answer. It says nothing about when panes or buckets are built or whether they -are stored. +Example 4 shows how stage 2 decides whether to store these window summaries. ### Example 4: Materialization of window summaries in physical planning Stage 2 takes the shared window summaries from Example 3 and decides whether to materialize them. That choice is driven by the workload's `recurrence`, -`predictability` and `data_workload.arrival`, none of which stage 1 reads. - -**Pattern B (sliding window, repeating).** - -| Candidate | Materialized | At ingestion time | At query time | -|---|---|---|---| -| B1 | 1-min KLL panes, retained 5 min | Build one KLL pane per minute | Merge the latest 5 panes, read p99 | -| B2 | Nothing | Nothing | Read 5 min of raw samples, build one KLL, read p99 | - -```mermaid -flowchart LR - subgraph B1["B1 · panes materialized at ingestion time"] - direction LR - subgraph B1I["Ingestion time"] - s1[("samples")]:::data --> p1["1-min KLL pane
stored 5 min"]:::summary - end - subgraph B1Q["Query time, every 1 min"] - g1["merge latest
5 panes"]:::exact --> r1(["p99"]):::readout - end - p1 --> g1 - end - subgraph B2["B2 · not materialized"] - direction LR - subgraph B2Q["Query time, every 1 min"] - s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::readout - end - end - classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; - classDef exact fill:#fff,stroke:#5f6368,color:#000; - classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef readout fill:#e6f4ea,stroke:#188038,color:#000; -``` - -The query repeats every minute and the data is continuously ingesting, so B1 -builds each pane once and reuses it in five evaluations, while B2 rescans raw -data every time. Selection usually picks B1. B2 wins only if storage is -expensive and raw data is available at query time. +`predictability` and `data_workload.arrival`. **Pattern A (sub-interval batch).** @@ -676,8 +690,44 @@ Which candidate wins depends on the workload: * **Data `"at_rest"`:** A2 is not generated, because there is no ingestion to maintain the histogram. -**What this shows.** The same logical candidate (a sliding window of KLL panes, -or one shared Exponential Histogram) yields different physical plans depending +**Pattern B (sliding window, repeating).** + +| Candidate | Materialized | At ingestion time | At query time | +|---|---|---|---| +| B1 | 1-min KLL panes, retained 5 min | Build one KLL pane per minute | Merge the latest 5 panes, read p99 | +| B2 | Nothing | Nothing | Read 5 min of raw samples, build one KLL, read p99 | + +```mermaid +flowchart LR + subgraph B1["B1 · panes materialized at ingestion time"] + direction LR + subgraph B1I["Ingestion time"] + s1[("samples")]:::data --> p1["1-min KLL pane
stored 5 min"]:::summary + end + subgraph B1Q["Query time, every 1 min"] + g1["merge latest
5 panes"]:::exact --> r1(["p99"]):::readout + end + p1 --> g1 + end + subgraph B2["B2 · not materialized"] + direction LR + subgraph B2Q["Query time, every 1 min"] + s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::readout + end + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef readout fill:#e6f4ea,stroke:#188038,color:#000; +``` + +The query repeats every minute and the data is continuously ingesting, so B1 +builds each pane once and reuses it in five evaluations, while B2 rescans raw +data every time. Selection usually picks B1. B2 wins only if storage is +expensive and raw data is available at query time. + +**What this shows.** The same logical candidate (one shared Exponential +Histogram, or a sliding window of KLL panes) yields different physical plans depending only on recurrence, predictability and data arrival. This is why window-summary replacement happens in logical planning, while materialization is decided separately in physical planning. From 7c5cc97b23923dd8467df471f59a2bf5a0273203 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 12:51:32 -0400 Subject: [PATCH 33/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 79ee18b3..103c31e1 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -61,7 +61,7 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu │ Physical planning — how to compute │ │ │ │ 2. Physical ASAP-aware optimization │ -│ Explore executable implementations of each logical candidate: │ +│ Explore physical implementations of each logical candidate: │ │ │ │ materialization decisions │ │ × physical operator implementations │ @@ -232,7 +232,7 @@ independent candidates, so selection can compare both. ### 2. Physical ASAP-aware optimization -Physical optimization turns each logical candidate into executable candidates. +Physical optimization turns each logical candidate into physical candidates. It makes two ASAP-specific decisions, described below. Parallelism, partitioning and resource management are TODO. From faf3c046008db0f1934fbbbe200703f3a4a31aa2 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:13:00 -0400 Subject: [PATCH 34/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 123 ++++++++++-------- 1 file changed, 67 insertions(+), 56 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 103c31e1..138f6baf 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -61,7 +61,7 @@ executes the plan: it supplies its own empirical cost estimation, empirical accu │ Physical planning — how to compute │ │ │ │ 2. Physical ASAP-aware optimization │ -│ Explore physical implementations of each logical candidate: │ +│ Explore executable implementations of each logical candidate: │ │ │ │ materialization decisions │ │ × physical operator implementations │ @@ -155,21 +155,27 @@ Example for summary candidates: | `Quantile(x, window)` | Exact quantile, KLL over the requested window | Each candidate records its input expression, filter, grouping, window, -supported readouts and accuracy requirement. Pass 2 uses these to decide +supported estimates and accuracy requirement. Pass 2 uses these to decide whether candidates can share a summary node. -A summary-based candidate has two kinds of nodes. The **summary node** builds -and maintains the summary from the input data, for example a KLL sketch over -`latency_ms`. A **readout node** computes a query's answer from that summary, -for example the p99 estimate from the KLL, or the entropy estimate from a -UnivMon. One summary node can feed several readout nodes, which is what Pass 2 +A summary-based candidate uses three kinds of summary nodes: + +* A **summary build node** builds and maintains a summary from input data, for + example a KLL sketch over `latency_ms`. +* A **summary merge node** combines summaries into one, for example merging + five 1-min KLL panes into one 5-min KLL, or merging lower-level summaries + into a coarser one. +* A **summary estimation node** computes an answer from a summary, for example + the p99 estimate from a KLL, or the entropy estimate from a UnivMon. + +One summary build node can feed several estimation nodes, which is what Pass 2 exploits. #### Pass 2: ASAP-aware common-subexpression elimination ASAP-aware CSE extends traditional CSE with summary-specific sharing rules. Computations can share work when they use identical expressions, when one -summary node supports several readouts, or when one window summary can answer +summary build node supports several estimates, or when one window summary can answer their overlapping windows. The rules compare computations by their **summary input data**: what a summary for @@ -198,8 +204,8 @@ into pieces that each carry their own summary: | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| | Identical-expression rule | The input and computation semantics are identical. | One common computation node serving multiple consumers. | -| Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary node feeding several readout nodes, e.g. UnivMon → distinct count, entropy, L2 norm. | -| Window-composition rule | The computations have the same summary input data, and one window summary can reconstruct the requested windows within their accuracy requirements. | One window summary feeding per-window reconstruction and readout nodes, e.g. KLL panes in a sliding window, or KLL buckets in an Exponential Histogram. | +| Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary build node feeding several estimation nodes, e.g. UnivMon → distinct count, entropy, L2 norm. | +| Window-composition rule | The computations have the same summary input data, and one window summary can reconstruct the requested windows within their accuracy requirements. | One window summary feeding per-window merge and estimation nodes, e.g. KLL panes in a sliding window, or KLL buckets in an Exponential Histogram. | The examples behind these rules: @@ -207,22 +213,22 @@ The examples behind these rules: `flows` in the last minute serves three queries refreshed every 10 s: `COUNT(DISTINCT src_ip)`, the entropy of the `src_ip` distribution, and the L2 norm of per-`src_ip` counts. Each flow record updates the UnivMon once; a - distinct-count, an entropy and an L2 readout node each read their statistic - from it. The UnivMon is sized for the strictest of the three accuracy + distinct-count, an entropy and an L2 estimation node each compute their + statistic from it. The UnivMon is sized for the strictest of the three accuracy requirements. * **Window-composition rule, sliding window (Example 3, Pattern B).** One sliding window of 1-min KLL panes serves every evaluation of `quantile_over_time(0.99, latency_ms[5m])`, repeated every minute. Each - evaluation's readout node merges the latest 5 panes, so consecutive - evaluations share 4 of their 5 panes. + evaluation merges the latest 5 panes with a merge node and computes p99 with + an estimation node, so consecutive evaluations share 4 of their 5 panes. * **Window-composition rule, Exponential Histogram (Example 3, Pattern A).** One Exponential Histogram of KLL buckets over the last 5 years serves the p99 queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and - `[3y] offset 2y`. Each query's reconstruction node merges the buckets - covering its interval, and its readout node reads p99 from the merged KLL. + `[3y] offset 2y`. Each query's merge node merges the buckets covering + its interval, and its estimation node computes p99 from the merged KLL. * **Other quantiles share for free.** One KLL answers every quantile, so adding `quantile_over_time(0.5, latency_ms[5m])` to the sliding-window dashboard - adds only a p50 readout node next to the p99 one, reading the same 5 merged + adds only a p50 estimation node next to the p99 one, reading the same 5 merged panes, with no new summary. Rules are defined by each summary family's capabilities and semantic @@ -232,7 +238,7 @@ independent candidates, so selection can compare both. ### 2. Physical ASAP-aware optimization -Physical optimization turns each logical candidate into physical candidates. +Physical optimization turns each logical candidate into executable candidates. It makes two ASAP-specific decisions, described below. Parallelism, partitioning and resource management are TODO. @@ -261,8 +267,6 @@ constraints: needs it. * Query-time work cannot feed ingestion-time work, so every node feeding an ingestion-time sub-DAG also runs at ingestion time. -* Readout nodes, which compute a query's answer from a summary, always run at - query time. The decision depends on the workload's `recurrence` and `predictability` and on the `DataWorkload`. Typical outcomes: @@ -279,7 +283,7 @@ A shared summary is materialized once for all its consumers. See Example 4. Physical operator implementation lowers every node to physical operators, for example TopK as a sort followed by a limit, or a KLL node as summary build, -merge and quantile readout operators. +merge and quantile estimation operators. ### 3. Plan selection @@ -299,8 +303,11 @@ given: it does not choose among summaries or decide what to materialize. Each example's workload is shown as tables. Field names in code font are the fields of [`workload.rs`](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). -An `as_of` of "evaluation time" means `as_of: None`: the window ends when the -query runs. Approximate accuracy targets are `EpsilonDelta`: the answer's error +Each query reads the event-time window [`as_of` − `lookback`, `as_of`] (fields +of `TimeSelection`). `as_of` is the window's end; `lookback` is its length. An +`as_of` of "evaluation time" means `as_of: None`: the window ends whenever the +query runs, so it moves forward with each evaluation. A fixed `as_of`, such as +T − 1 y, pins the window to a historical interval. Approximate accuracy targets are `EpsilonDelta`: the answer's error is at most ε with probability at least 1 − δ. | Example | Shows | @@ -311,21 +318,25 @@ is at most ε with probability at least 1 − δ. | 4. Materialization of window summaries | Stage 2: materialization, decoupled from stage 1 | In the diagrams below, grey cylinders are input data, white boxes are exact -operations, blue boxes are summary nodes, and green rounded boxes are readout -nodes. +operations and summary merges, blue boxes are summary build nodes, and green +rounded boxes are summary estimation nodes. ### Shared data workload Unless an example says otherwise, every example uses this data workload: -| `DataWorkload` field | Value | -|---|---| -| `arrival` | `continuously_ingesting` | -| `data_ingestion_interval` | 15 s (declared) | -| `ingestion_volume` | unknown | -| `ingestion_rate` | about 66,667 samples/s (declared) | -| `input_cardinality` | 1,000,000 series (declared) | -| `distribution` | `zipf` (declared) | +| `DataWorkload` field | Meaning | Value | +|---|---|---| +| `arrival` | Whether the data is at rest, still arriving, or both | `continuously_ingesting` | +| `data_ingestion_interval` | How often each series delivers one sample (the scrape interval in Prometheus). PromQL uses it as the look-back horizon of instant selectors. | 15 s (declared) | +| `ingestion_volume` | Total amount of ingested data | unknown | +| `ingestion_rate` | Samples arriving per second across all series | about 66,667 samples/s (declared) | +| `input_cardinality` | Number of distinct series (or keys) | 1,000,000 series (declared) | +| `distribution` | How samples are spread over keys | `zipf` (declared) | + +"Declared" is the value's `EvidenceSource`: the workload author stated it +rather than the planner observing it. With 1,000,000 series each sampled every +15 s, the ingestion rate is 1,000,000 / 15 ≈ 66,667 samples/s. ### Example 1: Aggregation over dimensions — summary replacement in Pass 1 @@ -379,11 +390,11 @@ flowchart LR end subgraph CM["Count-Min + heap per job"] direction LR - c1[("input")]:::data --> c2["Count-Min Sketch +
top-10 heap, one per job"]:::summary --> c3(["top 10
per job"]):::readout + c1[("input")]:::data --> c2["Count-Min Sketch +
top-10 heap, one per job"]:::summary --> c3(["top 10
per job"]):::estimate end subgraph H["Hydra"] direction LR - h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::readout + h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::estimate end end S{{"Stage 3 · cheapest valid candidate
many small jobs → Hydra
few large jobs → Count-Min per job"}} @@ -391,7 +402,7 @@ flowchart LR classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef readout fill:#e6f4ea,stroke:#188038,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` **Stage 3.** The deployment's accuracy model checks that each summary candidate @@ -458,8 +469,8 @@ UnivMon. **Pass 2.** All three computations have the same summary input data (`src_ip` from `flows`, no other filter) and the same 1-min window. UnivMon supports all three -readouts, so the summary-capability rule adds a shared candidate: **one -UnivMon node feeding three readout nodes**. It must be sized for the strictest +estimates, so the summary-capability rule adds a shared candidate: **one +UnivMon build node feeding three estimation nodes**. It must be sized for the strictest requirement, ε = 0.01. The independent candidates are kept as well. ```mermaid @@ -498,9 +509,9 @@ flowchart LR subgraph P2["Stage 1, Pass 2 · summary-capability rule adds a shared candidate"] direction LR u["one UnivMon
sized for ε = 0.01"]:::summary - u --> rd(["distinct count"]):::readout - u --> re(["entropy"]):::readout - u --> rl(["L2 norm"]):::readout + u --> rd(["distinct count"]):::estimate + u --> re(["entropy"]):::estimate + u --> rl(["L2 norm"]):::estimate end in --> P0 @@ -517,7 +528,7 @@ flowchart LR classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef readout fill:#e6f4ea,stroke:#188038,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` The three UnivMon options from Pass 1 (dashed arrows) are merged by Pass 2 into @@ -564,7 +575,7 @@ of data at rest, plus data still arriving. * **Pass 2.** Every interval is a sub-interval of [T − 5 y, T], and KLL is mergeable. The window-composition rule adds a shared candidate: **one Exponential Histogram of KLL buckets over [T − 5 y, T]**, with one - reconstruction and readout node per query that merges the buckets covering + merge and estimation node per query that merges the buckets covering its interval. The five independent candidates are kept. The five query intervals overlap, and all lie inside the last five years: @@ -583,20 +594,20 @@ gantt ``` The shared candidate replaces five KLL sketches with one Exponential Histogram -and a reconstruction and readout node per query: +and a merge and estimation node per query: ```mermaid flowchart LR in[("latency_ms
T − 5y to T")]:::data --> eh["Exponential Histogram
of KLL buckets"]:::summary - eh --> m1["merge buckets
T−5y … T"]:::exact --> o1(["q1 p99"]):::readout - eh --> m2["merge buckets
T−1y … T"]:::exact --> o2(["q2 p99"]):::readout - eh --> m3["merge buckets
T−2y … T−1y"]:::exact --> o3(["q3 p99"]):::readout - eh --> m4["merge buckets
T−3y … T−2y"]:::exact --> o4(["q4 p99"]):::readout - eh --> m5["merge buckets
T−5y … T−2y"]:::exact --> o5(["q5 p99"]):::readout + eh --> m1["merge buckets
T−5y … T"]:::exact --> o1(["q1 p99"]):::estimate + eh --> m2["merge buckets
T−1y … T"]:::exact --> o2(["q2 p99"]):::estimate + eh --> m3["merge buckets
T−2y … T−1y"]:::exact --> o3(["q3 p99"]):::estimate + eh --> m4["merge buckets
T−3y … T−2y"]:::exact --> o4(["q4 p99"]):::estimate + eh --> m5["merge buckets
T−5y … T−2y"]:::exact --> o5(["q5 p99"]):::estimate classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef readout fill:#e6f4ea,stroke:#188038,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` **Pattern B: one repeating query over a sliding window.** A real-time p99 panel @@ -662,7 +673,7 @@ flowchart LR direction LR subgraph A1Q["Query time, once at T"] s3[("5 years of
stored samples")]:::data --> h3["build Exponential
Histogram once"]:::summary - h3 --> r3(["q1 … q5
readouts"]):::readout + h3 --> r3(["q1 … q5
estimates"]):::estimate end end subgraph A2["A2 · materialized at ingestion time"] @@ -671,14 +682,14 @@ flowchart LR s4[("each new sample
+ one-time backfill")]:::data --> h4["maintain Exponential
Histogram"]:::summary end subgraph A2Q["Query time, at T"] - r4(["q1 … q5
readouts"]):::readout + r4(["q1 … q5
estimates"]):::estimate end h4 --> r4 end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef readout fill:#e6f4ea,stroke:#188038,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` Which candidate wins depends on the workload: @@ -705,20 +716,20 @@ flowchart LR s1[("samples")]:::data --> p1["1-min KLL pane
stored 5 min"]:::summary end subgraph B1Q["Query time, every 1 min"] - g1["merge latest
5 panes"]:::exact --> r1(["p99"]):::readout + g1["merge latest
5 panes"]:::exact --> r1(["p99"]):::estimate end p1 --> g1 end subgraph B2["B2 · not materialized"] direction LR subgraph B2Q["Query time, every 1 min"] - s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::readout + s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::estimate end end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef readout fill:#e6f4ea,stroke:#188038,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` The query repeats every minute and the data is continuously ingesting, so B1 From ada1b89a99d694fd79be9719b62a6b91451a1262 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:13:40 -0400 Subject: [PATCH 35/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 138f6baf..bb977f7b 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -167,6 +167,7 @@ A summary-based candidate uses three kinds of summary nodes: into a coarser one. * A **summary estimation node** computes an answer from a summary, for example the p99 estimate from a KLL, or the entropy estimate from a UnivMon. +* **summary subtract node** and **summary delete node** design is TODO. One summary build node can feed several estimation nodes, which is what Pass 2 exploits. From 659fd24da95eba54797a3603a86fb590db4fd0ea Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:42:44 -0400 Subject: [PATCH 36/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 7 ++----- 1 file changed, 2 insertions(+), 5 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index bb977f7b..081b7a24 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -131,11 +131,8 @@ evaluation timing and missing-data semantics. Logical optimization runs in two passes. Pass 1 generates candidates for each computation on its own; Pass 2 finds candidates that share computation across -sub-DAGs and queries. Neither pass decides materialization or execution -placement, and neither picks one summary per computation: every candidate is -kept for selection. Stage 1 reads the repetition interval only to detect -windows that overlap across evaluations; deciding when anything is computed is -left to stage 2. +sub-DAGs and queries. Decisions about materialization, execution +placement are taken in later stages. #### Pass 1: Local candidate generation From 0fff9b785b34a7ce28b8412cdd1009b18abf707a Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:47:34 -0400 Subject: [PATCH 37/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 081b7a24..c0b6fccf 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -263,8 +263,7 @@ constraints: * A materialized output is stored for as long as any of its consumers still needs it. -* Query-time work cannot feed ingestion-time work, so every node feeding an - ingestion-time sub-DAG also runs at ingestion time. +* Every node upstream of an ingestion-time node also runs at ingestion time. The decision depends on the workload's `recurrence` and `predictability` and on the `DataWorkload`. Typical outcomes: From cd10608973860bee652e4034edbd2f8be9221280 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:49:25 -0400 Subject: [PATCH 38/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index c0b6fccf..2a0a72c8 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -242,7 +242,7 @@ and resource management are TODO. #### Materialization -Materialization decides, for each sub-DAG, whether its output is kept across +Materialization decides, for each sub-DAG, whether its output is kept (persistent to disk or kept in memory) across (batch) query executions, and if so, when it is computed and how long it is stored. Materialization does not imply ingestion time; a sub-DAG has three options: From e0653823c8b7a1ef547eaf7863d2a5e75d9b331e Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:54:38 -0400 Subject: [PATCH 39/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 7 +------ 1 file changed, 1 insertion(+), 6 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 2a0a72c8..1b4b2f8c 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -185,12 +185,7 @@ separately. The window-composition rule shares a summary across windows by splitting time into pieces that each carry their own summary: -* A **pane** is a fixed-length, non-overlapping slice of time, for example one - minute, with one summary of the data that arrived in that slice. A query - window is answered by merging the summaries of the panes it covers. The pane - length is chosen so that every requested window is an exact union of panes: - for a 5-min window evaluated every 1 min, 1-min panes work, since each window - is exactly 5 consecutive panes. +* **Panes** split the time axis into consecutive fixed-length slices (for example, 1 min each); no two panes overlap. Each pane holds one summary of the data arriving in it. A window is answered by merging the summaries of its panes. The pane length must divide both the window length and the evaluation interval, so that every window is an exact run of panes: a 5-min window evaluated every 1 min uses 1-min panes, and each window is 5 consecutive panes. The windows overlap, not the panes: in a sliding window, consecutive windows share most of their panes (here 4 of 5), which is why one set of panes can serve every evaluation. * A **bucket** of an Exponential Histogram plays the same role, but bucket lengths grow with age: recent data sits in short buckets and older data in longer ones. This keeps few buckets over a long history, at the cost that old From 46c101472a1ed38e4af75415b61bed6fe994032c Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:59:02 -0400 Subject: [PATCH 40/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 34 +++++++++++++------ 1 file changed, 24 insertions(+), 10 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 1b4b2f8c..9bd568c0 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -112,17 +112,20 @@ query's accuracy target), and every rejected candidate carries a reason. | Stage | Input | Decides | Output | |---|---|---|---| -| 0. Frontends | Each `QueryWorkloadEntry.query` and `QueryWorkload.language` | Language semantics, and converting source-language queries into a common logical representation. Rejects constructs it cannot represent faithfully. | `CandidateLogicalDAGs`; nodes are logical operations with no summary operations | -| 1. Logical ASAP-aware optimization | Logical DAGs, and each query's accuracy requirement, `time_selection` and repetition interval (from `recurrence`) across the workload | **Pass 1 — Summary replacement and query rewriting:** apply query rewriting rules to each eligible sub-DAG and generate exact and summary-based candidates that satisfy its semantics and accuracy requirements. **Pass 2 — ASAP-aware CSE:** apply traditional CSE and summary-specific CSE rules across sub-DAGs and queries to generate shared computation candidates, while preserving independent candidates. | `CandidateLogicalASAPDAGs`; nodes include summary and window-summary operations | -| 2. Physical ASAP-aware optimization | Logical ASAP DAGs, each entry's `recurrence` and `predictability`, the `DataWorkload` | **Materialization:** for each sub-DAG, whether its output is materialized, when it is computed (ingestion time or query time), and how long it is retained. **Physical operator implementation:** physical operators for every node. **Parallelism, partitioning, resources:** TODO. | `CandidatePhysicalASAPDAGs` | -| 3. Plan selection | Physical ASAP DAG candidates, `requirements`, the deployment's cost model, accuracy model and capabilities | Rejects candidates that miss an accuracy target, a latency bound or a capability; picks the cheapest valid plan for the whole workload. A shared state is costed once with all its consumers' demand. | One `PhysicalASAPDAG` | -| 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Executes ingestion, storage, precomputation and query-time computation. | Query results | +| 0. Frontends | `query`, `language` | Parse and lower to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | +| 1. Logical ASAP-aware optimization | Logical DAGs; accuracy, `time_selection`, repetition interval | Summary replacement (Pass 1); ASAP-aware CSE (Pass 2) | `CandidateLogicalASAPDAGs` | +| 2. Physical ASAP-aware optimization | Logical ASAP DAGs; `recurrence`, `predictability`, `DataWorkload` | Materialization; physical operators; parallelism and resources (TODO) | `CandidatePhysicalASAPDAGs` | +| 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG` | +| 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Run ingestion, storage and query-time computation | Query results | + +The sections below describe each stage. ### 0. Language-specific frontends The frontend converts each query into a `LogicalDAG`. Nodes represent logical query operations, including selectors, transformations, aggregations, grouping -and window semantics. They contain no ASAP summary choices. +and window semantics. They contain no ASAP summary choices. A construct that cannot be represented +faithfully is rejected. The frontend preserves source-language behavior, including series identity, evaluation timing and missing-data semantics. @@ -185,7 +188,15 @@ separately. The window-composition rule shares a summary across windows by splitting time into pieces that each carry their own summary: -* **Panes** split the time axis into consecutive fixed-length slices (for example, 1 min each); no two panes overlap. Each pane holds one summary of the data arriving in it. A window is answered by merging the summaries of its panes. The pane length must divide both the window length and the evaluation interval, so that every window is an exact run of panes: a 5-min window evaluated every 1 min uses 1-min panes, and each window is 5 consecutive panes. The windows overlap, not the panes: in a sliding window, consecutive windows share most of their panes (here 4 of 5), which is why one set of panes can serve every evaluation. +* **Panes** split the time axis into consecutive fixed-length slices (for + example, 1 min each); no two panes overlap. Each pane holds one summary of + the data arriving in it. A window is answered by merging the summaries of + its panes. The pane length must divide both the window length and the + evaluation interval, so that every window is an exact run of panes: a 5-min + window evaluated every 1 min uses 1-min panes, and each window is 5 + consecutive panes. The *windows* overlap, not the panes: in a **sliding + window**, consecutive windows share most of their panes (here 4 of 5), which + is why one set of panes can serve every evaluation. * A **bucket** of an Exponential Histogram plays the same role, but bucket lengths grow with age: recent data sits in short buckets and older data in longer ones. This keeps few buckets over a long history, at the cost that old @@ -279,11 +290,14 @@ merge and quantile estimation operators. ### 3. Plan selection -Selection is the only stage that uses the deployment's cost and accuracy -models, and the only stage that discards valid candidates. Accuracy is +Selection rejects every candidate that misses an accuracy target or a latency +bound, or that needs a capability the deployment lacks, and then picks the +cheapest remaining plan. It is the only stage that uses the deployment's cost +and accuracy models, and the only stage that discards valid candidates. Accuracy is estimated by the deployment's accuracy model, not assumed from a summary's nominal bound. Cost is evaluated for the whole workload rather than per query, -which is what lets one shared summary beat several cheaper independent ones. +which is what lets one shared summary beat several cheaper independent ones: +a shared summary is costed once, with the demand of all its consumers. ### 4. Execution From 189fabb6d2e980b954d25a7a8c356e13adfa6d8f Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 14:19:17 -0400 Subject: [PATCH 41/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 9bd568c0..6b210771 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -413,7 +413,7 @@ flowchart LR **Stage 3.** The deployment's accuracy model checks that each summary candidate meets ε = 0.01, δ = 0.001, and its cost model compares per-`job` sketches with -one Hydra sketch. With many small jobs, one shared Hydra sketch is typically +one Hydra sketch. With many small jobs, one shared Hydra sketch can be cheaper; with a few large jobs, per-`job` Count-Min sketches may win. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 From de5e8867f007b911ffc71ed79be9e33004be807c Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 14:43:53 -0400 Subject: [PATCH 42/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 133 ++++++++++++++---- 1 file changed, 103 insertions(+), 30 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 6b210771..d5c681aa 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -334,15 +334,14 @@ Unless an example says otherwise, every example uses this data workload: | `DataWorkload` field | Meaning | Value | |---|---|---| | `arrival` | Whether the data is at rest, still arriving, or both | `continuously_ingesting` | -| `data_ingestion_interval` | How often each series delivers one sample (the scrape interval in Prometheus). PromQL uses it as the look-back horizon of instant selectors. | 15 s (declared) | +| `data_ingestion_interval` | How often each series delivers one sample (the scrape interval in Prometheus). PromQL uses it as the look-back horizon of instant selectors. | 15 s | | `ingestion_volume` | Total amount of ingested data | unknown | -| `ingestion_rate` | Samples arriving per second across all series | about 66,667 samples/s (declared) | -| `input_cardinality` | Number of distinct series (or keys) | 1,000,000 series (declared) | -| `distribution` | How samples are spread over keys | `zipf` (declared) | +| `ingestion_rate` | Samples arriving per second across all series | about 66,667 samples/s | +| `input_cardinality` | Number of distinct series (or keys) | 1,000,000 series | +| `distribution` | How samples are spread over keys | `zipf` | -"Declared" is the value's `EvidenceSource`: the workload author stated it -rather than the planner observing it. With 1,000,000 series each sampled every -15 s, the ingestion rate is 1,000,000 / 15 ≈ 66,667 samples/s. +With 1,000,000 series each sampled every 15 s, the ingestion rate is +1,000,000 / 15 ≈ 66,667 samples/s. ### Example 1: Aggregation over dimensions — summary replacement in Pass 1 @@ -357,21 +356,37 @@ first needs an exact total; the second tolerates error. | `predictability` | `predictable` | | `time_selection.scope` | `real_time` | -| Query | `lookback` | `as_of` | Accuracy | Latency | +| Query | `lookback` | `as_of` | Accuracy requirement | Latency requirement | |---|---|---|---|---| | `sum by (job) (rate(http_requests_total[1m]))` | 1 m | evaluation time | exact (`implicit_exact`) | unspecified | | `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | -**Stage 0.** The two `LogicalDAG`s are -`range http_requests_total[1m] → rate → sum by (job)` and -`range http_requests_total[1m] → sum_over_time → topk by (job) (10)`. +**Stage 0.** The frontend produces one `LogicalDAG` per query. Neither +contains a summary: + +```mermaid +flowchart LR + subgraph Q1["Q1 · sum by (job) (rate(http_requests_total[1m]))"] + direction LR + x1[("http_requests_total")]:::data --> x2["range 1m"]:::exact --> x3["rate"]:::exact --> x4["sum by (job)"]:::exact + end + subgraph Q2["Q2 · topk by (job) (10, sum_over_time(http_requests_total[1m]))"] + direction LR + y1[("http_requests_total")]:::data --> y2["range 1m"]:::exact --> y3["sum_over_time"]:::exact --> y4["topk by (job) (10)"]:::exact + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; +``` **Stage 1, Pass 1.** -* `sum by (job) (rate(...))`: the accuracy requirement is `implicit_exact`, and - an exact grouped sum already keeps one value per `job`. The only candidate is - a per-series Rate feeding an exact per-`job` Sum. No summary helps here. -* `topk by (job) (10, ...)`: the exact candidate keeps a per-series sum and +* Q1, `sum by (job) (rate(...))`: the accuracy requirement is `implicit_exact`, + and an exact grouped sum already keeps one value per `job`. The only + candidate is a per-series Rate feeding an exact per-`job` Sum. No summary + helps here. +* Q2, `topk by (job) (10, ...)`: the exact candidate keeps a per-series sum and sorts within each `job`, which is costly at one million Zipf-distributed series. The `EpsilonDelta` target admits two summary candidates: 1. **Count-Min Sketch with a top-*k* heap per `job`.** One sketch per group; @@ -379,16 +394,12 @@ first needs an exact total; the second tolerates error. 2. **Hydra over the whole `job` column.** One sketch covers every (`job`, series) key and answers the top 10 for any `job`. - Pass 1 keeps all three candidates. The better summary depends on the number - of jobs and on costs that only the deployment knows. + Pass 1 keeps all three candidates for Q2. The better summary depends on the + number of jobs and on costs that only the deployment knows. ```mermaid flowchart LR - subgraph L["Stage 0 · LogicalDAG"] - direction LR - a1[("http_requests_total
last 1m")]:::data --> a2["sum_over_time"]:::exact --> a3["topk by (job) (10)"]:::exact - end - subgraph C["Stage 1, Pass 1 · CandidateLogicalASAPDAGs"] + subgraph C["Stage 1, Pass 1 · candidates for Q2"] direction TB subgraph E["Exact"] direction LR @@ -403,14 +414,76 @@ flowchart LR h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::estimate end end - S{{"Stage 3 · cheapest valid candidate
many small jobs → Hydra
few large jobs → Count-Min per job"}} - L --> C --> S classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` +**Stage 1, Pass 2.** Both queries read the same range selector, +`http_requests_total[1m]`, so the identical-expression rule adds a candidate +in which Q1 and Q2 share one input node. No summary is shared: Q1 needs an +exact Rate and Sum, and no summary supports both Q1 and Q2. Both queries also +evaluate a 1-min window every 10 s, so consecutive windows overlap; the +window-composition rule could add 10-s panes for the mergeable candidates. +Example 3 walks through that rule, so it is not expanded here. + +```mermaid +flowchart LR + in[("http_requests_total")]:::data --> r["range 1m
one shared input node"]:::exact + subgraph Q1P["Q1 · exact candidate"] + direction LR + q1a["rate
per series"]:::exact --> q1b["sum by (job)"]:::exact + end + subgraph Q2P["Q2 · a summary candidate from Pass 1"] + direction LR + q2a["Count-Min + heap per job
or Hydra"]:::summary --> q2b(["top 10
per job"]):::estimate + end + r --> q1a + r --> q2a + pn["optional: 10-s panes
window-composition rule, Example 3"]:::note -.-> q2a + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; + classDef note fill:#fff,stroke:#9aa0a6,stroke-dasharray:4 3,color:#5f6368; +``` + +**Stage 2.** Both panels repeat every 10 s over continuously arriving data, so +stage 2 generates candidates that materialize each Q2 summary at ingestion time +next to candidates that build it at query time. The exact Q1 path is lowered to +a per-series rate and a grouped sum. Example 4 shows this materialization +decision in detail. + +```mermaid +flowchart LR + subgraph C1["Candidate 1 · Q2 summary materialized at ingestion time"] + direction LR + subgraph C1I["Ingestion time"] + s1[("samples")]:::data --> b1["maintain Q2 summary
as data arrives"]:::summary + end + subgraph C1Q["Query time, every 10 s"] + e1(["top 10
per job"]):::estimate + end + b1 --> e1 + end + subgraph C2["Candidate 2 · Q2 summary built at query time"] + direction LR + subgraph C2Q["Query time, every 10 s"] + s2[("last 1 min of
raw samples")]:::data --> b2["build Q2 summary"]:::summary --> e2(["top 10
per job"]):::estimate + end + end + subgraph Q1L["Q1 in both candidates · lowered exact path"] + direction LR + s3[("samples")]:::data --> r3["per-series rate"]:::exact --> g3["grouped sum
by (job)"]:::exact + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; + classDef note fill:#fff,stroke:#9aa0a6,stroke-dasharray:4 3,color:#5f6368; +``` + **Stage 3.** The deployment's accuracy model checks that each summary candidate meets ε = 0.01, δ = 0.001, and its cost model compares per-`job` sketches with one Hydra sketch. With many small jobs, one shared Hydra sketch can be @@ -429,7 +502,7 @@ source IPs over the last minute. | `predictability` | `predictable` | | `time_selection.scope` | `real_time` | | `as_of` | evaluation time | -| Latency | unspecified | +| Latency requirement | unspecified | ```sql -- Q1: Distinct(src_ip) @@ -456,7 +529,7 @@ FROM ( ); ``` -| Query | Computation | `lookback` | Accuracy | +| Query | Computation | `lookback` | Accuracy requirement | |---|---|---|---| | Q1 | `Distinct(src_ip)` | 1 m | ε = 0.02, δ = 0.01 | | Q2 | `Entropy(src_ip)` | 1 m | ε = 0.05, δ = 0.01 | @@ -466,7 +539,7 @@ The data workload differs from the shared one in two fields: | `DataWorkload` field | Value | |---|---| -| `input_cardinality` | 10,000,000 distinct source IPs (declared) | +| `input_cardinality` | 10,000,000 distinct source IPs | | `data_ingestion_interval` | not needed for SQL | **Pass 1.** Rewrite rules recognize the three computations, and each gets its @@ -562,8 +635,8 @@ all executed together at time T. | `execute_at` | T | | `predictability` | `ad_hoc` | | `time_selection.scope` | `longitudinal` | -| Accuracy | ε = 0.005, δ = 0.01 | -| Latency | unspecified | +| Accuracy requirement | ε = 0.005, δ = 0.01 | +| Latency requirement | unspecified | | Query | `lookback` | `as_of` | |---|---|---| @@ -627,7 +700,7 @@ over the last 5 min, refreshed every minute. | `predictability` | `predictable` | | `time_selection.scope` | `real_time` | -| Query | `lookback` | `as_of` | Accuracy | Latency | +| Query | `lookback` | `as_of` | Accuracy requirement | Latency requirement | |---|---|---|---|---| | `quantile_over_time(0.99, latency_ms[5m])` | 5 m | evaluation time | ε = 0.01, δ = 0.01 | ≤ 200 ms | From 289838f42bffad8f9dbec7bfe6e2ae0f41340acc Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 14:48:42 -0400 Subject: [PATCH 43/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 10 ++++++---- 1 file changed, 6 insertions(+), 4 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index d5c681aa..da798f12 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -451,8 +451,10 @@ flowchart LR **Stage 2.** Both panels repeat every 10 s over continuously arriving data, so stage 2 generates candidates that materialize each Q2 summary at ingestion time -next to candidates that build it at query time. The exact Q1 path is lowered to -a per-series rate and a grouped sum. Example 4 shows this materialization +next to candidates that do not materialize it and rebuild it from raw samples +at every evaluation. Q1 has only its exact +candidate, so both candidates compute it the same way: a per-series rate +followed by a grouped sum. Example 4 shows this materialization decision in detail. ```mermaid @@ -467,13 +469,13 @@ flowchart LR end b1 --> e1 end - subgraph C2["Candidate 2 · Q2 summary built at query time"] + subgraph C2["Candidate 2 · Q2 summary not materialized, rebuilt every evaluation"] direction LR subgraph C2Q["Query time, every 10 s"] s2[("last 1 min of
raw samples")]:::data --> b2["build Q2 summary"]:::summary --> e2(["top 10
per job"]):::estimate end end - subgraph Q1L["Q1 in both candidates · lowered exact path"] + subgraph Q1L["Q1 · exact, computed the same way in Candidates 1 and 2"] direction LR s3[("samples")]:::data --> r3["per-series rate"]:::exact --> g3["grouped sum
by (job)"]:::exact end From c31782d126966479acaadca60245fb65eb4e7834 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:00:04 -0400 Subject: [PATCH 44/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 100 ++++++++++++------ 1 file changed, 67 insertions(+), 33 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index da798f12..98101162 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -185,18 +185,22 @@ or value being summarized together with its grouping. The summary input data doe not include the window; the window-composition rule compares windows separately. -The window-composition rule shares a summary across windows by splitting time -into pieces that each carry their own summary: - -* **Panes** split the time axis into consecutive fixed-length slices (for - example, 1 min each); no two panes overlap. Each pane holds one summary of - the data arriving in it. A window is answered by merging the summaries of - its panes. The pane length must divide both the window length and the - evaluation interval, so that every window is an exact run of panes: a 5-min - window evaluated every 1 min uses 1-min panes, and each window is 5 - consecutive panes. The *windows* overlap, not the panes: in a **sliding - window**, consecutive windows share most of their panes (here 4 of 5), which - is why one set of panes can serve every evaluation. +The window-composition rule distinguishes what a query reads from what the +planner stores: + +* A **window** is the time range one query evaluation reads, for example the + last 5 min. Consecutive evaluations of a repeating query read overlapping + windows; this is a **sliding window**. +* A **pane** is what the planner summarizes and stores. The time axis is cut + into back-to-back panes of equal length (a tumbling window), each with one + summary. A window is answered by merging the summaries of the panes it + covers, so overlapping windows reuse the same panes instead of each building + its own summary. Most summaries cannot remove old data, so one summary cannot + simply slide forward; panes avoid that, because the oldest pane is dropped + rather than subtracted. The pane length must divide both the window length + and the evaluation interval, so that every window starts and ends on a pane + boundary and never needs part of a pane: a 5-min window evaluated every + 1 min uses 1-min panes, and each window merges exactly 5 whole panes. * A **bucket** of an Exponential Histogram plays the same role, but bucket lengths grow with age: recent data sits in short buckets and older data in longer ones. This keeps few buckets over a long history, at the cost that old @@ -441,51 +445,81 @@ flowchart LR end r --> q1a r --> q2a - pn["optional: 10-s panes
window-composition rule, Example 3"]:::note -.-> q2a classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; - classDef note fill:#fff,stroke:#9aa0a6,stroke-dasharray:4 3,color:#5f6368; ``` -**Stage 2.** Both panels repeat every 10 s over continuously arriving data, so -stage 2 generates candidates that materialize each Q2 summary at ingestion time -next to candidates that do not materialize it and rebuild it from raw samples -at every evaluation. Q1 has only its exact -candidate, so both candidates compute it the same way: a per-series rate -followed by a grouped sum. Example 4 shows this materialization -decision in detail. +**Stage 2.** Stage 2 decides, for each query, whether to materialize its +computation at ingestion time or to recompute it from raw samples at every +refresh. Each query has its own two options: ```mermaid flowchart LR - subgraph C1["Candidate 1 · Q2 summary materialized at ingestion time"] + in[("http_requests_total
samples")]:::data + + subgraph Q1M["Q1 · materialized"] direction LR - subgraph C1I["Ingestion time"] - s1[("samples")]:::data --> b1["maintain Q2 summary
as data arrives"]:::summary + subgraph Q1MI["Ingestion time"] + m1["maintain per-series rate
and sum by (job)"]:::exact end - subgraph C1Q["Query time, every 10 s"] - e1(["top 10
per job"]):::estimate + subgraph Q1MQ["Query time, every 10 s"] + m1r["read Q1 result"]:::exact end - b1 --> e1 + m1 --> m1r end - subgraph C2["Candidate 2 · Q2 summary not materialized, rebuilt every evaluation"] + + subgraph Q1N["Q1 · not materialized"] direction LR - subgraph C2Q["Query time, every 10 s"] - s2[("last 1 min of
raw samples")]:::data --> b2["build Q2 summary"]:::summary --> e2(["top 10
per job"]):::estimate + subgraph Q1NQ["Query time, every 10 s"] + n1a["range 1m"]:::exact --> n1b["rate
per series"]:::exact --> n1c["sum by (job)"]:::exact end end - subgraph Q1L["Q1 · exact, computed the same way in Candidates 1 and 2"] + + subgraph Q2M["Q2 · materialized"] direction LR - s3[("samples")]:::data --> r3["per-series rate"]:::exact --> g3["grouped sum
by (job)"]:::exact + subgraph Q2MI["Ingestion time"] + m2["maintain Q2 summary
Count-Min + heap or Hydra"]:::summary + end + subgraph Q2MQ["Query time, every 10 s"] + m2e(["top 10
per job"]):::estimate + end + m2 --> m2e end + + subgraph Q2N["Q2 · not materialized"] + direction LR + subgraph Q2NQ["Query time, every 10 s"] + n2a["range 1m"]:::exact --> n2b["build Q2 summary
Count-Min + heap or Hydra"]:::summary --> n2c(["top 10
per job"]):::estimate + end + end + + in --> m1 + in --> n1a + in --> m2 + in --> n2a classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; - classDef note fill:#fff,stroke:#9aa0a6,stroke-dasharray:4 3,color:#5f6368; ``` +Combining the two options per query gives four candidates: + +| Candidate | Q1 | Q2 | Shared input node from Pass 2 | +|---|---|---|---| +| 1 | materialized | materialized | not used: neither query reads the range at query time | +| 2 | materialized | not materialized | not used: only Q2 reads the range | +| 3 | not materialized | materialized | not used: only Q1 reads the range | +| 4 | not materialized | not materialized | used: both read the same last minute of samples | + +Both panels refresh every 10 s over arriving data, so candidates that +materialize usually cost less per refresh. Candidate 4 wins only when storage +is expensive and raw data is available at query time; it is also the only +candidate where the shared input node saves work. Example 4 covers +materialization in more detail, including materializing at query time. + **Stage 3.** The deployment's accuracy model checks that each summary candidate meets ε = 0.01, δ = 0.001, and its cost model compares per-`job` sketches with one Hydra sketch. With many small jobs, one shared Hydra sketch can be From 749723ee7b2da152e06ebed8d261b1f7fbe3cdc9 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:05:17 -0400 Subject: [PATCH 45/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 164 +++++++++++------- 1 file changed, 101 insertions(+), 63 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 98101162..21f5bf82 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -110,6 +110,11 @@ be shared or enumerated lazily. A stage may prune a candidate early only when it is provably invalid (for example, a summary family that cannot meet the query's accuracy target), and every rejected candidate carries a reason. +A candidate is a DAG for the **whole workload**, not for one query. Each stage +combines its choices for every sub-DAG with the candidates it receives (the × +in the diagram), so the candidate set grows from stage to stage until +selection picks one. Example 1 traces this growth step by step. + | Stage | Input | Decides | Output | |---|---|---|---| | 0. Frontends | `query`, `language` | Parse and lower to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | @@ -322,7 +327,7 @@ is at most ε with probability at least 1 − δ. | Example | Shows | |---|---| -| 1. Aggregation over dimensions | Pass 1: summary replacement | +| 1. Aggregation over dimensions | How the candidate set grows through every stage | | 2. One summary for several computations | Pass 2: summary-capability rule | | 3. Aggregation over windows | Pass 1 and Pass 2: window-composition rule | | 4. Materialization of window summaries | Stage 2: materialization, decoupled from stage 1 | @@ -347,7 +352,7 @@ Unless an example says otherwise, every example uses this data workload: With 1,000,000 series each sampled every 15 s, the ingestion rate is 1,000,000 / 15 ≈ 66,667 samples/s. -### Example 1: Aggregation over dimensions — summary replacement in Pass 1 +### Example 1: Aggregation over dimensions — the candidate set through every stage **Query workload.** Two PromQL dashboard panels over the last minute. The first needs an exact total; the second tolerates error. @@ -365,8 +370,27 @@ first needs an exact total; the second tolerates error. | `sum by (job) (rate(http_requests_total[1m]))` | 1 m | evaluation time | exact (`implicit_exact`) | unspecified | | `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | -**Stage 0.** The frontend produces one `LogicalDAG` per query. Neither -contains a summary: +This example follows the workload's candidate set through every stage. Each +candidate covers both queries. To keep the counts small, it leaves out the +pane candidates that the window-composition rule would also add (Example 3 +covers them). + +```mermaid +flowchart TB + s0["Stage 0 · 1 candidate
the workload's LogicalDAG"] + p1["Stage 1, Pass 1 · 3 candidates
Q2 exact, Count-Min + heap, or Hydra"] + p2["Stage 1, Pass 2 · 6 candidates
each with separate or shared input"] + s2["Stage 2 · 15 candidates
Q1 and Q2 each materialized or not"] + s3(["Stage 3 · 1 plan"]):::estimate + s0 -- "× summary choices" --> p1 -- "× sharing choices" --> p2 -- "× materialization choices" --> s2 -- "select" --> s3 + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; +``` + +**Stage 0: 1 candidate.** The frontend lowers both queries into one workload +`LogicalDAG` with no summaries: ```mermaid flowchart LR @@ -384,26 +408,20 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Stage 1, Pass 1.** - -* Q1, `sum by (job) (rate(...))`: the accuracy requirement is `implicit_exact`, - and an exact grouped sum already keeps one value per `job`. The only - candidate is a per-series Rate feeding an exact per-`job` Sum. No summary - helps here. -* Q2, `topk by (job) (10, ...)`: the exact candidate keeps a per-series sum and - sorts within each `job`, which is costly at one million Zipf-distributed - series. The `EpsilonDelta` target admits two summary candidates: - 1. **Count-Min Sketch with a top-*k* heap per `job`.** One sketch per group; - each answers its own top 10. - 2. **Hydra over the whole `job` column.** One sketch covers every (`job`, - series) key and answers the top 10 for any `job`. +**Stage 1, Pass 1: 3 candidates.** Pass 1 finds local options for each query: - Pass 1 keeps all three candidates for Q2. The better summary depends on the - number of jobs and on costs that only the deployment knows. +* **Q1** has one option, the exact per-series rate and per-`job` sum. Its + accuracy requirement is `implicit_exact`, and an exact grouped sum is + already small. +* **Q2** has three options. The exact one keeps a per-series sum and sorts + within each `job`, which is costly at one million Zipf-distributed series. + The `EpsilonDelta` target also admits a **Count-Min Sketch with a top-*k* + heap per `job`**, and **Hydra over the whole `job` column**, where one sketch + covers every (`job`, series) key. ```mermaid flowchart LR - subgraph C["Stage 1, Pass 1 · candidates for Q2"] + subgraph C["Q2's three local options"] direction TB subgraph E["Exact"] direction LR @@ -424,36 +442,42 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Stage 1, Pass 2.** Both queries read the same range selector, -`http_requests_total[1m]`, so the identical-expression rule adds a candidate -in which Q1 and Q2 share one input node. No summary is shared: Q1 needs an -exact Rate and Sum, and no summary supports both Q1 and Q2. Both queries also -evaluate a 1-min window every 10 s, so consecutive windows overlap; the -window-composition rule could add 10-s panes for the mergeable candidates. -Example 3 walks through that rule, so it is not expanded here. +Combining them gives 1 × 3 = 3 workload candidates: + +| Candidate | Q1 | Q2 | +|---|---|---| +| L1 | exact | exact | +| L2 | exact | Count-Min + heap | +| L3 | exact | Hydra | + +**Stage 1, Pass 2: 6 candidates.** Both queries read the same range selector, +`http_requests_total[1m]`, so the identical-expression rule adds, for each of +L1–L3, a variant in which Q1 and Q2 share one input node: L1s, L2s and L3s. The +originals are kept. No summary is shared, because Q1 must be exact and no +summary supports both queries. ```mermaid flowchart LR - in[("http_requests_total")]:::data --> r["range 1m
one shared input node"]:::exact - subgraph Q1P["Q1 · exact candidate"] + subgraph SEP["L3 · separate inputs"] direction LR - q1a["rate
per series"]:::exact --> q1b["sum by (job)"]:::exact + a1[("http_requests_total")]:::data --> a2["range 1m"]:::exact --> a3["rate → sum by (job)"]:::exact + b1[("http_requests_total")]:::data --> b2["range 1m"]:::exact --> b3["Hydra"]:::summary --> b4(["top 10"]):::estimate end - subgraph Q2P["Q2 · a summary candidate from Pass 1"] + subgraph SH["L3s · shared input"] direction LR - q2a["Count-Min + heap per job
or Hydra"]:::summary --> q2b(["top 10
per job"]):::estimate + c1[("http_requests_total")]:::data --> c2["range 1m
one shared input node"]:::exact + c2 --> c3["rate → sum by (job)"]:::exact + c2 --> c4["Hydra"]:::summary --> c5(["top 10"]):::estimate end - r --> q1a - r --> q2a classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Stage 2.** Stage 2 decides, for each query, whether to materialize its -computation at ingestion time or to recompute it from raw samples at every -refresh. Each query has its own two options: +**Stage 2: 15 candidates.** For every logical candidate, stage 2 decides +separately for Q1 and Q2 whether to materialize the computation at ingestion +time or recompute it from raw samples at every refresh: ```mermaid flowchart LR @@ -480,7 +504,7 @@ flowchart LR subgraph Q2M["Q2 · materialized"] direction LR subgraph Q2MI["Ingestion time"] - m2["maintain Q2 summary
Count-Min + heap or Hydra"]:::summary + m2["maintain Q2's summary
or exact state"]:::summary end subgraph Q2MQ["Query time, every 10 s"] m2e(["top 10
per job"]):::estimate @@ -491,7 +515,7 @@ flowchart LR subgraph Q2N["Q2 · not materialized"] direction LR subgraph Q2NQ["Query time, every 10 s"] - n2a["range 1m"]:::exact --> n2b["build Q2 summary
Count-Min + heap or Hydra"]:::summary --> n2c(["top 10
per job"]):::estimate + n2a["range 1m"]:::exact --> n2b["build Q2's summary
or exact state"]:::summary --> n2c(["top 10
per job"]):::estimate end end @@ -505,25 +529,27 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -Combining the two options per query gives four candidates: - -| Candidate | Q1 | Q2 | Shared input node from Pass 2 | -|---|---|---|---| -| 1 | materialized | materialized | not used: neither query reads the range at query time | -| 2 | materialized | not materialized | not used: only Q2 reads the range | -| 3 | not materialized | materialized | not used: only Q1 reads the range | -| 4 | not materialized | not materialized | used: both read the same last minute of samples | - -Both panels refresh every 10 s over arriving data, so candidates that -materialize usually cost less per refresh. Candidate 4 wins only when storage -is expensive and raw data is available at query time; it is also the only -candidate where the shared input node saves work. Example 4 covers -materialization in more detail, including materializing at query time. - -**Stage 3.** The deployment's accuracy model checks that each summary candidate -meets ε = 0.01, δ = 0.001, and its cost model compares per-`job` sketches with -one Hydra sketch. With many small jobs, one shared Hydra sketch can be -cheaper; with a few large jobs, per-`job` Count-Min sketches may win. +That gives 2 × 2 = 4 materialization choices per logical candidate. A shared +input node only matters when both queries read raw samples at query time; in +the other three choices, the shared-input variant is the same plan as its +original. So each Q2 option yields 5 distinct physical candidates, 15 in all: + +| Q2 option (logical candidates) | Q1 mat., Q2 mat. | Q1 mat., Q2 not | Q1 not, Q2 mat. | Q1 not, Q2 not, separate inputs | Q1 not, Q2 not, shared input | +|---|---|---|---|---|---| +| Exact (L1, L1s) | P1 | P2 | P3 | P4 | P5 | +| Count-Min + heap (L2, L2s) | P6 | P7 | P8 | P9 | P10 | +| Hydra (L3, L3s) | P11 | P12 | P13 | P14 | P15 | + +**Stage 3: 1 plan.** Selection first rejects invalid candidates. The +deployment's accuracy model checks the summary candidates against ε = 0.01, +δ = 0.001, and its cost model estimates Q2's latency against the 100 ms bound; +for example, an exact top-k sorted over one million series at query time (P2, +P4, P5) may miss it. Among the rest, it picks the cheapest for the whole +workload. Because both panels refresh every 10 s over arriving data, a +candidate that materializes both queries usually wins: with many small jobs, +P11 (Hydra); with a few large jobs, P6 (Count-Min per `job`). The +not-materialized candidates win only when storage is expensive and raw data is +available at query time. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 @@ -580,13 +606,16 @@ The data workload differs from the shared one in two fields: **Pass 1.** Rewrite rules recognize the three computations, and each gets its local candidates from the Pass 1 table: exact, a specialized summary, or -UnivMon. +UnivMon. Combined, that is 3 × 3 × 3 = 27 workload candidates. **Pass 2.** All three computations have the same summary input data (`src_ip` from `flows`, no other filter) and the same 1-min window. UnivMon supports all three estimates, so the summary-capability rule adds a shared candidate: **one UnivMon build node feeding three estimation nodes**. It must be sized for the strictest -requirement, ε = 0.01. The independent candidates are kept as well. +requirement, ε = 0.01. The independent candidates are kept as well. Pass 2 +also adds candidates where only two of the three share a UnivMon and the third +keeps any of its own 3 options (3 pairs × 3 = 9), so stage 1 outputs +27 + 1 + 9 = 37 candidates. The figure shows the all-three case. ```mermaid flowchart LR @@ -691,7 +720,13 @@ of data at rest, plus data still arriving. mergeable. The window-composition rule adds a shared candidate: **one Exponential Histogram of KLL buckets over [T − 5 y, T]**, with one merge and estimation node per query that merges the buckets covering - its interval. The five independent candidates are kept. + its interval. The independent candidates are kept. + +Counted at the workload level, Pass 1 gives each query 2 options (exact or +KLL), so 2⁵ = 32 candidates. Pass 2 adds a candidate for every subset of two +or more queries that shares one Exponential Histogram while the others keep +their own options, 131 more, for 163 in total. Counts like this are why +candidate sets may be enumerated lazily. The five query intervals overlap, and all lie inside the last five years: @@ -743,7 +778,8 @@ over the last 5 min, refreshed every minute. * **Pass 1.** One KLL over 5 min for each evaluation. * **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The window-composition rule adds a shared candidate: **a sliding window of - 1-min KLL panes**, where each evaluation merges the latest 5 panes. + 1-min KLL panes**, where each evaluation merges the latest 5 panes. With the + exact candidate, stage 1 outputs 3 candidates. Each evaluation reads five 1-min panes, and consecutive evaluations share four of them: @@ -772,7 +808,9 @@ Example 4 shows how stage 2 decides whether to store these window summaries. ### Example 4: Materialization of window summaries in physical planning Stage 2 takes the shared window summaries from Example 3 and decides whether to -materialize them. That choice is driven by the workload's `recurrence`, +materialize them. This example shows only the physical candidates of the +shared logical candidate; every other logical candidate from Example 3 gets +its own physical candidates the same way. That choice is driven by the workload's `recurrence`, `predictability` and `data_workload.arrival`. **Pattern A (sub-interval batch).** From a159181ef70bf7206f5f2c64a6bb81dee9a7c133 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:06:21 -0400 Subject: [PATCH 46/56] Update planner-layering.md --- docs/design_docs/proposals/planner-layering.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 21f5bf82..d3883d7e 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -113,7 +113,7 @@ query's accuracy target), and every rejected candidate carries a reason. A candidate is a DAG for the **whole workload**, not for one query. Each stage combines its choices for every sub-DAG with the candidates it receives (the × in the diagram), so the candidate set grows from stage to stage until -selection picks one. Example 1 traces this growth step by step. +selection picks one; in the implementation, each stage can early prune invalid candidates. Example 1 traces this growth step by step. | Stage | Input | Decides | Output | |---|---|---|---| From b6fa552ead571a1e46084751b87faabec07c3d20 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:27:48 -0400 Subject: [PATCH 47/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 277 +++++++++--------- 1 file changed, 141 insertions(+), 136 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index d3883d7e..f85773c1 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -11,11 +11,12 @@ examples for why the rules are needed. ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the deployment's inputs (TODO: define this data structure in a follow-up PR), and returns one optimal physical plan. It decides what is computed, how it is computed, and which plan is best. The deployment only supplies inputs and -executes the plan: it supplies its own empirical cost estimation, empirical accuracy estimation and capabilities of deployment but never does the query planning or plan selection. +executes the plan: it provides its empirical cost model, empirical accuracy +model and capabilities, but never plans queries or selects plans. ## Stages -`x` represents Cartesian product for enumerating and combining different optimization angles in planning. +In the diagram, × means the Cartesian product: each stage combines every option along one dimension with every option along the others. ```text Query workload @@ -113,12 +114,12 @@ query's accuracy target), and every rejected candidate carries a reason. A candidate is a DAG for the **whole workload**, not for one query. Each stage combines its choices for every sub-DAG with the candidates it receives (the × in the diagram), so the candidate set grows from stage to stage until -selection picks one; in the implementation, each stage can early prune invalid candidates. Example 1 traces this growth step by step. +selection picks one. Example 1 traces this growth step by step. | Stage | Input | Decides | Output | |---|---|---|---| | 0. Frontends | `query`, `language` | Parse and lower to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | -| 1. Logical ASAP-aware optimization | Logical DAGs; accuracy, `time_selection`, repetition interval | Summary replacement (Pass 1); ASAP-aware CSE (Pass 2) | `CandidateLogicalASAPDAGs` | +| 1. Logical ASAP-aware optimization | Logical DAGs; accuracy requirements, `time_selection`, repetition interval | Summary replacement (Pass 1); ASAP-aware CSE (Pass 2) | `CandidateLogicalASAPDAGs` | | 2. Physical ASAP-aware optimization | Logical ASAP DAGs; `recurrence`, `predictability`, `DataWorkload` | Materialization; physical operators; parallelism and resources (TODO) | `CandidatePhysicalASAPDAGs` | | 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG` | | 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Run ingestion, storage and query-time computation | Query results | @@ -139,8 +140,8 @@ evaluation timing and missing-data semantics. Logical optimization runs in two passes. Pass 1 generates candidates for each computation on its own; Pass 2 finds candidates that share computation across -sub-DAGs and queries. Decisions about materialization, execution -placement are taken in later stages. +sub-DAGs and queries. Materialization and execution +placement are decided in later stages. #### Pass 1: Local candidate generation @@ -357,37 +358,13 @@ With 1,000,000 series each sampled every 15 s, the ingestion rate is **Query workload.** Two PromQL dashboard panels over the last minute. The first needs an exact total; the second tolerates error. -| Workload field | Value (both queries) | -|---|---| -| `language` | `promql` | -| Entry type | `repeating_queries` | -| `demand` | every 10 s (`fixed_interval`) | -| `predictability` | `predictable` | -| `time_selection.scope` | `real_time` | - -| Query | `lookback` | `as_of` | Accuracy requirement | Latency requirement | -|---|---|---|---|---| -| `sum by (job) (rate(http_requests_total[1m]))` | 1 m | evaluation time | exact (`implicit_exact`) | unspecified | -| `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | +| Query | Repeats | `lookback` | `as_of` | Accuracy requirement | Latency requirement | +|---|---|---|---|---|---| +| Q1: `sum by (job) (rate(http_requests_total[1m]))` | every 10 s | 1 m | evaluation time | exact | none | +| Q2: `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | every 10 s | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | This example follows the workload's candidate set through every stage. Each -candidate covers both queries. To keep the counts small, it leaves out the -pane candidates that the window-composition rule would also add (Example 3 -covers them). - -```mermaid -flowchart TB - s0["Stage 0 · 1 candidate
the workload's LogicalDAG"] - p1["Stage 1, Pass 1 · 3 candidates
Q2 exact, Count-Min + heap, or Hydra"] - p2["Stage 1, Pass 2 · 6 candidates
each with separate or shared input"] - s2["Stage 2 · 15 candidates
Q1 and Q2 each materialized or not"] - s3(["Stage 3 · 1 plan"]):::estimate - s0 -- "× summary choices" --> p1 -- "× sharing choices" --> p2 -- "× materialization choices" --> s2 -- "select" --> s3 - classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; - classDef exact fill:#fff,stroke:#5f6368,color:#000; - classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; -``` +candidate covers both queries. **Stage 0: 1 candidate.** The frontend lowers both queries into one workload `LogicalDAG` with no summaries: @@ -411,8 +388,7 @@ flowchart LR **Stage 1, Pass 1: 3 candidates.** Pass 1 finds local options for each query: * **Q1** has one option, the exact per-series rate and per-`job` sum. Its - accuracy requirement is `implicit_exact`, and an exact grouped sum is - already small. + accuracy requirement is exact, so no summary qualifies. * **Q2** has three options. The exact one keeps a per-series sum and sorts within each `job`, which is costly at one million Zipf-distributed series. The `EpsilonDelta` target also admits a **Count-Min Sketch with a top-*k* @@ -444,128 +420,153 @@ flowchart LR Combining them gives 1 × 3 = 3 workload candidates: -| Candidate | Q1 | Q2 | +| Logical candidate | Q1 | Q2 | |---|---|---| -| L1 | exact | exact | -| L2 | exact | Count-Min + heap | -| L3 | exact | Hydra | - -**Stage 1, Pass 2: 6 candidates.** Both queries read the same range selector, -`http_requests_total[1m]`, so the identical-expression rule adds, for each of -L1–L3, a variant in which Q1 and Q2 share one input node: L1s, L2s and L3s. The -originals are kept. No summary is shared, because Q1 must be exact and no -summary supports both queries. +| Exact | exact | exact | +| Count-Min | exact | Count-Min + heap | +| Hydra | exact | Hydra | + +**Stage 1, Pass 2: 24 candidates.** Pass 2 applies two ASAP-aware CSE rules +to each of the 3 candidates, and keeps every original: + +* **Identical-expression rule.** Both queries read the same range selector, + `http_requests_total[1m]`, so Pass 2 adds a variant in which Q1 and Q2 share + one input node. No summary is shared, because Q1 must be exact and no + summary supports both queries. +* **Window-composition rule.** Each query reads a 1-min window every 10 s, so + consecutive evaluations overlap by 50 s. For each query, Pass 2 adds a + variant that computes its window from **10-s panes**: 6 per-pane states, + merged at every refresh. Q1's per-series rates and sums, Q2's exact sums, + Count-Min sketches and Hydra sketches can all be built per pane and merged. + For Count-Min + heap, the per-pane heaps only approximate the merged top 10, + which the accuracy model accounts for in stage 3. + +Each Pass 1 candidate therefore has 2 input choices (separate or shared) × +2 choices for Q1 (no panes or panes) × 2 for Q2, so stage 1 outputs +3 × 2 × 2 × 2 = 24 logical candidates. A candidate is named by its choices, +for example *Hydra, Q1 panes, Q2 panes, shared input*. + +The two rules applied to the Hydra candidate: ```mermaid flowchart LR - subgraph SEP["L3 · separate inputs"] + subgraph SEP["Hydra · separate inputs, no panes"] direction LR a1[("http_requests_total")]:::data --> a2["range 1m"]:::exact --> a3["rate → sum by (job)"]:::exact b1[("http_requests_total")]:::data --> b2["range 1m"]:::exact --> b3["Hydra"]:::summary --> b4(["top 10"]):::estimate end - subgraph SH["L3s · shared input"] + subgraph SH["Hydra · shared input, no panes"] direction LR c1[("http_requests_total")]:::data --> c2["range 1m
one shared input node"]:::exact c2 --> c3["rate → sum by (job)"]:::exact c2 --> c4["Hydra"]:::summary --> c5(["top 10"]):::estimate end + subgraph PN["Hydra · Q2 panes"] + direction LR + d1[("http_requests_total")]:::data --> d2["Hydra per
10-s pane"]:::summary --> d3["merge latest
6 panes"]:::exact --> d4(["top 10"]):::estimate + end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Stage 2: 15 candidates.** For every logical candidate, stage 2 decides -separately for Q1 and Q2 whether to materialize the computation at ingestion -time or recompute it from raw samples at every refresh: +**Stage 2: 75 candidates.** Stage 2 picks, for each query, when its state is +computed and whether it is kept. What it can choose depends on the query's +logical form from Pass 2: + +| Query's logical form | Physical options | +|---|---| +| No panes | **Raw:** rebuild the state from the last 1 min of raw samples at every refresh. It cannot be materialized: the window slides every 10 s, and most summaries cannot drop old data. | +| 10-s panes | **Ingestion panes:** build each pane as samples arrive and keep the last 6 (materialized at ingestion time). **Query panes:** at each refresh, build only the newest pane from the last 10 s of raw samples and reuse the 5 kept panes (materialized at query time). **Rebuilt panes:** at each refresh, build all 6 panes from the last 1 min of raw samples, merge them, and discard them (not materialized). | ```mermaid flowchart LR in[("http_requests_total
samples")]:::data - subgraph Q1M["Q1 · materialized"] + subgraph R["Raw · no panes"] direction LR - subgraph Q1MI["Ingestion time"] - m1["maintain per-series rate
and sum by (job)"]:::exact - end - subgraph Q1MQ["Query time, every 10 s"] - m1r["read Q1 result"]:::exact + subgraph RQ["Query time, every 10 s"] + r1["range 1m"]:::exact --> r2["build state"]:::summary --> r3(["answer"]):::estimate end - m1 --> m1r end - subgraph Q1N["Q1 · not materialized"] + subgraph IP["Ingestion panes"] direction LR - subgraph Q1NQ["Query time, every 10 s"] - n1a["range 1m"]:::exact --> n1b["rate
per series"]:::exact --> n1c["sum by (job)"]:::exact + subgraph IPI["Ingestion time"] + i1["build 10-s pane
keep last 6"]:::summary end + subgraph IPQ["Query time, every 10 s"] + i2["merge 6 panes"]:::exact --> i3(["answer"]):::estimate + end + i1 --> i2 end - subgraph Q2M["Q2 · materialized"] + subgraph QP["Query panes"] direction LR - subgraph Q2MI["Ingestion time"] - m2["maintain Q2's summary
or exact state"]:::summary - end - subgraph Q2MQ["Query time, every 10 s"] - m2e(["top 10
per job"]):::estimate + subgraph QPQ["Query time, every 10 s"] + q1["range 10 s"]:::exact --> q2["build newest pane"]:::summary --> q3["merge with
5 kept panes"]:::exact --> q4(["answer"]):::estimate end - m2 --> m2e end - subgraph Q2N["Q2 · not materialized"] + subgraph RB["Rebuilt panes"] direction LR - subgraph Q2NQ["Query time, every 10 s"] - n2a["range 1m"]:::exact --> n2b["build Q2's summary
or exact state"]:::summary --> n2c(["top 10
per job"]):::estimate + subgraph RBQ["Query time, every 10 s"] + b1["range 1m"]:::exact --> b2["build 6 panes"]:::summary --> b3["merge 6 panes"]:::exact --> b4(["answer"]):::estimate end end - in --> m1 - in --> n1a - in --> m2 - in --> n2a + in --> r1 + in --> i1 + in --> q1 + in --> b1 classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -That gives 2 × 2 = 4 materialization choices per logical candidate. A shared -input node only matters when both queries read raw samples at query time; in -the other three choices, the shared-input variant is the same plan as its -original. So each Q2 option yields 5 distinct physical candidates, 15 in all: +Each query thus ends up with one of four options: Raw, Ingestion panes, Query +panes or Rebuilt panes, giving 4 × 4 = 16 combinations per Q2 option. A shared +input node only changes the plan when both queries read raw samples at query +time, that is, when neither uses Ingestion panes; those 9 combinations also +have a shared-input version: -| Q2 option (logical candidates) | Q1 mat., Q2 mat. | Q1 mat., Q2 not | Q1 not, Q2 mat. | Q1 not, Q2 not, separate inputs | Q1 not, Q2 not, shared input | -|---|---|---|---|---|---| -| Exact (L1, L1s) | P1 | P2 | P3 | P4 | P5 | -| Count-Min + heap (L2, L2s) | P6 | P7 | P8 | P9 | P10 | -| Hydra (L3, L3s) | P11 | P12 | P13 | P14 | P15 | +| Q1 \ Q2 | Raw | Ingestion panes | Query panes | Rebuilt panes | +|---|---|---|---|---| +| **Raw** | separate or shared input | one plan | separate or shared input | separate or shared input | +| **Ingestion panes** | one plan | one plan | one plan | one plan | +| **Query panes** | separate or shared input | one plan | separate or shared input | separate or shared input | +| **Rebuilt panes** | separate or shared input | one plan | separate or shared input | separate or shared input | + +That is 16 + 9 = 25 plans for each of Q2's 3 options, or 75 physical +candidates. A candidate is named by Q2's option and each query's physical +option, for example *Hydra, Q1 ingestion panes, Q2 ingestion panes*. **Stage 3: 1 plan.** Selection first rejects invalid candidates. The deployment's accuracy model checks the summary candidates against ε = 0.01, -δ = 0.001, and its cost model estimates Q2's latency against the 100 ms bound; -for example, an exact top-k sorted over one million series at query time (P2, -P4, P5) may miss it. Among the rest, it picks the cheapest for the whole -workload. Because both panels refresh every 10 s over arriving data, a -candidate that materializes both queries usually wins: with many small jobs, -P11 (Hydra); with a few large jobs, P6 (Count-Min per `job`). The -not-materialized candidates win only when storage is expensive and raw data is -available at query time. +δ = 0.001, including the approximate heap merge of Count-Min panes. Its cost +model estimates Q2's latency against the 100 ms bound; for example, an exact +top 10 rebuilt with Raw over one million series at every refresh may miss it. +Among the rest, selection picks the cheapest plan for the whole workload: + +* **Usually:** both queries on ingestion panes, with Hydra when there are many + small jobs or Count-Min when there are a few large jobs. Each sample is + processed once, and each refresh only merges 6 small panes. +* **When ingestion-time work is expensive:** query panes, which still build + each pane only once but do it at query time. +* **When storage is expensive and raw data is available at query time:** Raw + for both queries with a shared input, which keeps nothing and reads the + last minute of samples once for both queries. +* **Rarely:** Rebuilt panes. They do the same raw read as Raw plus extra merge + work, so the cost model usually ranks them below Raw. They stay in the + candidate set because they are valid; only selection rules them out. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 **Query workload.** A network-monitoring dashboard computes three statistics of source IPs over the last minute. -| Workload field | Value (all three queries) | -|---|---| -| `language` | `sql` (`datafusion_sql`) | -| Entry type | `repeating_queries` | -| `demand` | every 10 s (`fixed_interval`) | -| `predictability` | `predictable` | -| `time_selection.scope` | `real_time` | -| `as_of` | evaluation time | -| Latency requirement | unspecified | - ```sql -- Q1: Distinct(src_ip) SELECT COUNT(DISTINCT src_ip) @@ -591,11 +592,11 @@ FROM ( ); ``` -| Query | Computation | `lookback` | Accuracy requirement | -|---|---|---|---| -| Q1 | `Distinct(src_ip)` | 1 m | ε = 0.02, δ = 0.01 | -| Q2 | `Entropy(src_ip)` | 1 m | ε = 0.05, δ = 0.01 | -| Q3 | `L2(src_ip)` | 1 m | ε = 0.01, δ = 0.01 | +| Query | Computation | Repeats | `lookback` | `as_of` | Accuracy requirement | +|---|---|---|---|---|---| +| Q1 | `Distinct(src_ip)` | every 10 s | 1 m | evaluation time | ε = 0.02, δ = 0.01 | +| Q2 | `Entropy(src_ip)` | every 10 s | 1 m | evaluation time | ε = 0.05, δ = 0.01 | +| Q3 | `L2(src_ip)` | every 10 s | 1 m | evaluation time | ε = 0.01, δ = 0.01 | The data workload differs from the shared one in two fields: @@ -666,7 +667,7 @@ flowchart LR e2 -.-> u l2 -.-> u - OUT[["CandidateLogicalASAPDAGs:
all Pass 1 candidates + the shared candidate"]] + OUT[["CandidateLogicalASAPDAGs:
27 Pass 1 candidates + 10 shared candidates"]] P1 --> OUT P2 --> OUT classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; @@ -677,7 +678,12 @@ flowchart LR The three UnivMon options from Pass 1 (dashed arrows) are merged by Pass 2 into one shared UnivMon. The independent candidates are kept, so stage 1 outputs -both. +both kinds. + +**Stage 2.** As in Example 1, every summary in every candidate can be +materialized at ingestion time or rebuilt at each 10-s refresh. Because the +dashboard repeats over arriving data, materializing at ingestion time is +usually cheaper. **Stage 3.** Selection compares one UnivMon sized for ε = 0.01 against three separate summaries, each sized for its own requirement. The shared candidate @@ -694,14 +700,9 @@ all executed together at time T. | Workload field | Value (all five queries) | |---|---| -| `language` | `promql` | -| Entry type | `query_batch` | -| `invocations` | 1 | -| `execute_at` | T | +| Entry type | `query_batch`, run once (`invocations: 1`) at T | | `predictability` | `ad_hoc` | -| `time_selection.scope` | `longitudinal` | | Accuracy requirement | ε = 0.005, δ = 0.01 | -| Latency requirement | unspecified | | Query | `lookback` | `as_of` | |---|---|---| @@ -723,9 +724,10 @@ of data at rest, plus data still arriving. its interval. The independent candidates are kept. Counted at the workload level, Pass 1 gives each query 2 options (exact or -KLL), so 2⁵ = 32 candidates. Pass 2 adds a candidate for every subset of two -or more queries that shares one Exponential Histogram while the others keep -their own options, 131 more, for 163 in total. Counts like this are why +KLL), so 2⁵ = 32 candidates. Pass 2 adds a candidate for every way of grouping +two or more queries onto shared Exponential Histograms (one group of five, or +several smaller groups such as two pairs) while the remaining queries keep +their own options: 171 more, for 203 in total. Counts like this are why candidate sets may be enumerated lazily. The five query intervals overlap, and all lie inside the last five years: @@ -763,17 +765,9 @@ flowchart LR **Pattern B: one repeating query over a sliding window.** A real-time p99 panel over the last 5 min, refreshed every minute. -| Workload field | Value | -|---|---| -| `language` | `promql` | -| Entry type | `repeating_queries` | -| `demand` | every 1 min (`fixed_interval`) | -| `predictability` | `predictable` | -| `time_selection.scope` | `real_time` | - -| Query | `lookback` | `as_of` | Accuracy requirement | Latency requirement | -|---|---|---|---|---| -| `quantile_over_time(0.99, latency_ms[5m])` | 5 m | evaluation time | ε = 0.01, δ = 0.01 | ≤ 200 ms | +| Query | Repeats | `lookback` | `as_of` | Accuracy requirement | Latency requirement | +|---|---|---|---|---|---| +| `quantile_over_time(0.99, latency_ms[5m])` | every 1 min | 5 m | evaluation time | ε = 0.01, δ = 0.01 | ≤ 200 ms | * **Pass 1.** One KLL over 5 min for each evaluation. * **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The @@ -808,10 +802,10 @@ Example 4 shows how stage 2 decides whether to store these window summaries. ### Example 4: Materialization of window summaries in physical planning Stage 2 takes the shared window summaries from Example 3 and decides whether to -materialize them. This example shows only the physical candidates of the -shared logical candidate; every other logical candidate from Example 3 gets -its own physical candidates the same way. That choice is driven by the workload's `recurrence`, -`predictability` and `data_workload.arrival`. +materialize them. That choice is driven by the workload's `recurrence`, +`predictability` and `data_workload.arrival`. This example shows only the +physical candidates of the shared logical candidate; every other logical +candidate from Example 3 gets its own physical candidates the same way. **Pattern A (sub-interval batch).** @@ -860,6 +854,7 @@ Which candidate wins depends on the workload: |---|---|---|---| | B1 | 1-min KLL panes, retained 5 min | Build one KLL pane per minute | Merge the latest 5 panes, read p99 | | B2 | Nothing | Nothing | Read 5 min of raw samples, build one KLL, read p99 | +| B3 | 1-min KLL panes, at query time, retained 5 min | Nothing | Build only the newest pane from raw samples, merge it with the 4 kept panes, read p99 | ```mermaid flowchart LR @@ -879,6 +874,13 @@ flowchart LR s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::estimate end end + subgraph B3["B3 · panes materialized at query time"] + direction LR + subgraph B3Q["Query time, every 1 min"] + s5[("last 1 min of
raw samples")]:::data --> k5["build newest
1-min KLL pane"]:::summary --> g5["merge with 4
kept panes"]:::exact --> r5(["p99"]):::estimate + kp["4 kept panes
from earlier evaluations"]:::summary --> g5 + end + end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; @@ -887,8 +889,11 @@ flowchart LR The query repeats every minute and the data is continuously ingesting, so B1 builds each pane once and reuses it in five evaluations, while B2 rescans raw -data every time. Selection usually picks B1. B2 wins only if storage is -expensive and raw data is available at query time. +data every time. B3 also builds each pane once, but at query time, so it needs +raw data at query time and adds the newest pane's build to each evaluation's +latency. Selection usually picks B1. B3 can win when ingestion-time work is +expensive, and B2 only when storage is expensive and raw data is available at +query time. **What this shows.** The same logical candidate (one shared Exponential Histogram, or a sliding window of KLL panes) yields different physical plans depending From 47e705a4bf9c372c62a1cca25054f059d57e336b Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:42:57 -0400 Subject: [PATCH 48/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 379 ++++++++++-------- 1 file changed, 212 insertions(+), 167 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index f85773c1..c2de34c7 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -146,8 +146,8 @@ placement are decided in later stages. #### Pass 1: Local candidate generation For each eligible sub-DAG, Pass 1 identifies its computation semantics, applies -rewrite rules, and generates every candidate that can meet its accuracy -requirement. +rewrite rules, and generates every candidate that is not provably unable to +meet its accuracy requirement. Example for summary candidates: @@ -169,8 +169,8 @@ A summary-based candidate uses three kinds of summary nodes: * A **summary build node** builds and maintains a summary from input data, for example a KLL sketch over `latency_ms`. * A **summary merge node** combines summaries into one, for example merging - five 1-min KLL panes into one 5-min KLL, or merging lower-level summaries - into a coarser one. + five 1-min tumbling-window KLLs into one 5-min KLL, or merging lower-level + summaries into a coarser one. * A **summary estimation node** computes an answer from a summary, for example the p99 estimate from a KLL, or the entropy estimate from a UnivMon. * **summary subtract node** and **summary delete node** design is TODO. @@ -191,35 +191,44 @@ or value being summarized together with its grouping. The summary input data doe not include the window; the window-composition rule compares windows separately. -The window-composition rule distinguishes what a query reads from what the -planner stores: +The window-composition rule distinguishes the window a query reads from the +window summary that answers it: * A **window** is the time range one query evaluation reads, for example the last 5 min. Consecutive evaluations of a repeating query read overlapping - windows; this is a **sliding window**. -* A **pane** is what the planner summarizes and stores. The time axis is cut - into back-to-back panes of equal length (a tumbling window), each with one - summary. A window is answered by merging the summaries of the panes it - covers, so overlapping windows reuse the same panes instead of each building - its own summary. Most summaries cannot remove old data, so one summary cannot - simply slide forward; panes avoid that, because the oldest pane is dropped - rather than subtracted. The pane length must divide both the window length - and the evaluation interval, so that every window starts and ends on a pane - boundary and never needs part of a pane: a 5-min window evaluated every - 1 min uses 1-min panes, and each window merges exactly 5 whole panes. -* A **bucket** of an Exponential Histogram plays the same role, but bucket - lengths grow with age: recent data sits in short buckets and older data in - longer ones. This keeps few buckets over a long history, at the cost that old - window boundaries may fall inside a bucket and are then approximate. -* A **window summary** is a summary organized as panes or buckets so that it - can answer many windows, for example a sliding window of panes, a tumbling - window, or an Exponential Histogram. + windows. Most summaries cannot remove old data, so one summary cannot simply + slide forward with the window. +* A **window summary** keeps summaries so that many windows can be answered. + Three window summaries are considered for now: + * **Sliding window:** one summary per active window. Each arriving sample is + inserted into every active window that contains it, and each evaluation + reads the window that has just completed, with no merge. For a 5-min + window evaluated every 1 min, 5 windows are active and each sample updates + all 5. It works for any summary, including ones that cannot be merged, at + the cost of more ingestion work and memory. + * **Tumbling window:** back-to-back, non-overlapping windows of one fixed + length, each with one summary. A longer query window is answered by + merging the tumbling windows it covers. The tumbling length must divide + both the query window length and the evaluation interval, so that every + query window starts and ends on a tumbling boundary: a 5-min window + evaluated every 1 min uses 1-min tumbling windows and merges exactly 5 of + them. It needs a mergeable summary. + * **Exponential Histogram (EH):** a sequence of EH buckets that covers a + long history. A query window is answered by merging the EH buckets it + covers. Few EH buckets cover a long history, at the cost that an old + query-window boundary may fall inside an EH bucket and is then + approximate. + * An **EH bucket** is one non-overlapping time range of the history with + one summary of the data in it. Unlike tumbling windows, EH buckets are + not all the same length: they grow with age, so recent data sits in + short EH buckets and older data in longer ones. Adjacent EH buckets are + merged into a longer one as they age. | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| | Identical-expression rule | The input and computation semantics are identical. | One common computation node serving multiple consumers. | | Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary build node feeding several estimation nodes, e.g. UnivMon → distinct count, entropy, L2 norm. | -| Window-composition rule | The computations have the same summary input data, and one window summary can reconstruct the requested windows within their accuracy requirements. | One window summary feeding per-window merge and estimation nodes, e.g. KLL panes in a sliding window, or KLL buckets in an Exponential Histogram. | +| Window-composition rule | The computations have the same summary input data, and one window summary can answer the requested windows within their accuracy requirements. | One window summary feeding per-query merge (where needed) and estimation nodes, e.g. a sliding-window or tumbling-window KLL, or an Exponential Histogram with a KLL per EH bucket. | The examples behind these rules: @@ -230,20 +239,23 @@ The examples behind these rules: distinct-count, an entropy and an L2 estimation node each compute their statistic from it. The UnivMon is sized for the strictest of the three accuracy requirements. -* **Window-composition rule, sliding window (Example 3, Pattern B).** One - sliding window of 1-min KLL panes serves every evaluation of - `quantile_over_time(0.99, latency_ms[5m])`, repeated every minute. Each - evaluation merges the latest 5 panes with a merge node and computes p99 with - an estimation node, so consecutive evaluations share 4 of their 5 panes. +* **Window-composition rule, sliding or tumbling window (Example 3, + Pattern B).** For `quantile_over_time(0.99, latency_ms[5m])` repeated every + minute, one window summary serves every evaluation. With a sliding-window + KLL, each sample updates the 5 active windows, and each evaluation reads the + one that has just completed. With 1-min tumbling-window KLLs, each sample + updates one window, and each evaluation merges the latest 5 with a merge + node; consecutive evaluations share 4 of them. * **Window-composition rule, Exponential Histogram (Example 3, Pattern A).** - One Exponential Histogram of KLL buckets over the last 5 years serves the p99 + One Exponential Histogram over the last 5 years, with a KLL per EH bucket, + serves the p99 queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and - `[3y] offset 2y`. Each query's merge node merges the buckets covering + `[3y] offset 2y`. Each query's merge node merges the EH buckets covering its interval, and its estimation node computes p99 from the merged KLL. * **Other quantiles share for free.** One KLL answers every quantile, so adding `quantile_over_time(0.5, latency_ms[5m])` to the sliding-window dashboard - adds only a p50 estimation node next to the p99 one, reading the same 5 merged - panes, with no new summary. + adds only a p50 estimation node next to the p99 one, reading the same KLL, + with no new summary. Rules are defined by each summary family's capabilities and semantic requirements. A shared summary must meet the strictest accuracy requirement @@ -264,8 +276,8 @@ stored. Materialization does not imply ingestion time; a sub-DAG has three options: * **Materialized at ingestion time:** the sub-DAG runs as data arrives, and its - output is stored before any query asks for it. For example, the 1-min KLL - panes in Example 4, Pattern B. + output is stored before any query asks for it. For example, the 1-min + tumbling-window KLLs in Example 4, Pattern B. * **Materialized at query time:** the sub-DAG runs when a query first needs it, and its output is stored so that later executions, or other queries in the same batch, reuse it instead of recomputing it. For example, an @@ -426,7 +438,7 @@ Combining them gives 1 × 3 = 3 workload candidates: | Count-Min | exact | Count-Min + heap | | Hydra | exact | Hydra | -**Stage 1, Pass 2: 24 candidates.** Pass 2 applies two ASAP-aware CSE rules +**Stage 1, Pass 2: 54 candidates.** Pass 2 applies two ASAP-aware CSE rules to each of the 3 candidates, and keeps every original: * **Identical-expression rule.** Both queries read the same range selector, @@ -434,36 +446,33 @@ to each of the 3 candidates, and keeps every original: one input node. No summary is shared, because Q1 must be exact and no summary supports both queries. * **Window-composition rule.** Each query reads a 1-min window every 10 s, so - consecutive evaluations overlap by 50 s. For each query, Pass 2 adds a - variant that computes its window from **10-s panes**: 6 per-pane states, - merged at every refresh. Q1's per-series rates and sums, Q2's exact sums, - Count-Min sketches and Hydra sketches can all be built per pane and merged. - For Count-Min + heap, the per-pane heaps only approximate the merged top 10, - which the accuracy model accounts for in stage 3. + consecutive evaluations overlap by 50 s. For each query, Pass 2 adds two + variants: a **sliding window**, where each sample updates the 6 active 1-min + windows, and **10-s tumbling windows**, merged 6 at a time at every refresh. + Tumbling windows need a mergeable summary. Q1's rates and sums, Q2's exact + sums and Hydra merge exactly; Count-Min sketches do too, but their top-10 + heaps merge only approximately, so that candidate is kept and the accuracy + model judges it in stage 3. Each Pass 1 candidate therefore has 2 input choices (separate or shared) × -2 choices for Q1 (no panes or panes) × 2 for Q2, so stage 1 outputs -3 × 2 × 2 × 2 = 24 logical candidates. A candidate is named by its choices, -for example *Hydra, Q1 panes, Q2 panes, shared input*. +3 window forms for Q1 (none, sliding, tumbling) × 3 for Q2, so stage 1 outputs +3 × 2 × 3 × 3 = 54 logical candidates. -The two rules applied to the Hydra candidate: +The window-composition rule applied to Q2 in the Hydra candidate: ```mermaid flowchart LR - subgraph SEP["Hydra · separate inputs, no panes"] + subgraph NO["Hydra · no window summary"] direction LR - a1[("http_requests_total")]:::data --> a2["range 1m"]:::exact --> a3["rate → sum by (job)"]:::exact - b1[("http_requests_total")]:::data --> b2["range 1m"]:::exact --> b3["Hydra"]:::summary --> b4(["top 10"]):::estimate + a1[("http_requests_total")]:::data --> a2["range 1m"]:::exact --> a3["Hydra"]:::summary --> a4(["top 10"]):::estimate end - subgraph SH["Hydra · shared input, no panes"] + subgraph SL["Hydra · sliding window"] direction LR - c1[("http_requests_total")]:::data --> c2["range 1m
one shared input node"]:::exact - c2 --> c3["rate → sum by (job)"]:::exact - c2 --> c4["Hydra"]:::summary --> c5(["top 10"]):::estimate + b1[("http_requests_total")]:::data --> b2["insert each sample into
6 active 1-min windows"]:::summary --> b3(["top 10 from the
completed window"]):::estimate end - subgraph PN["Hydra · Q2 panes"] + subgraph TU["Hydra · 10-s tumbling windows"] direction LR - d1[("http_requests_total")]:::data --> d2["Hydra per
10-s pane"]:::summary --> d3["merge latest
6 panes"]:::exact --> d4(["top 10"]):::estimate + c1[("http_requests_total")]:::data --> c2["Hydra per
10-s window"]:::summary --> c3["merge latest 6"]:::exact --> c4(["top 10"]):::estimate end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; @@ -471,96 +480,110 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Stage 2: 75 candidates.** Stage 2 picks, for each query, when its state is +**Stage 2: 156 candidates.** Stage 2 picks, for each query, when its state is computed and whether it is kept. What it can choose depends on the query's -logical form from Pass 2: +window form from Pass 2: -| Query's logical form | Physical options | +| Window form | Physical options | |---|---| -| No panes | **Raw:** rebuild the state from the last 1 min of raw samples at every refresh. It cannot be materialized: the window slides every 10 s, and most summaries cannot drop old data. | -| 10-s panes | **Ingestion panes:** build each pane as samples arrive and keep the last 6 (materialized at ingestion time). **Query panes:** at each refresh, build only the newest pane from the last 10 s of raw samples and reuse the 5 kept panes (materialized at query time). **Rebuilt panes:** at each refresh, build all 6 panes from the last 1 min of raw samples, merge them, and discard them (not materialized). | +| None | **Raw:** rebuild the state from the last 1 min of raw samples at every refresh. It cannot be materialized: the window slides every 10 s, and most summaries cannot drop old data. | +| Sliding window | **Sliding, ingestion time:** insert each sample into the 6 active windows as it arrives (materialized at ingestion time). **Sliding, query time:** at each refresh, insert the last 10 s of raw samples into the 6 kept active windows (materialized at query time). Building the window from raw samples at query time without keeping it is the same plan as Raw. | +| 10-s tumbling windows | **Tumbling, ingestion time:** build each tumbling window as samples arrive and keep the last 6. **Tumbling, query time:** at each refresh, build only the newest tumbling window from the last 10 s of raw samples and reuse the 5 kept ones. **Tumbling, rebuilt:** at each refresh, build all 6 tumbling windows from the last 1 min of raw samples, merge them, and discard them (not materialized). | ```mermaid flowchart LR in[("http_requests_total
samples")]:::data - subgraph R["Raw · no panes"] + subgraph R["Raw"] direction LR subgraph RQ["Query time, every 10 s"] - r1["range 1m"]:::exact --> r2["build state"]:::summary --> r3(["answer"]):::estimate + r1["last 1 min"]:::exact --> r2["build state"]:::summary --> r3(["answer"]):::estimate end end - subgraph IP["Ingestion panes"] + subgraph SI["Sliding, ingestion time"] direction LR - subgraph IPI["Ingestion time"] - i1["build 10-s pane
keep last 6"]:::summary + subgraph SII["Ingestion time"] + s1["update 6 active
windows per sample"]:::summary end - subgraph IPQ["Query time, every 10 s"] - i2["merge 6 panes"]:::exact --> i3(["answer"]):::estimate + subgraph SIQ["Query time, every 10 s"] + s2(["answer from the
completed window"]):::estimate end - i1 --> i2 + s1 --> s2 end - subgraph QP["Query panes"] + subgraph SQ["Sliding, query time"] direction LR - subgraph QPQ["Query time, every 10 s"] - q1["range 10 s"]:::exact --> q2["build newest pane"]:::summary --> q3["merge with
5 kept panes"]:::exact --> q4(["answer"]):::estimate + subgraph SQQ["Query time, every 10 s"] + t1["last 10 s"]:::exact --> t2["update 6 kept
active windows"]:::summary --> t3(["answer from the
completed window"]):::estimate end end - subgraph RB["Rebuilt panes"] + subgraph TI["Tumbling, ingestion time"] direction LR - subgraph RBQ["Query time, every 10 s"] - b1["range 1m"]:::exact --> b2["build 6 panes"]:::summary --> b3["merge 6 panes"]:::exact --> b4(["answer"]):::estimate + subgraph TII["Ingestion time"] + u1["build 10-s window
keep last 6"]:::summary + end + subgraph TIQ["Query time, every 10 s"] + u2["merge 6"]:::exact --> u3(["answer"]):::estimate + end + u1 --> u2 + end + + subgraph TQ["Tumbling, query time"] + direction LR + subgraph TQQ["Query time, every 10 s"] + v1["last 10 s"]:::exact --> v2["build newest
10-s window"]:::summary --> v3["merge with
5 kept"]:::exact --> v4(["answer"]):::estimate + end + end + + subgraph TR["Tumbling, rebuilt"] + direction LR + subgraph TRQ["Query time, every 10 s"] + w1["last 1 min"]:::exact --> w2["build 6
10-s windows"]:::summary --> w3["merge 6"]:::exact --> w4(["answer"]):::estimate end end in --> r1 - in --> i1 - in --> q1 - in --> b1 + in --> s1 + in --> t1 + in --> u1 + in --> v1 + in --> w1 classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -Each query thus ends up with one of four options: Raw, Ingestion panes, Query -panes or Rebuilt panes, giving 4 × 4 = 16 combinations per Q2 option. A shared -input node only changes the plan when both queries read raw samples at query -time, that is, when neither uses Ingestion panes; those 9 combinations also -have a shared-input version: - -| Q1 \ Q2 | Raw | Ingestion panes | Query panes | Rebuilt panes | -|---|---|---|---|---| -| **Raw** | separate or shared input | one plan | separate or shared input | separate or shared input | -| **Ingestion panes** | one plan | one plan | one plan | one plan | -| **Query panes** | separate or shared input | one plan | separate or shared input | separate or shared input | -| **Rebuilt panes** | separate or shared input | one plan | separate or shared input | separate or shared input | - -That is 16 + 9 = 25 plans for each of Q2's 3 options, or 75 physical -candidates. A candidate is named by Q2's option and each query's physical -option, for example *Hydra, Q1 ingestion panes, Q2 ingestion panes*. +Across the three window forms, each query has 1 + 2 + 3 = 6 physical options, +so each Q2 option gives 6 × 6 = 36 combinations. The shared-input variant only +changes the plan when both queries read raw samples at query time, which 4 of +the 6 options do (all except the two ingestion-time ones). That adds +4 × 4 = 16 shared-input plans, for 52 plans per Q2 option and +3 × 52 = 156 physical candidates. **Stage 3: 1 plan.** Selection first rejects invalid candidates. The deployment's accuracy model checks the summary candidates against ε = 0.01, -δ = 0.001, including the approximate heap merge of Count-Min panes. Its cost -model estimates Q2's latency against the 100 ms bound; for example, an exact -top 10 rebuilt with Raw over one million series at every refresh may miss it. -Among the rest, selection picks the cheapest plan for the whole workload: - -* **Usually:** both queries on ingestion panes, with Hydra when there are many - small jobs or Count-Min when there are a few large jobs. Each sample is - processed once, and each refresh only merges 6 small panes. -* **When ingestion-time work is expensive:** query panes, which still build - each pane only once but do it at query time. -* **When storage is expensive and raw data is available at query time:** Raw - for both queries with a shared input, which keeps nothing and reads the - last minute of samples once for both queries. -* **Rarely:** Rebuilt panes. They do the same raw read as Raw plus extra merge - work, so the cost model usually ranks them below Raw. They stay in the - candidate set because they are valid; only selection rules them out. +δ = 0.001, including the approximate heap merge of Count-Min with tumbling +windows. Its cost model estimates Q2's latency against the 100 ms bound; for +example, an exact top 10 rebuilt from one million series at every refresh may +miss it. Among the rest, selection picks the cheapest plan for the whole +workload. Typical winners: + +* **Hydra for Q2, with both queries on tumbling windows at ingestion time**, + when there are many small jobs: each sample updates one window, and each + refresh merges 6 small summaries. +* **Count-Min + heap for Q2 on a sliding window at ingestion time**, with Q1 + on tumbling windows, when there are a few large jobs: the heaps do not merge + cleanly, so updating 6 active windows per sample is worth it. +* **Query-time variants** of either, when ingestion-time work is expensive. +* **Raw with a shared input** for both queries, when storage is expensive and + raw data is available at query time. + +The other candidates, such as tumbling windows rebuilt at every refresh, stay +in the candidate set because they are valid, and selection rules them out on +cost. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 @@ -616,7 +639,10 @@ UnivMon build node feeding three estimation nodes**. It must be sized for the st requirement, ε = 0.01. The independent candidates are kept as well. Pass 2 also adds candidates where only two of the three share a UnivMon and the third keeps any of its own 3 options (3 pairs × 3 = 9), so stage 1 outputs -27 + 1 + 9 = 37 candidates. The figure shows the all-three case. +27 + 1 + 9 = 37 candidates. The figure shows the all-three case. The +window-composition rule would also add sliding-window and 10-s tumbling-window +variants, exactly as in Example 1; they are left out here to keep the focus on +the summary-capability rule. ```mermaid flowchart LR @@ -680,10 +706,12 @@ The three UnivMon options from Pass 1 (dashed arrows) are merged by Pass 2 into one shared UnivMon. The independent candidates are kept, so stage 1 outputs both kinds. -**Stage 2.** As in Example 1, every summary in every candidate can be -materialized at ingestion time or rebuilt at each 10-s refresh. Because the -dashboard repeats over arriving data, materializing at ingestion time is -usually cheaper. +**Stage 2.** The same options as in Example 1 apply: a summary without a +window summary is rebuilt from raw samples at every refresh, while a +sliding-window or tumbling-window UnivMon can be kept from ingestion time or +from query time (and tumbling windows can also be rebuilt). Because the +dashboard repeats over arriving data, ingestion-time window summaries are +usually cheapest. UnivMon merges exactly, so tumbling windows suit it. **Stage 3.** Selection compares one UnivMon sized for ε = 0.01 against three separate summaries, each sized for its own requirement. The shared candidate @@ -719,16 +747,17 @@ of data at rest, plus data still arriving. its own interval: five independent KLL candidates over overlapping data. * **Pass 2.** Every interval is a sub-interval of [T − 5 y, T], and KLL is mergeable. The window-composition rule adds a shared candidate: **one - Exponential Histogram of KLL buckets over [T − 5 y, T]**, with one - merge and estimation node per query that merges the buckets covering + Exponential Histogram over [T − 5 y, T], with a KLL per EH bucket**, with one + merge and estimation node per query that merges the EH buckets covering its interval. The independent candidates are kept. -Counted at the workload level, Pass 1 gives each query 2 options (exact or -KLL), so 2⁵ = 32 candidates. Pass 2 adds a candidate for every way of grouping -two or more queries onto shared Exponential Histograms (one group of five, or -several smaller groups such as two pairs) while the remaining queries keep -their own options: 171 more, for 203 in total. Counts like this are why -candidate sets may be enumerated lazily. +Yearly tumbling windows would also work here, since every interval is a whole +number of years; the Exponential Histogram is shown because it also handles +intervals that are not. Counted at the workload level, Pass 1 gives each query +2 options (exact or KLL), so 2⁵ = 32 candidates, and Pass 2 adds one candidate +for every way of grouping two or more queries onto shared window summaries. +That quickly reaches hundreds of candidates, which is why candidate sets may be +enumerated lazily. The five query intervals overlap, and all lie inside the last five years: @@ -750,19 +779,19 @@ and a merge and estimation node per query: ```mermaid flowchart LR - in[("latency_ms
T − 5y to T")]:::data --> eh["Exponential Histogram
of KLL buckets"]:::summary - eh --> m1["merge buckets
T−5y … T"]:::exact --> o1(["q1 p99"]):::estimate - eh --> m2["merge buckets
T−1y … T"]:::exact --> o2(["q2 p99"]):::estimate - eh --> m3["merge buckets
T−2y … T−1y"]:::exact --> o3(["q3 p99"]):::estimate - eh --> m4["merge buckets
T−3y … T−2y"]:::exact --> o4(["q4 p99"]):::estimate - eh --> m5["merge buckets
T−5y … T−2y"]:::exact --> o5(["q5 p99"]):::estimate + in[("latency_ms
T − 5y to T")]:::data --> eh["Exponential Histogram
KLL per EH bucket"]:::summary + eh --> m1["merge EH buckets
T−5y … T"]:::exact --> o1(["q1 p99"]):::estimate + eh --> m2["merge EH buckets
T−1y … T"]:::exact --> o2(["q2 p99"]):::estimate + eh --> m3["merge EH buckets
T−2y … T−1y"]:::exact --> o3(["q3 p99"]):::estimate + eh --> m4["merge EH buckets
T−3y … T−2y"]:::exact --> o4(["q4 p99"]):::estimate + eh --> m5["merge EH buckets
T−5y … T−2y"]:::exact --> o5(["q5 p99"]):::estimate classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Pattern B: one repeating query over a sliding window.** A real-time p99 panel +**Pattern B: one repeating query with overlapping windows.** A real-time p99 panel over the last 5 min, refreshed every minute. | Query | Repeats | `lookback` | `as_of` | Accuracy requirement | Latency requirement | @@ -771,30 +800,32 @@ over the last 5 min, refreshed every minute. * **Pass 1.** One KLL over 5 min for each evaluation. * **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The - window-composition rule adds a shared candidate: **a sliding window of - 1-min KLL panes**, where each evaluation merges the latest 5 panes. With the - exact candidate, stage 1 outputs 3 candidates. + window-composition rule adds two shared candidates for each Pass 1 option: + a **sliding window**, where each sample updates the 5 active 5-min windows, + and **1-min tumbling windows**, where each evaluation merges the latest 5. + With exact and KLL from Pass 1, each in 3 window forms (none, sliding, + tumbling), stage 1 outputs 2 × 3 = 6 candidates. -Each evaluation reads five 1-min panes, and consecutive evaluations share -four of them: +With 1-min tumbling windows, each evaluation merges five of them, and +consecutive evaluations share four: ```mermaid gantt - title Pattern B · 1-min KLL panes and 5-min evaluations + title Pattern B · 1-min tumbling KLL windows and 5-min evaluations dateFormat HH:mm axisFormat %H:%M - section KLL panes - pane 1 :p1, 00:00, 1m - pane 2 :p2, 00:01, 1m - pane 3 :p3, 00:02, 1m - pane 4 :p4, 00:03, 1m - pane 5 :p5, 00:04, 1m - pane 6 :p6, 00:05, 1m - pane 7 :p7, 00:06, 1m + section 1-min tumbling windows + window 1 :p1, 00:00, 1m + window 2 :p2, 00:01, 1m + window 3 :p3, 00:02, 1m + window 4 :p4, 00:03, 1m + window 5 :p5, 00:04, 1m + window 6 :p6, 00:05, 1m + window 7 :p7, 00:06, 1m section Evaluations - eval at 00:05 (panes 1–5) :e1, 00:00, 5m - eval at 00:06 (panes 2–6) :e2, 00:01, 5m - eval at 00:07 (panes 3–7) :e3, 00:02, 5m + eval at 00:05 (windows 1–5) :e1, 00:00, 5m + eval at 00:06 (windows 2–6) :e2, 00:01, 5m + eval at 00:07 (windows 3–7) :e3, 00:02, 5m ``` Example 4 shows how stage 2 decides whether to store these window summaries. @@ -813,6 +844,7 @@ candidate from Example 3 gets its own physical candidates the same way. |---|---|---| | A1 | The Exponential Histogram, at query time | At query time, when the batch runs at T; read by all five queries, then discarded | | A2 | The Exponential Histogram, at ingestion time | At ingestion time, with each new sample; old data backfilled once | +| A3 | Nothing | At query time, once per query: each of the five queries rebuilds it for itself and discards it | ```mermaid flowchart LR @@ -833,6 +865,12 @@ flowchart LR end h4 --> r4 end + subgraph A3["A3 · not materialized"] + direction LR + subgraph A3Q["Query time, once per query at T"] + s6[("5 years of
stored samples")]:::data --> h6["build Exponential
Histogram, ×5"]:::summary --> r6(["one estimate
per rebuild"]):::estimate + end + end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; @@ -842,43 +880,44 @@ flowchart LR Which candidate wins depends on the workload: * **As given** (`invocations: 1`, `ad_hoc`): selection picks A1. A2 would - maintain the histogram for years only to serve one batch. + maintain the histogram for years only to serve one batch, and A3 builds it + five times instead of once. * **Repeated monthly and `Predictable { known_at }`:** A2 can win, because its maintenance cost is shared by many batches. * **Data `"at_rest"`:** A2 is not generated, because there is no ingestion to maintain the histogram. -**Pattern B (sliding window, repeating).** +**Pattern B (overlapping windows, repeating).** | Candidate | Materialized | At ingestion time | At query time | |---|---|---|---| -| B1 | 1-min KLL panes, retained 5 min | Build one KLL pane per minute | Merge the latest 5 panes, read p99 | -| B2 | Nothing | Nothing | Read 5 min of raw samples, build one KLL, read p99 | -| B3 | 1-min KLL panes, at query time, retained 5 min | Nothing | Build only the newest pane from raw samples, merge it with the 4 kept panes, read p99 | +| B1 | 1-min tumbling KLLs, kept 5 min | Build one tumbling KLL per minute | Merge the latest 5, read p99 | +| B2 | Nothing | Nothing | Read 5 min of raw samples, rebuild all 5 tumbling KLLs, merge them, read p99 | +| B3 | 1-min tumbling KLLs, at query time, kept 5 min | Nothing | Build only the newest tumbling KLL from raw samples, merge it with the 4 kept ones, read p99 | ```mermaid flowchart LR - subgraph B1["B1 · panes materialized at ingestion time"] + subgraph B1["B1 · tumbling KLLs materialized at ingestion time"] direction LR subgraph B1I["Ingestion time"] - s1[("samples")]:::data --> p1["1-min KLL pane
stored 5 min"]:::summary + s1[("samples")]:::data --> p1["1-min tumbling KLL
kept 5 min"]:::summary end subgraph B1Q["Query time, every 1 min"] - g1["merge latest
5 panes"]:::exact --> r1(["p99"]):::estimate + g1["merge latest 5"]:::exact --> r1(["p99"]):::estimate end p1 --> g1 end subgraph B2["B2 · not materialized"] direction LR subgraph B2Q["Query time, every 1 min"] - s2[("5 min of
raw samples")]:::data --> k2["build one KLL"]:::summary --> r2(["p99"]):::estimate + s2[("5 min of
raw samples")]:::data --> k2["rebuild 5
tumbling KLLs"]:::summary --> m2b["merge 5"]:::exact --> r2(["p99"]):::estimate end end - subgraph B3["B3 · panes materialized at query time"] + subgraph B3["B3 · tumbling KLLs materialized at query time"] direction LR subgraph B3Q["Query time, every 1 min"] - s5[("last 1 min of
raw samples")]:::data --> k5["build newest
1-min KLL pane"]:::summary --> g5["merge with 4
kept panes"]:::exact --> r5(["p99"]):::estimate - kp["4 kept panes
from earlier evaluations"]:::summary --> g5 + s5[("last 1 min of
raw samples")]:::data --> k5["build newest
1-min tumbling KLL"]:::summary --> g5["merge with
4 kept"]:::exact --> r5(["p99"]):::estimate + kp["4 kept tumbling KLLs
from earlier evaluations"]:::summary --> g5 end end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; @@ -887,16 +926,22 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -The query repeats every minute and the data is continuously ingesting, so B1 -builds each pane once and reuses it in five evaluations, while B2 rescans raw -data every time. B3 also builds each pane once, but at query time, so it needs -raw data at query time and adds the newest pane's build to each evaluation's -latency. Selection usually picks B1. B3 can win when ingestion-time work is -expensive, and B2 only when storage is expensive and raw data is available at -query time. +This table covers the tumbling-window candidate. The query repeats every +minute and the data is continuously ingesting, so B1 builds each tumbling KLL +once and reuses it in five evaluations, while B2 rescans raw data every time. +B3 also builds each tumbling KLL once, but at query time, so it needs raw data +at query time and adds the newest build to each evaluation's latency. +Selection usually picks B1. B3 can win when ingestion-time work is expensive, +and B2 only when storage is expensive and raw data is available at query time. + +The sliding-window candidate from Example 3 gets its own physical candidates +the same way: kept from ingestion time (each sample updates the 5 active +windows) or from query time (each evaluation inserts the last minute of raw +samples into the 5 kept windows). It does more ingestion work than B1 but +needs no merge, so it wins only for a summary that merges poorly. **What this shows.** The same logical candidate (one shared Exponential -Histogram, or a sliding window of KLL panes) yields different physical plans depending +Histogram, or 1-min tumbling KLL windows) yields different physical plans depending only on recurrence, predictability and data arrival. This is why window-summary replacement happens in logical planning, while materialization is decided separately in physical planning. From 8c29c2e376800ad0823d4917e121eed3b60ed3ab Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 15:47:48 -0400 Subject: [PATCH 49/56] Update planner-layering.md From 369f2e62357e5d45b880baf6f4d57806554b02d5 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:54:15 +0000 Subject: [PATCH 50/56] docs: clarify planning contracts and simplify examples --- .../design_docs/proposals/planner-layering.md | 766 ++++++------------ 1 file changed, 262 insertions(+), 504 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index c2de34c7..8697eddf 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -3,95 +3,38 @@ Status: proposal. Audience: designers and developers of ASAPPlanner and of deployments such as ASAPQuery-backend. -Read Stages for the overview, the stage sections for the rules, and the -examples for why the rules are needed. - ## Goal ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the deployment's -inputs (TODO: define this data structure in a follow-up PR), and returns one optimal physical plan. It decides what is computed, how -it is computed, and which plan is best. The deployment only supplies inputs and -executes the plan: it provides its empirical cost model, empirical accuracy -model and capabilities, but never plans queries or selects plans. +inputs (concrete deployment-input types remain a follow-up), and returns the +lowest-cost feasible physical plan within the explored search space, or an +explanation that no candidate qualifies. It decides what is computed, how +it is computed, and which plan is best. The deployment supplies cost/accuracy models and capabilities, then executes the +selected plan. Planning and selection remain inside ASAPPlanner. ## Stages -In the diagram, × means the Cartesian product: each stage combines every option along one dimension with every option along the others. - ```text - Query workload - (PromQL / SQL / MetricsQL, - query recurrence, - accuracy requirements, - latency requirements) - - + Data workload - (streaming vs. data at rest, - data distribution, - cardinality) - - + Deployment inputs - (cost model, - accuracy model, - deployment capabilities) +Query workload + Data workload + Deployment inputs + │ + ▼ + 0. Language frontends + CandidateLogicalDAGs + │ + ▼ + 1. Logical ASAP-aware optimization + CandidateLogicalASAPDAGs + │ + ▼ + 2. Physical ASAP-aware optimization + CandidatePhysicalASAPDAGs │ ▼ -┌────────────────────────────── ASAPPlanner ──────────────────────────────┐ -│ │ -│ 0. Language-specific frontends │ -│ Parse and convert source-language queries into a common logical │ -│ representation. Reject unsupported query expressions. │ -│ │ -│ Output: CandidateLogicalDAGs │ -│ │ -│ │ │ -│ ▼ │ -│ Logical planning — what to compute │ -│ │ -│ 1. Logical ASAP-aware optimization │ -│ Explore semantically equivalent and legal logical candidates: │ -│ │ -│ summary families │ -│ × query rewrites │ -│ × sharing one summary across multiple computations │ -│ │ -│ Output: CandidateLogicalASAPDAGs │ -│ │ -│ │ │ -│ ▼ │ -│ Physical planning — how to compute │ -│ │ -│ 2. Physical ASAP-aware optimization │ -│ Explore executable implementations of each logical candidate: │ -│ │ -│ materialization decisions │ -│ × physical operator implementations │ -│ × parallelism and partitioning │ -│ × resource management │ -│ │ -│ Output: CandidatePhysicalASAPDAGs │ -│ │ -│ │ │ -│ ▼ │ -│ 3. Plan selection │ -│ Evaluate complete physical candidates using the deployment's │ -│ empirical cost and accuracy models. Reject candidates that violate │ -│ accuracy, latency, or capability constraints. │ -│ │ -│ Choose the cheapest valid plan for the whole workload. │ -│ │ -└────────────────────────────────┬───────────────────────────────────────┘ - │ - ▼ - one selected PhysicalASAPDAG - │ - ▼ -┌────────────────────────────── Deployment ───────────────────────────────┐ -│ │ -│ 4. Execution │ -│ deployment executes the selected DAG (plan). │ -│ │ -└────────────────────────────────────────────────────────────────────────┘ + 3. Plan selection + one PhysicalASAPDAG, or failure reasons + │ success + ▼ + 4. Execution by the deployment ``` The planner receives three groups of inputs: @@ -104,34 +47,54 @@ The planner receives three groups of inputs: ## Stages and their decisions -Stages 0 to 2 each output a candidate set holding every semantically equivalent -and legal candidate DAG of that stage; stage 3 is the only step that chooses -one candidate DAG as output. Candidate sets are internal to ASAPPlanner and may -be shared or enumerated lazily. A stage may prune a candidate early only when -it is provably invalid (for example, a summary family that cannot meet the -query's accuracy target), and every rejected candidate carries a reason. +Stages 0–2 retain alternatives generated by the supported rewrite rules and +parameter domains; stage 3 selects a plan. This is not enumeration of every +possible equivalent program. Candidate sets may be represented lazily. Early +rejection requires a structural or semantic proof, such as a family lacking a +required operation; unknown empirical accuracy is deferred to selection, not +assumed acceptable. Rejections carry reasons. Valid alternatives are not removed +on estimated cost before stage 3. -A candidate is a DAG for the **whole workload**, not for one query. Each stage -combines its choices for every sub-DAG with the candidates it receives (the × -in the diagram), so the candidate set grows from stage to stage until -selection picks one. Example 1 traces this growth step by step. +Each candidate covers the **whole workload**, preserving the mapping from query +identities to their results. Alternatives combine only when their dependencies, +input contracts and evaluation contexts are compatible. Equivalent generated +candidates may be deduplicated, so counts need not grow monotonically. | Stage | Input | Decides | Output | |---|---|---|---| | 0. Frontends | `query`, `language` | Parse and lower to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | | 1. Logical ASAP-aware optimization | Logical DAGs; accuracy requirements, `time_selection`, repetition interval | Summary replacement (Pass 1); ASAP-aware CSE (Pass 2) | `CandidateLogicalASAPDAGs` | | 2. Physical ASAP-aware optimization | Logical ASAP DAGs; `recurrence`, `predictability`, `DataWorkload` | Materialization; physical operators; parallelism and resources (TODO) | `CandidatePhysicalASAPDAGs` | -| 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG` | +| 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG`, or no feasible candidate with reasons | | 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Run ingestion, storage and query-time computation | Query results | -The sections below describe each stage. +### Proposed stage interfaces + +These are design contracts, not new Rust API declarations. The stage names below +refer to sets of workload plans; they do not require separate operator enums. +Use the common `OperatorNode` / `Operator` / `ScalarExpr` model proposed in +[#511](https://github.com/ProjectASAP/ASAPPlanner/pull/511). That proposal owns node +structure and validation; this one owns planning responsibilities. + +| Interface value | Required contents and invariant | +|---|---| +| `CandidateLogicalDAGs` | Resolved ordinary operator graphs and scalar expressions; a result binding for every workload query; source-language evaluation contexts, schemas and requirements. No ASAP choices. | +| `CandidateLogicalASAPDAGs` | The same query bindings, plus summary families, update/readout semantics, window composition and logical sharing. Retain each consumer's accuracy requirement and any remaining family-parameter alternatives; execution timing may be unassigned. | +| `CandidatePhysicalASAPDAGs` | Concrete implementations and summary parameters; assigned execution timing; materialization lifetimes, startup/backfill and raw-data dependencies. Each graph is structurally executable subject to deployment validation; empirical accuracy, latency and capability feasibility are checked at selection. | +| Selection result | One complete `PhysicalASAPDAG` with its evaluated cost and per-query feasibility evidence, or a no-feasible-candidate outcome with rejection reasons. Unknown required accuracy/capability evidence is not success. | + +Logical planning chooses a family and its semantic parameters (such as requested +quantile). Physical planning enumerates concrete size/precision configurations +within the supported parameter domain. Selection evaluates those configurations +with deployment models; it does not silently resize a selected plan. A search +limit must be reported: the result is optimal only among the explored candidates. +Concrete containers, deployment-model signatures and diagnostic types remain TODO. ### 0. Language-specific frontends -The frontend converts each query into a `LogicalDAG`. Nodes represent logical -query operations, including selectors, transformations, aggregations, grouping -and window semantics. They contain no ASAP summary choices. A construct that cannot be represented -faithfully is rejected. +The frontend lowers each query and assembles workload `LogicalDAG` candidates. +Nodes represent selectors, transformations, aggregations, grouping and windows, +without ASAP choices. Constructs that cannot be represented faithfully are rejected. The frontend preserves source-language behavior, including series identity, evaluation timing and missing-data semantics. @@ -140,16 +103,15 @@ evaluation timing and missing-data semantics. Logical optimization runs in two passes. Pass 1 generates candidates for each computation on its own; Pass 2 finds candidates that share computation across -sub-DAGs and queries. Materialization and execution -placement are decided in later stages. +sub-DAGs and queries. Physical planning assigns materialization and execution timing. #### Pass 1: Local candidate generation For each eligible sub-DAG, Pass 1 identifies its computation semantics, applies -rewrite rules, and generates every candidate that is not provably unable to -meet its accuracy requirement. +registered rewrite rules, and retains alternatives whose semantic contracts +are established and whose accuracy is not provably infeasible. -Example for summary candidates: +Illustrative families, admitted only with registered semantic contracts: | Original computation | Local candidates | |---|---| @@ -164,7 +126,7 @@ Each candidate records its input expression, filter, grouping, window, supported estimates and accuracy requirement. Pass 2 uses these to decide whether candidates can share a summary node. -A summary-based candidate uses three kinds of summary nodes: +Summary-based candidates distinguish these operations: * A **summary build node** builds and maintains a summary from input data, for example a KLL sketch over `latency_ms`. @@ -173,17 +135,15 @@ A summary-based candidate uses three kinds of summary nodes: summaries into a coarser one. * A **summary estimation node** computes an answer from a summary, for example the p99 estimate from a KLL, or the entropy estimate from a UnivMon. -* **summary subtract node** and **summary delete node** design is TODO. - -One summary build node can feed several estimation nodes, which is what Pass 2 -exploits. +* **Summary subtract** and **summary delete** remain TODO. Merge is also a + capability-gated operation: #511 reserves its payload pending a complete + family-specific contract; the examples here do not establish runtime support. #### Pass 2: ASAP-aware common-subexpression elimination -ASAP-aware CSE extends traditional CSE with summary-specific sharing rules. -Computations can share work when they use identical expressions, when one -summary build node supports several estimates, or when one window summary can answer -their overlapping windows. +ASAP-aware CSE adds summary-capability and window-composition rules to +identical-expression reuse. Reuse must preserve evaluation context, volatility, +series identity and missing-data behavior, not merely match expression text. The rules compare computations by their **summary input data**: what a summary for that computation would ingest, namely the data source, the filters, and the key @@ -191,38 +151,18 @@ or value being summarized together with its grouping. The summary input data doe not include the window; the window-composition rule compares windows separately. -The window-composition rule distinguishes the window a query reads from the -window summary that answers it: - -* A **window** is the time range one query evaluation reads, for example the - last 5 min. Consecutive evaluations of a repeating query read overlapping - windows. Most summaries cannot remove old data, so one summary cannot simply - slide forward with the window. -* A **window summary** keeps summaries so that many windows can be answered. - Three window summaries are considered for now: - * **Sliding window:** one summary per active window. Each arriving sample is - inserted into every active window that contains it, and each evaluation - reads the window that has just completed, with no merge. For a 5-min - window evaluated every 1 min, 5 windows are active and each sample updates - all 5. It works for any summary, including ones that cannot be merged, at - the cost of more ingestion work and memory. - * **Tumbling window:** back-to-back, non-overlapping windows of one fixed - length, each with one summary. A longer query window is answered by - merging the tumbling windows it covers. The tumbling length must divide - both the query window length and the evaluation interval, so that every - query window starts and ends on a tumbling boundary: a 5-min window - evaluated every 1 min uses 1-min tumbling windows and merges exactly 5 of - them. It needs a mergeable summary. - * **Exponential Histogram (EH):** a sequence of EH buckets that covers a - long history. A query window is answered by merging the EH buckets it - covers. Few EH buckets cover a long history, at the cost that an old - query-window boundary may fall inside an EH bucket and is then - approximate. - * An **EH bucket** is one non-overlapping time range of the history with - one summary of the data in it. Unlike tumbling windows, EH buckets are - not all the same length: they grow with age, so recent data sits in - short EH buckets and older data in longer ones. Adjacent EH buckets are - merged into a longer one as they age. +A **window** is the time range read by one query evaluation. A **window summary** +organizes state across windows so overlapping evaluations can reuse work: + +| Window organization | Update and readout | Required contract | +|---|---|---| +| Sliding | Keep one summary per active window; update every window containing an arriving sample. Read the completed window without merging. A 5-min window refreshed every minute has 5 active summaries. | Valid per-window updates and bounded state lifetime; mergeability is unnecessary. | +| Tumbling | Keep nonoverlapping fixed-length buckets; merge those covering each query window. A 5-min window refreshed every minute can merge 5 one-minute buckets. | Merge semantics for the complete state. Bucket width divides window length and refresh interval, and bucket origin aligns with evaluation boundaries. | +| Exponential Histogram (EH) | Keep nonoverlapping buckets of varying size, merging adjacent buckets as they age; read an interval from its covering buckets. | A defined boundary-bucket rule and error bound, including both boundaries of historical intervals. Coarser old buckets may include data outside the requested interval. | + +The EH+KLL case is conditional: KLL mergeability alone does not bound interval +boundary error. Missing boundary semantics must be defined before candidate +generation; empirical accuracy assessment cannot supply an undefined operation. | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| @@ -230,37 +170,12 @@ window summary that answers it: | Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary build node feeding several estimation nodes, e.g. UnivMon → distinct count, entropy, L2 norm. | | Window-composition rule | The computations have the same summary input data, and one window summary can answer the requested windows within their accuracy requirements. | One window summary feeding per-query merge (where needed) and estimation nodes, e.g. a sliding-window or tumbling-window KLL, or an Exponential Histogram with a KLL per EH bucket. | -The examples behind these rules: - -* **Summary-capability rule (Example 2).** One UnivMon over `src_ip` from - `flows` in the last minute serves three queries refreshed every 10 s: - `COUNT(DISTINCT src_ip)`, the entropy of the `src_ip` distribution, and the - L2 norm of per-`src_ip` counts. Each flow record updates the UnivMon once; a - distinct-count, an entropy and an L2 estimation node each compute their - statistic from it. The UnivMon is sized for the strictest of the three accuracy - requirements. -* **Window-composition rule, sliding or tumbling window (Example 3, - Pattern B).** For `quantile_over_time(0.99, latency_ms[5m])` repeated every - minute, one window summary serves every evaluation. With a sliding-window - KLL, each sample updates the 5 active windows, and each evaluation reads the - one that has just completed. With 1-min tumbling-window KLLs, each sample - updates one window, and each evaluation merges the latest 5 with a merge - node; consecutive evaluations share 4 of them. -* **Window-composition rule, Exponential Histogram (Example 3, Pattern A).** - One Exponential Histogram over the last 5 years, with a KLL per EH bucket, - serves the p99 - queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and - `[3y] offset 2y`. Each query's merge node merges the EH buckets covering - its interval, and its estimation node computes p99 from the merged KLL. -* **Other quantiles share for free.** One KLL answers every quantile, so adding - `quantile_over_time(0.5, latency_ms[5m])` to the sliding-window dashboard - adds only a p50 estimation node next to the p99 one, reading the same KLL, - with no new summary. - -Rules are defined by each summary family's capabilities and semantic -requirements. A shared summary must meet the strictest accuracy requirement -among its consumers. Applying a rule adds a shared candidate and keeps the -independent candidates, so selection can compare both. +Examples 2 and 3 illustrate summary-capability and window-composition sharing. +A shared configuration must satisfy **each** consumer's metric and confidence +requirement; ε values for different statistics cannot be ordered as one common +precision target. A second quantile may reuse a KLL if the existing configuration +meets its requirement, but its estimation work still contributes to cost. +Sharing adds alternatives; independent plans remain available to selection. ### 2. Physical ASAP-aware optimization @@ -270,56 +185,50 @@ and resource management are TODO. #### Materialization -Materialization decides, for each sub-DAG, whether its output is kept (persistent to disk or kept in memory) across -(batch) query executions, and if so, when it is computed and how long it is -stored. Materialization does not imply ingestion time; a sub-DAG has three -options: - -* **Materialized at ingestion time:** the sub-DAG runs as data arrives, and its - output is stored before any query asks for it. For example, the 1-min - tumbling-window KLLs in Example 4, Pattern B. -* **Materialized at query time:** the sub-DAG runs when a query first needs - it, and its output is stored so that later executions, or other queries in - the same batch, reuse it instead of recomputing it. For example, an - Exponential Histogram built when the batch in Example 4, Pattern A runs and - read by all five of its queries. -* **Not materialized:** the sub-DAG runs at query time for each execution, and - its output is discarded afterward. - -Whichever option is chosen for each sub-DAG, the plan must also satisfy these -constraints: - -* A materialized output is stored for as long as any of its consumers still - needs it. -* Every node upstream of an ingestion-time node also runs at ingestion time. - -The decision depends on the workload's `recurrence` and `predictability` and on -the `DataWorkload`. Typical outcomes: - -* Read by repeated queries while data keeps arriving: materialize at ingestion - time. -* Read by several queries in one batch, or over data at rest: materialize at - query time. -* Read once by an ad hoc query: do not materialize. +Materialization records which intermediate outputs are retained, their lifetime, +and when they are computed: -A shared summary is materialized once for all its consumers. See Example 4. +| Choice | Computation and reuse | +|---|---| +| Ingestion-time materialization | Build/update as data arrives; keep output for later query evaluations. | +| Query-time materialization | Build when first needed; retain for other consumers in the batch or subsequent evaluations. | +| No materialization | Execute at query time without a retained result for later reuse. A producer may still feed multiple consumers during that execution. | + +Logical producer identity exposes possible reuse; it does not force storage. +Physical planning may duplicate a producer only when repeated evaluation preserves +semantics, including volatility and randomized-summary guarantees. Such copies +are explicit in the physical graph and are costed separately. A producer executed +once is costed once with all consumer demand. Example 4 contrasts these choices. + +Retain output until its last planned consumer has finished. An ingestion-time node +cannot depend on query-time work; its upstream computation must be available in +the ingestion phase. Every retained-state plan specifies initialization/backfill, +retention, and the handling of missed evaluations or late data. Incremental costs +in the examples describe steady state after initialization. Ingestion-time +maintenance requires data arrival and sufficient advance notice; an ad hoc query +cannot assume that years of maintenance have already occurred. #### Physical operator implementation Physical operator implementation lowers every node to physical operators, for -example TopK as a sort followed by a limit, or a KLL node as summary build, -merge and quantile estimation operators. +example choosing a sorting implementation for an exact top-k plan or a backend +implementation for each KLL build, merge and readout. Logical build/merge/readout +semantics are already explicit; physical lowering implements those operations. ### 3. Plan selection Selection rejects every candidate that misses an accuracy target or a latency bound, or that needs a capability the deployment lacks, and then picks the cheapest remaining plan. It is the only stage that uses the deployment's cost -and accuracy models, and the only stage that discards valid candidates. Accuracy is -estimated by the deployment's accuracy model, not assumed from a summary's -nominal bound. Cost is evaluated for the whole workload rather than per query, +and accuracy models, and the only stage that discards valid candidates. Its +accuracy evidence must use the requested metric; +a nominal family bound or an empirical point estimate alone is not a confidence +guarantee. Cost is evaluated for the whole workload rather than per query, which is what lets one shared summary beat several cheaper independent ones: -a shared summary is costed once, with the demand of all its consumers. +a physically shared summary is costed once, with the demand of all its consumers. +Cost includes startup/backfill, maintenance, storage and query execution over the +workload horizon. If no candidate satisfies every requirement, return reasons +rather than a best-effort plan that violates the workload contract. ### 4. Execution @@ -328,26 +237,40 @@ given: it does not choose among summaries or decide what to materialize. ## End-to-end examples -Each example's workload is shown as tables. Field names in code font are the -fields of -[`workload.rs`](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). -Each query reads the event-time window [`as_of` − `lookback`, `as_of`] (fields -of `TimeSelection`). `as_of` is the window's end; `lookback` is its length. An -`as_of` of "evaluation time" means `as_of: None`: the window ends whenever the -query runs, so it moves forward with each evaluation. A fixed `as_of`, such as -T − 1 y, pins the window to a historical interval. Approximate accuracy targets are `EpsilonDelta`: the answer's error -is at most ε with probability at least 1 − δ. +Field names in code font refer to +[`workload.rs`](../../../crates/types/src/workload.rs). `TimeSelection.lookback` +is the window length and `as_of` its end; `as_of: None` uses the evaluation time. +Endpoint inclusion follows the source language: PromQL ranges use +`(as_of − lookback, as_of]`; SQL follows its predicates. PromQL `1y` is 365 days, +not a calendar year. The diagrams use years only as schematic labels. + +The examples reuse `AccuracyTarget::EpsilonDelta` and existing +[`ErrorMetric` / `ResultGuarantee`](../../../crates/types/src/post_asap/guarantee.rs). +Each target needs its result's metric, normalization and failure-probability scope: + +| Result | Metric used in these examples | +|---|---| +| Quantile | `Rank`: normalized rank error, not error in the returned latency value. | +| Distinct count | `Cardinality`: relative count error. | +| Entropy | `AbsoluteValue`: absolute error in nats, matching `LN`. | +| L2 norm | `RelativeValue`: relative norm error; the model must cover the zero case. | +| Top-k | A per-key `Frequency` bound does not establish `TopKMembership`. Selection needs membership evidence as well as score bounds, or a separately specified approximate-membership contract. | + +Here δ applies per query result per evaluation; simultaneous guarantees across +queries or time require an explicit joint failure budget. No independence is +assumed for shared summaries. A model must cover update, merge, boundary selection +and readout errors before a candidate is selectable. | Example | Shows | |---|---| -| 1. Aggregation over dimensions | How the candidate set grows through every stage | +| 1. A workload through every stage | Candidate generation, physical alternatives and selection | | 2. One summary for several computations | Pass 2: summary-capability rule | | 3. Aggregation over windows | Pass 1 and Pass 2: window-composition rule | | 4. Materialization of window summaries | Stage 2: materialization, decoupled from stage 1 | -In the diagrams below, grey cylinders are input data, white boxes are exact -operations and summary merges, blue boxes are summary build nodes, and green -rounded boxes are summary estimation nodes. +In the colored diagrams, grey cylinders are inputs, blue boxes build summaries, +and green nodes read estimates. White merge nodes do not imply exact estimates; +readout guarantees include the merge contract. ### Shared data workload @@ -356,234 +279,83 @@ Unless an example says otherwise, every example uses this data workload: | `DataWorkload` field | Meaning | Value | |---|---|---| | `arrival` | Whether the data is at rest, still arriving, or both | `continuously_ingesting` | -| `data_ingestion_interval` | How often each series delivers one sample (the scrape interval in Prometheus). PromQL uses it as the look-back horizon of instant selectors. | 15 s | +| `data_ingestion_interval` | Sampling cadence, analogous to the Prometheus scrape interval; distinct from instant-selector lookback. | 15 s | | `ingestion_volume` | Total amount of ingested data | unknown | | `ingestion_rate` | Samples arriving per second across all series | about 66,667 samples/s | | `input_cardinality` | Number of distinct series (or keys) | 1,000,000 series | | `distribution` | How samples are spread over keys | `zipf` | With 1,000,000 series each sampled every 15 s, the ingestion rate is -1,000,000 / 15 ≈ 66,667 samples/s. +1,000,000 / 15 ≈ 66,667 samples/s. Instant-selector lookback is a separate +query setting (Prometheus defaults to 5 min), not derived from that cadence. +See the [Prometheus selector semantics](https://prometheus.io/docs/prometheus/latest/querying/basics/#staleness). -### Example 1: Aggregation over dimensions — the candidate set through every stage +### Example 1: A workload through every stage -**Query workload.** Two PromQL dashboard panels over the last minute. The -first needs an exact total; the second tolerates error. +Two PromQL panels refresh every 10 s over the same 1-min input range: -| Query | Repeats | `lookback` | `as_of` | Accuracy requirement | Latency requirement | -|---|---|---|---|---|---| -| Q1: `sum by (job) (rate(http_requests_total[1m]))` | every 10 s | 1 m | evaluation time | exact | none | -| Q2: `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | every 10 s | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | - -This example follows the workload's candidate set through every stage. Each -candidate covers both queries. - -**Stage 0: 1 candidate.** The frontend lowers both queries into one workload -`LogicalDAG` with no summaries: - -```mermaid -flowchart LR - subgraph Q1["Q1 · sum by (job) (rate(http_requests_total[1m]))"] - direction LR - x1[("http_requests_total")]:::data --> x2["range 1m"]:::exact --> x3["rate"]:::exact --> x4["sum by (job)"]:::exact - end - subgraph Q2["Q2 · topk by (job) (10, sum_over_time(http_requests_total[1m]))"] - direction LR - y1[("http_requests_total")]:::data --> y2["range 1m"]:::exact --> y3["sum_over_time"]:::exact --> y4["topk by (job) (10)"]:::exact - end - classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; - classDef exact fill:#fff,stroke:#5f6368,color:#000; - classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; -``` - -**Stage 1, Pass 1: 3 candidates.** Pass 1 finds local options for each query: - -* **Q1** has one option, the exact per-series rate and per-`job` sum. Its - accuracy requirement is exact, so no summary qualifies. -* **Q2** has three options. The exact one keeps a per-series sum and sorts - within each `job`, which is costly at one million Zipf-distributed series. - The `EpsilonDelta` target also admits a **Count-Min Sketch with a top-*k* - heap per `job`**, and **Hydra over the whole `job` column**, where one sketch - covers every (`job`, series) key. - -```mermaid -flowchart LR - subgraph C["Q2's three local options"] - direction TB - subgraph E["Exact"] - direction LR - e1[("input")]:::data --> e2["sum_over_time
per series"]:::exact --> e3["sort + limit 10
per job"]:::exact - end - subgraph CM["Count-Min + heap per job"] - direction LR - c1[("input")]:::data --> c2["Count-Min Sketch +
top-10 heap, one per job"]:::summary --> c3(["top 10
per job"]):::estimate - end - subgraph H["Hydra"] - direction LR - h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::estimate - end - end - classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; - classDef exact fill:#fff,stroke:#5f6368,color:#000; - classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; -``` +| Query | Accuracy requirement | Latency requirement | +|---|---|---| +| Q1: `sum by (job) (rate(http_requests_total[1m]))` | Exact | None | +| Q2: `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | Score ε = 0.01, δ = 0.001 under `Frequency`, plus `TopKMembership` evidence | ≤ 100 ms | -Combining them gives 1 × 3 = 3 workload candidates: +Assume nonnegative finite float samples for Q2. Its sum ranks sampled counter +values; it is not a request-rate ranking. Retain the complete series labels. +The membership requirement is additional to score error; unknown membership +evidence prevents selection of a sketch plan. -| Logical candidate | Q1 | Q2 | -|---|---|---| -| Exact | exact | exact | -| Count-Min | exact | Count-Min + heap | -| Hydra | exact | Hydra | - -**Stage 1, Pass 2: 54 candidates.** Pass 2 applies two ASAP-aware CSE rules -to each of the 3 candidates, and keeps every original: - -* **Identical-expression rule.** Both queries read the same range selector, - `http_requests_total[1m]`, so Pass 2 adds a variant in which Q1 and Q2 share - one input node. No summary is shared, because Q1 must be exact and no - summary supports both queries. -* **Window-composition rule.** Each query reads a 1-min window every 10 s, so - consecutive evaluations overlap by 50 s. For each query, Pass 2 adds two - variants: a **sliding window**, where each sample updates the 6 active 1-min - windows, and **10-s tumbling windows**, merged 6 at a time at every refresh. - Tumbling windows need a mergeable summary. Q1's rates and sums, Q2's exact - sums and Hydra merge exactly; Count-Min sketches do too, but their top-10 - heaps merge only approximately, so that candidate is kept and the accuracy - model judges it in stage 3. - -Each Pass 1 candidate therefore has 2 input choices (separate or shared) × -3 window forms for Q1 (none, sliding, tumbling) × 3 for Q2, so stage 1 outputs -3 × 2 × 3 × 3 = 54 logical candidates. - -The window-composition rule applied to Q2 in the Hydra candidate: +**Stage 0.** One workload graph contains both query results: ```mermaid flowchart LR - subgraph NO["Hydra · no window summary"] - direction LR - a1[("http_requests_total")]:::data --> a2["range 1m"]:::exact --> a3["Hydra"]:::summary --> a4(["top 10"]):::estimate - end - subgraph SL["Hydra · sliding window"] - direction LR - b1[("http_requests_total")]:::data --> b2["insert each sample into
6 active 1-min windows"]:::summary --> b3(["top 10 from the
completed window"]):::estimate - end - subgraph TU["Hydra · 10-s tumbling windows"] - direction LR - c1[("http_requests_total")]:::data --> c2["Hydra per
10-s window"]:::summary --> c3["merge latest 6"]:::exact --> c4(["top 10"]):::estimate - end - classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; - classDef exact fill:#fff,stroke:#5f6368,color:#000; - classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; + input["http_requests_total: range 1m"] --> rate["rate per series"] --> sum["sum by job: Q1"] + input --> values["sum_over_time per series"] --> top["topk by job: Q2"] ``` -**Stage 2: 156 candidates.** Stage 2 picks, for each query, when its state is -computed and whether it is kept. What it can choose depends on the query's -window form from Pass 2: - -| Window form | Physical options | +The common input is a possible sharing point, shown once for readability; +frontend lowering need not perform CSE. + +**Stage 1, Pass 1.** Q1 keeps its exact evaluation. Q2 has an exact sort/limit +option and possible Count-Min+heap or Hydra alternatives. The sketch alternatives +require registered update/readout contracts for these weighted series values and +membership evidence at selection. Family names alone do not qualify a rewrite. +If all three alternatives have semantic contracts, there are initially +`1 × 3 = 3` workload candidates, not three selected plans. + +**Stage 1, Pass 2.** Add an alternative sharing the raw range input. Window +composition adds further alternatives only for operations with valid update and +merge contracts. Q1 stays on raw samples in this example: +[`rate()`](https://prometheus.io/docs/prometheus/latest/querying/functions/#rate) +handles counter resets and extrapolation at the full window boundaries; six +10-s rates cannot simply be merged into one 1-min rate. A future exact-state +rewrite must retain enough ordered boundary/reset information and prove the same +result before becoming an alternative. + +For Q2, compare a direct 1-min computation with active sliding windows or aligned +10-s tumbling windows. A mergeable frequency sketch does not automatically make +its candidate-key heap mergeable or preserve top-k membership. A tumbling option +requires a contract for the complete state and readout, including labels, absent +series and numerical behavior. Undefined merges are excluded; empirical accuracy +assessment is not a substitute for defined semantics. + +**Stage 2.** Each admitted Q2 window form has these physical alternatives: + +| Logical form | Physical alternatives | |---|---| -| None | **Raw:** rebuild the state from the last 1 min of raw samples at every refresh. It cannot be materialized: the window slides every 10 s, and most summaries cannot drop old data. | -| Sliding window | **Sliding, ingestion time:** insert each sample into the 6 active windows as it arrives (materialized at ingestion time). **Sliding, query time:** at each refresh, insert the last 10 s of raw samples into the 6 kept active windows (materialized at query time). Building the window from raw samples at query time without keeping it is the same plan as Raw. | -| 10-s tumbling windows | **Tumbling, ingestion time:** build each tumbling window as samples arrive and keep the last 6. **Tumbling, query time:** at each refresh, build only the newest tumbling window from the last 10 s of raw samples and reuse the 5 kept ones. **Tumbling, rebuilt:** at each refresh, build all 6 tumbling windows from the last 1 min of raw samples, merge them, and discard them (not materialized). | - -```mermaid -flowchart LR - in[("http_requests_total
samples")]:::data +| Direct 1-min computation | Rebuild at each refresh; no cross-refresh state reuse. | +| Active sliding windows | Retain and update at ingestion time, or at query time from newly available raw samples. | +| Aligned tumbling windows | Retain at ingestion time, retain at query time, or rebuild the required buckets per evaluation. | - subgraph R["Raw"] - direction LR - subgraph RQ["Query time, every 10 s"] - r1["last 1 min"]:::exact --> r2["build state"]:::summary --> r3(["answer"]):::estimate - end - end +If all forms qualify, that is `1 + 2 + 3 = 6` options for a Q2 family before +combining with input-sharing and parameter choices. This is illustrative +branching, not a fixed total: shared inputs must have compatible evaluation times, +phases and fetched intervals, and duplicate physical plans count once. - subgraph SI["Sliding, ingestion time"] - direction LR - subgraph SII["Ingestion time"] - s1["update 6 active
windows per sample"]:::summary - end - subgraph SIQ["Query time, every 10 s"] - s2(["answer from the
completed window"]):::estimate - end - s1 --> s2 - end - - subgraph SQ["Sliding, query time"] - direction LR - subgraph SQQ["Query time, every 10 s"] - t1["last 10 s"]:::exact --> t2["update 6 kept
active windows"]:::summary --> t3(["answer from the
completed window"]):::estimate - end - end - - subgraph TI["Tumbling, ingestion time"] - direction LR - subgraph TII["Ingestion time"] - u1["build 10-s window
keep last 6"]:::summary - end - subgraph TIQ["Query time, every 10 s"] - u2["merge 6"]:::exact --> u3(["answer"]):::estimate - end - u1 --> u2 - end - - subgraph TQ["Tumbling, query time"] - direction LR - subgraph TQQ["Query time, every 10 s"] - v1["last 10 s"]:::exact --> v2["build newest
10-s window"]:::summary --> v3["merge with
5 kept"]:::exact --> v4(["answer"]):::estimate - end - end - - subgraph TR["Tumbling, rebuilt"] - direction LR - subgraph TRQ["Query time, every 10 s"] - w1["last 1 min"]:::exact --> w2["build 6
10-s windows"]:::summary --> w3["merge 6"]:::exact --> w4(["answer"]):::estimate - end - end - - in --> r1 - in --> s1 - in --> t1 - in --> u1 - in --> v1 - in --> w1 - classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; - classDef exact fill:#fff,stroke:#5f6368,color:#000; - classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; - classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; -``` - -Across the three window forms, each query has 1 + 2 + 3 = 6 physical options, -so each Q2 option gives 6 × 6 = 36 combinations. The shared-input variant only -changes the plan when both queries read raw samples at query time, which 4 of -the 6 options do (all except the two ingestion-time ones). That adds -4 × 4 = 16 shared-input plans, for 52 plans per Q2 option and -3 × 52 = 156 physical candidates. - -**Stage 3: 1 plan.** Selection first rejects invalid candidates. The -deployment's accuracy model checks the summary candidates against ε = 0.01, -δ = 0.001, including the approximate heap merge of Count-Min with tumbling -windows. Its cost model estimates Q2's latency against the 100 ms bound; for -example, an exact top 10 rebuilt from one million series at every refresh may -miss it. Among the rest, selection picks the cheapest plan for the whole -workload. Typical winners: - -* **Hydra for Q2, with both queries on tumbling windows at ingestion time**, - when there are many small jobs: each sample updates one window, and each - refresh merges 6 small summaries. -* **Count-Min + heap for Q2 on a sliding window at ingestion time**, with Q1 - on tumbling windows, when there are a few large jobs: the heaps do not merge - cleanly, so updating 6 active windows per sample is worth it. -* **Query-time variants** of either, when ingestion-time work is expensive. -* **Raw with a shared input** for both queries, when storage is expensive and - raw data is available at query time. - -The other candidates, such as tumbling windows rebuilt at every refresh, stay -in the candidate set because they are valid, and selection rules them out on -cost. +**Stage 3.** Evaluate each complete workload plan. Reject Q2 candidates without +both score and membership evidence or exceeding 100 ms. Compare total costs of +the remaining raw, maintained and rebuilt variants; Q1's exact computation remains +part of every cost. Return the cheapest qualifying plan, or the no-feasible-plan +outcome. No winner is implied without deployment model results. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 @@ -594,14 +366,14 @@ source IPs over the last minute. -- Q1: Distinct(src_ip) SELECT COUNT(DISTINCT src_ip) FROM flows -WHERE ts >= now() - INTERVAL '1 minute'; +WHERE ts > now() - INTERVAL '1 minute' AND ts <= now(); -- Q2: Entropy(src_ip) SELECT -SUM(p * LN(p)) FROM ( SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows - WHERE ts >= now() - INTERVAL '1 minute' + WHERE ts > now() - INTERVAL '1 minute' AND ts <= now() GROUP BY src_ip ); @@ -610,7 +382,7 @@ SELECT SQRT(SUM(c * c)) FROM ( SELECT src_ip, COUNT(*) AS c FROM flows - WHERE ts >= now() - INTERVAL '1 minute' + WHERE ts > now() - INTERVAL '1 minute' AND ts <= now() GROUP BY src_ip ); ``` @@ -628,21 +400,24 @@ The data workload differs from the shared one in two fields: | `input_cardinality` | 10,000,000 distinct source IPs | | `data_ingestion_interval` | not needed for SQL | -**Pass 1.** Rewrite rules recognize the three computations, and each gets its +Assume non-null `src_ip`, a nonempty window and arithmetic that does not overflow. +All queries use the same bound evaluation time. Empty-input and NULL cases require +their own semantics-preserving lowering; this example does not infer them. + +**Pass 1.** Given registered contracts for these computations, each gets its local candidates from the Pass 1 table: exact, a specialized summary, or UnivMon. Combined, that is 3 × 3 × 3 = 27 workload candidates. **Pass 2.** All three computations have the same summary input data (`src_ip` from `flows`, no other filter) and the same 1-min window. UnivMon supports all three estimates, so the summary-capability rule adds a shared candidate: **one -UnivMon build node feeding three estimation nodes**. It must be sized for the strictest -requirement, ε = 0.01. The independent candidates are kept as well. Pass 2 -also adds candidates where only two of the three share a UnivMon and the third -keeps any of its own 3 options (3 pairs × 3 = 9), so stage 1 outputs -27 + 1 + 9 = 37 candidates. The figure shows the all-three case. The -window-composition rule would also add sliding-window and 10-s tumbling-window -variants, exactly as in Example 1; they are left out here to keep the focus on -the summary-capability rule. +UnivMon build node feeding three estimation nodes**. Physical configuration must +satisfy cardinality, entropy and L2 requirements separately; choosing the smallest +ε is not sufficient. Pass 2 also adds candidates where two queries share a +UnivMon and the third keeps any of its own 3 options (3 pairs × 3 = 9), so stage 1 outputs +27 + 1 + 9 = 37 structural candidates, before parameter choices and feasibility +assessment, assuming all illustrated contracts are available. The figure shows +the all-three case; window alternatives are omitted to isolate this sharing rule. ```mermaid flowchart LR @@ -679,7 +454,7 @@ flowchart LR subgraph P2["Stage 1, Pass 2 · summary-capability rule adds a shared candidate"] direction LR - u["one UnivMon
sized for ε = 0.01"]:::summary + u["one UnivMon
three metric requirements"]:::summary u --> rd(["distinct count"]):::estimate u --> re(["entropy"]):::estimate u --> rl(["L2 norm"]):::estimate @@ -702,25 +477,22 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -The three UnivMon options from Pass 1 (dashed arrows) are merged by Pass 2 into -one shared UnivMon. The independent candidates are kept, so stage 1 outputs -both kinds. - -**Stage 2.** The same options as in Example 1 apply: a summary without a -window summary is rebuilt from raw samples at every refresh, while a -sliding-window or tumbling-window UnivMon can be kept from ingestion time or -from query time (and tumbling windows can also be rebuilt). Because the -dashboard repeats over arriving data, ingestion-time window summaries are -usually cheapest. UnivMon merges exactly, so tumbling windows suit it. +**Stage 2.** Enumerate retained or rebuilt implementations as in Example 1. +Tumbling-window options require a registered merge contract, including compatible +configuration and randomness; merging state does not make estimates exact. -**Stage 3.** Selection compares one UnivMon sized for ε = 0.01 against three -separate summaries, each sized for its own requirement. The shared candidate -usually wins because each flow record updates one summary instead of three. +**Stage 3.** Compare shared configurations that meet all three requirements with +independent summaries sized for their own requirements. Updating one structure +may save work, but its required size, readout costs and maintenance determine +whether sharing wins. ### Example 3: Aggregation over windows — the window-composition rule in Pass 2 -This example has two workload patterns that both lead to a shared window -summary. +The two patterns below share window state across queries or evaluations. +`quantile_over_time` operates per series: each depicted KLL or window structure +is instantiated per series, preserving its labels; it never mixes different +series into one quantile. Assume finite float samples and nonempty windows for +the illustrated readouts. **Pattern A: a batch of sub-interval queries over historical data.** An analyst submits a batch of p99 latency reports over different historical intervals, @@ -749,15 +521,16 @@ of data at rest, plus data still arriving. mergeable. The window-composition rule adds a shared candidate: **one Exponential Histogram over [T − 5 y, T], with a KLL per EH bucket**, with one merge and estimation node per query that merges the EH buckets covering - its interval. The independent candidates are kept. + its interval. This is conditional on a defined two-boundary interval contract + and composed rank-error bound; without them, only the aligned tumbling option + below is justified. -Yearly tumbling windows would also work here, since every interval is a whole -number of years; the Exponential Histogram is shown because it also handles -intervals that are not. Counted at the workload level, Pass 1 gives each query +Tumbling buckets of 365 days, aligned to T, cover these intervals without +boundary approximation. EH is an additional proposal for nonaligned intervals, +not an automatic consequence of KLL mergeability. Pass 1 gives each query 2 options (exact or KLL), so 2⁵ = 32 candidates, and Pass 2 adds one candidate -for every way of grouping two or more queries onto shared window summaries. -That quickly reaches hundreds of candidates, which is why candidate sets may be -enumerated lazily. +for each admitted grouping of two or more queries onto shared window summaries. +The total depends on admitted window contracts and parameter choices. The five query intervals overlap, and all lie inside the last five years: @@ -774,8 +547,7 @@ gantt q5 · [3y] offset 2y :q5, 2021, 2024 ``` -The shared candidate replaces five KLL sketches with one Exponential Histogram -and a merge and estimation node per query: +The conditional EH alternative has a merge and estimation node per query: ```mermaid flowchart LR @@ -798,13 +570,14 @@ over the last 5 min, refreshed every minute. |---|---|---|---|---|---| | `quantile_over_time(0.99, latency_ms[5m])` | every 1 min | 5 m | evaluation time | ε = 0.01, δ = 0.01 | ≤ 200 ms | -* **Pass 1.** One KLL over 5 min for each evaluation. +* **Pass 1.** Exact quantile or one KLL over 5 min for each evaluation. * **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The window-composition rule adds two shared candidates for each Pass 1 option: a **sliding window**, where each sample updates the 5 active 5-min windows, and **1-min tumbling windows**, where each evaluation merges the latest 5. - With exact and KLL from Pass 1, each in 3 window forms (none, sliding, - tumbling), stage 1 outputs 2 × 3 = 6 candidates. + With a valid merge contract for both exact state (for example, retaining + samples) and KLL, this illustrates 2 × 3 = 6 structural alternatives before + parameter choices. A compact exact-quantile accumulator is not assumed. With 1-min tumbling windows, each evaluation merges five of them, and consecutive evaluations share four: @@ -835,8 +608,8 @@ Example 4 shows how stage 2 decides whether to store these window summaries. Stage 2 takes the shared window summaries from Example 3 and decides whether to materialize them. That choice is driven by the workload's `recurrence`, `predictability` and `data_workload.arrival`. This example shows only the -physical candidates of the shared logical candidate; every other logical -candidate from Example 3 gets its own physical candidates the same way. +physical alternatives of the shared logical candidate. The conditional EH +case assumes the interval/error contract required in Example 3 has been supplied. **Pattern A (sub-interval batch).** @@ -844,7 +617,7 @@ candidate from Example 3 gets its own physical candidates the same way. |---|---|---| | A1 | The Exponential Histogram, at query time | At query time, when the batch runs at T; read by all five queries, then discarded | | A2 | The Exponential Histogram, at ingestion time | At ingestion time, with each new sample; old data backfilled once | -| A3 | Nothing | At query time, once per query: each of the five queries rebuilds it for itself and discards it | +| A3 | Nothing | Physical planning explicitly duplicates the producer: each query rebuilds and discards its own state, if reevaluation preserves its contract | ```mermaid flowchart LR @@ -877,15 +650,12 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -Which candidate wins depends on the workload: - -* **As given** (`invocations: 1`, `ad_hoc`): selection picks A1. A2 would - maintain the histogram for years only to serve one batch, and A3 builds it - five times instead of once. -* **Repeated monthly and `Predictable { known_at }`:** A2 can win, because its - maintenance cost is shared by many batches. -* **Data `"at_rest"`:** A2 is not generated, because there is no ingestion to - maintain the histogram. +For the one-off ad hoc batch, A1 avoids repeated builds; selection still compares +its storage and readout costs against A3. A2 is eligible only if advance notice or +existing state and a feasible backfill schedule make it available by T. It cannot +retroactively maintain five years of history. A recurring predictable workload +may amortize that startup cost. For data entirely at rest there is no ongoing +ingestion-maintenance option; query-time builds remain available. **Pattern B (overlapping windows, repeating).** @@ -926,22 +696,10 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -This table covers the tumbling-window candidate. The query repeats every -minute and the data is continuously ingesting, so B1 builds each tumbling KLL -once and reuses it in five evaluations, while B2 rescans raw data every time. -B3 also builds each tumbling KLL once, but at query time, so it needs raw data -at query time and adds the newest build to each evaluation's latency. -Selection usually picks B1. B3 can win when ingestion-time work is expensive, -and B2 only when storage is expensive and raw data is available at query time. - -The sliding-window candidate from Example 3 gets its own physical candidates -the same way: kept from ingestion time (each sample updates the 5 active -windows) or from query time (each evaluation inserts the last minute of raw -samples into the 5 kept windows). It does more ingestion work than B1 but -needs no merge, so it wins only for a summary that merges poorly. - -**What this shows.** The same logical candidate (one shared Exponential -Histogram, or 1-min tumbling KLL windows) yields different physical plans depending -only on recurrence, predictability and data arrival. This is why window-summary -replacement happens in logical planning, while materialization is decided -separately in physical planning. +B1 shifts bucket construction out of query latency; B2 trades retention for +repeated scans; B3 adds only the newest bucket build in steady state. B3's first +execution must initialize all five buckets, and a missed evaluation may require +several new buckets. Retention must cover the last consumer, not expire a bucket +before its final read. Selection compares these costs under the same latency and +accuracy requirements. Sliding-window variants use the materialization choices +already listed in Example 1. From 8db2c17ed5e59ff564370f698414f08f73c04ea1 Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 21:57:10 +0000 Subject: [PATCH 51/56] docs: restore detailed planning stages diagram --- .../design_docs/proposals/planner-layering.md | 91 +++++++++++++++---- 1 file changed, 73 insertions(+), 18 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 8697eddf..134ddbc7 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -14,27 +14,82 @@ selected plan. Planning and selection remain inside ASAPPlanner. ## Stages +In the diagram, × means the Cartesian product: each stage combines every option along one dimension with every option along the others. + ```text -Query workload + Data workload + Deployment inputs - │ - ▼ - 0. Language frontends - CandidateLogicalDAGs - │ - ▼ - 1. Logical ASAP-aware optimization - CandidateLogicalASAPDAGs - │ - ▼ - 2. Physical ASAP-aware optimization - CandidatePhysicalASAPDAGs + Query workload + (PromQL / SQL / MetricsQL, + query recurrence, + accuracy requirements, + latency requirements) + + + Data workload + (streaming vs. data at rest, + data distribution, + cardinality) + + + Deployment inputs + (cost model, + accuracy model, + deployment capabilities) │ ▼ - 3. Plan selection - one PhysicalASAPDAG, or failure reasons - │ success - ▼ - 4. Execution by the deployment +┌────────────────────────────── ASAPPlanner ──────────────────────────────┐ +│ │ +│ 0. Language-specific frontends │ +│ Parse and convert source-language queries into a common logical │ +│ representation. Reject unsupported query expressions. │ +│ │ +│ Output: CandidateLogicalDAGs │ +│ │ +│ │ │ +│ ▼ │ +│ Logical planning — what to compute │ +│ │ +│ 1. Logical ASAP-aware optimization │ +│ Explore semantically equivalent and legal logical candidates: │ +│ │ +│ summary families │ +│ × query rewrites │ +│ × sharing one summary across multiple computations │ +│ │ +│ Output: CandidateLogicalASAPDAGs │ +│ │ +│ │ │ +│ ▼ │ +│ Physical planning — how to compute │ +│ │ +│ 2. Physical ASAP-aware optimization │ +│ Explore executable implementations of each logical candidate: │ +│ │ +│ materialization decisions │ +│ × physical operator implementations │ +│ × parallelism and partitioning │ +│ × resource management │ +│ │ +│ Output: CandidatePhysicalASAPDAGs │ +│ │ +│ │ │ +│ ▼ │ +│ 3. Plan selection │ +│ Evaluate complete physical candidates using the deployment's │ +│ empirical cost and accuracy models. Reject candidates that violate │ +│ accuracy, latency, or capability constraints. │ +│ │ +│ Choose the cheapest valid plan for the whole workload. │ +│ │ +└────────────────────────────────┬───────────────────────────────────────┘ + │ + ▼ + one selected PhysicalASAPDAG + │ + ▼ +┌────────────────────────────── Deployment ───────────────────────────────┐ +│ │ +│ 4. Execution │ +│ deployment executes the selected DAG (plan). │ +│ │ +└────────────────────────────────────────────────────────────────────────┘ ``` The planner receives three groups of inputs: From 9ca78c9180cfbe39b9ee95f95338dc65416ba69b Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 22:03:13 +0000 Subject: [PATCH 52/56] docs: align stage diagram with planning contracts --- .../design_docs/proposals/planner-layering.md | 40 +++++++++++++------ 1 file changed, 28 insertions(+), 12 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 134ddbc7..655b0379 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -9,12 +9,14 @@ ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/ inputs (concrete deployment-input types remain a follow-up), and returns the lowest-cost feasible physical plan within the explored search space, or an explanation that no candidate qualifies. It decides what is computed, how -it is computed, and which plan is best. The deployment supplies cost/accuracy models and capabilities, then executes the -selected plan. Planning and selection remain inside ASAPPlanner. +it is computed, and which plan is best. The deployment supplies cost/accuracy +models and capabilities, then executes the selected plan. Planning and selection remain inside ASAPPlanner. ## Stages -In the diagram, × means the Cartesian product: each stage combines every option along one dimension with every option along the others. +In the diagram, × denotes exploring combinations across the listed dimensions. +Each combination must satisfy the stage contracts below; incompatible combinations +are rejected and equivalent candidates may be deduplicated. ```text Query workload @@ -47,7 +49,7 @@ In the diagram, × means the Cartesian product: each stage combines every option │ Logical planning — what to compute │ │ │ │ 1. Logical ASAP-aware optimization │ -│ Explore semantically equivalent and legal logical candidates: │ +│ Explore exact rewrites and permitted approximate alternatives: │ │ │ │ summary families │ │ × query rewrites │ @@ -60,12 +62,12 @@ In the diagram, × means the Cartesian product: each stage combines every option │ Physical planning — how to compute │ │ │ │ 2. Physical ASAP-aware optimization │ -│ Explore executable implementations of each logical candidate: │ +│ Specify complete implementations of each logical candidate: │ │ │ │ materialization decisions │ │ × physical operator implementations │ -│ × parallelism and partitioning │ -│ × resource management │ +│ × parallelism and partitioning (design TODO) │ +│ × resource management (design TODO) │ │ │ │ Output: CandidatePhysicalASAPDAGs │ │ │ @@ -76,7 +78,7 @@ In the diagram, × means the Cartesian product: each stage combines every option │ empirical cost and accuracy models. Reject candidates that violate │ │ accuracy, latency, or capability constraints. │ │ │ -│ Choose the cheapest valid plan for the whole workload. │ +│ Choose the cheapest feasible workload plan among candidates. │ │ │ └────────────────────────────────┬───────────────────────────────────────┘ │ @@ -92,6 +94,12 @@ In the diagram, × means the Cartesian product: each stage combines every option └────────────────────────────────────────────────────────────────────────┘ ``` +The diagram shows the successful selection path. If no candidate qualifies, +selection returns rejection reasons and execution does not begin. Physical +candidates have complete execution decisions; deployment feasibility is still +checked at selection. Dimensions marked TODO belong to physical planning, but +their detailed contracts are outside this proposal. + The planner receives three groups of inputs: | Input | Contents | @@ -158,7 +166,12 @@ evaluation timing and missing-data semantics. Logical optimization runs in two passes. Pass 1 generates candidates for each computation on its own; Pass 2 finds candidates that share computation across -sub-DAGs and queries. Physical planning assigns materialization and execution timing. +sub-DAGs and queries. Exact rewrites preserve source-language results; summary +replacements may approximate the requested result under its accuracy requirement. +They must preserve the defined input, grouping, window and result contracts. +Logical planning records the approximation obligations; selection establishes +whether a concrete physical candidate meets them using deployment models. +Physical planning assigns materialization and execution timing. #### Pass 1: Local candidate generation @@ -234,9 +247,12 @@ Sharing adds alternatives; independent plans remain available to selection. ### 2. Physical ASAP-aware optimization -Physical optimization turns each logical candidate into executable candidates. -It makes two ASAP-specific decisions, described below. Parallelism, partitioning -and resource management are TODO. +Physical optimization gives each logical candidate concrete implementations, +parameters, materialization and execution timing. Its output is structurally +complete, but has not yet passed deployment capability, accuracy and latency +checks. Selection performs those checks. The two decisions detailed below are +materialization and operator implementation; parallelism, partitioning and resource +management remain physical-planning responsibilities with contracts still TODO. #### Materialization From 0afde73aada683e482cc0e194dfe06d2b5fd59ea Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 22:09:51 +0000 Subject: [PATCH 53/56] docs: restore planner layering proposal to 8c29c2e3 --- .../design_docs/proposals/planner-layering.md | 709 +++++++++++------- 1 file changed, 440 insertions(+), 269 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index 655b0379..c2de34c7 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -3,20 +3,20 @@ Status: proposal. Audience: designers and developers of ASAPPlanner and of deployments such as ASAPQuery-backend. +Read Stages for the overview, the stage sections for the rules, and the +examples for why the rules are needed. + ## Goal ASAPPlanner takes a [query workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs), a [data workload](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs#L531) and the deployment's -inputs (concrete deployment-input types remain a follow-up), and returns the -lowest-cost feasible physical plan within the explored search space, or an -explanation that no candidate qualifies. It decides what is computed, how -it is computed, and which plan is best. The deployment supplies cost/accuracy -models and capabilities, then executes the selected plan. Planning and selection remain inside ASAPPlanner. +inputs (TODO: define this data structure in a follow-up PR), and returns one optimal physical plan. It decides what is computed, how +it is computed, and which plan is best. The deployment only supplies inputs and +executes the plan: it provides its empirical cost model, empirical accuracy +model and capabilities, but never plans queries or selects plans. ## Stages -In the diagram, × denotes exploring combinations across the listed dimensions. -Each combination must satisfy the stage contracts below; incompatible combinations -are rejected and equivalent candidates may be deduplicated. +In the diagram, × means the Cartesian product: each stage combines every option along one dimension with every option along the others. ```text Query workload @@ -49,7 +49,7 @@ are rejected and equivalent candidates may be deduplicated. │ Logical planning — what to compute │ │ │ │ 1. Logical ASAP-aware optimization │ -│ Explore exact rewrites and permitted approximate alternatives: │ +│ Explore semantically equivalent and legal logical candidates: │ │ │ │ summary families │ │ × query rewrites │ @@ -62,12 +62,12 @@ are rejected and equivalent candidates may be deduplicated. │ Physical planning — how to compute │ │ │ │ 2. Physical ASAP-aware optimization │ -│ Specify complete implementations of each logical candidate: │ +│ Explore executable implementations of each logical candidate: │ │ │ │ materialization decisions │ │ × physical operator implementations │ -│ × parallelism and partitioning (design TODO) │ -│ × resource management (design TODO) │ +│ × parallelism and partitioning │ +│ × resource management │ │ │ │ Output: CandidatePhysicalASAPDAGs │ │ │ @@ -78,7 +78,7 @@ are rejected and equivalent candidates may be deduplicated. │ empirical cost and accuracy models. Reject candidates that violate │ │ accuracy, latency, or capability constraints. │ │ │ -│ Choose the cheapest feasible workload plan among candidates. │ +│ Choose the cheapest valid plan for the whole workload. │ │ │ └────────────────────────────────┬───────────────────────────────────────┘ │ @@ -94,12 +94,6 @@ are rejected and equivalent candidates may be deduplicated. └────────────────────────────────────────────────────────────────────────┘ ``` -The diagram shows the successful selection path. If no candidate qualifies, -selection returns rejection reasons and execution does not begin. Physical -candidates have complete execution decisions; deployment feasibility is still -checked at selection. Dimensions marked TODO belong to physical planning, but -their detailed contracts are outside this proposal. - The planner receives three groups of inputs: | Input | Contents | @@ -110,54 +104,34 @@ The planner receives three groups of inputs: ## Stages and their decisions -Stages 0–2 retain alternatives generated by the supported rewrite rules and -parameter domains; stage 3 selects a plan. This is not enumeration of every -possible equivalent program. Candidate sets may be represented lazily. Early -rejection requires a structural or semantic proof, such as a family lacking a -required operation; unknown empirical accuracy is deferred to selection, not -assumed acceptable. Rejections carry reasons. Valid alternatives are not removed -on estimated cost before stage 3. +Stages 0 to 2 each output a candidate set holding every semantically equivalent +and legal candidate DAG of that stage; stage 3 is the only step that chooses +one candidate DAG as output. Candidate sets are internal to ASAPPlanner and may +be shared or enumerated lazily. A stage may prune a candidate early only when +it is provably invalid (for example, a summary family that cannot meet the +query's accuracy target), and every rejected candidate carries a reason. -Each candidate covers the **whole workload**, preserving the mapping from query -identities to their results. Alternatives combine only when their dependencies, -input contracts and evaluation contexts are compatible. Equivalent generated -candidates may be deduplicated, so counts need not grow monotonically. +A candidate is a DAG for the **whole workload**, not for one query. Each stage +combines its choices for every sub-DAG with the candidates it receives (the × +in the diagram), so the candidate set grows from stage to stage until +selection picks one. Example 1 traces this growth step by step. | Stage | Input | Decides | Output | |---|---|---|---| | 0. Frontends | `query`, `language` | Parse and lower to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | | 1. Logical ASAP-aware optimization | Logical DAGs; accuracy requirements, `time_selection`, repetition interval | Summary replacement (Pass 1); ASAP-aware CSE (Pass 2) | `CandidateLogicalASAPDAGs` | | 2. Physical ASAP-aware optimization | Logical ASAP DAGs; `recurrence`, `predictability`, `DataWorkload` | Materialization; physical operators; parallelism and resources (TODO) | `CandidatePhysicalASAPDAGs` | -| 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG`, or no feasible candidate with reasons | +| 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG` | | 4. Execution (deployment) | The selected `PhysicalASAPDAG` | Run ingestion, storage and query-time computation | Query results | -### Proposed stage interfaces - -These are design contracts, not new Rust API declarations. The stage names below -refer to sets of workload plans; they do not require separate operator enums. -Use the common `OperatorNode` / `Operator` / `ScalarExpr` model proposed in -[#511](https://github.com/ProjectASAP/ASAPPlanner/pull/511). That proposal owns node -structure and validation; this one owns planning responsibilities. - -| Interface value | Required contents and invariant | -|---|---| -| `CandidateLogicalDAGs` | Resolved ordinary operator graphs and scalar expressions; a result binding for every workload query; source-language evaluation contexts, schemas and requirements. No ASAP choices. | -| `CandidateLogicalASAPDAGs` | The same query bindings, plus summary families, update/readout semantics, window composition and logical sharing. Retain each consumer's accuracy requirement and any remaining family-parameter alternatives; execution timing may be unassigned. | -| `CandidatePhysicalASAPDAGs` | Concrete implementations and summary parameters; assigned execution timing; materialization lifetimes, startup/backfill and raw-data dependencies. Each graph is structurally executable subject to deployment validation; empirical accuracy, latency and capability feasibility are checked at selection. | -| Selection result | One complete `PhysicalASAPDAG` with its evaluated cost and per-query feasibility evidence, or a no-feasible-candidate outcome with rejection reasons. Unknown required accuracy/capability evidence is not success. | - -Logical planning chooses a family and its semantic parameters (such as requested -quantile). Physical planning enumerates concrete size/precision configurations -within the supported parameter domain. Selection evaluates those configurations -with deployment models; it does not silently resize a selected plan. A search -limit must be reported: the result is optimal only among the explored candidates. -Concrete containers, deployment-model signatures and diagnostic types remain TODO. +The sections below describe each stage. ### 0. Language-specific frontends -The frontend lowers each query and assembles workload `LogicalDAG` candidates. -Nodes represent selectors, transformations, aggregations, grouping and windows, -without ASAP choices. Constructs that cannot be represented faithfully are rejected. +The frontend converts each query into a `LogicalDAG`. Nodes represent logical +query operations, including selectors, transformations, aggregations, grouping +and window semantics. They contain no ASAP summary choices. A construct that cannot be represented +faithfully is rejected. The frontend preserves source-language behavior, including series identity, evaluation timing and missing-data semantics. @@ -166,20 +140,16 @@ evaluation timing and missing-data semantics. Logical optimization runs in two passes. Pass 1 generates candidates for each computation on its own; Pass 2 finds candidates that share computation across -sub-DAGs and queries. Exact rewrites preserve source-language results; summary -replacements may approximate the requested result under its accuracy requirement. -They must preserve the defined input, grouping, window and result contracts. -Logical planning records the approximation obligations; selection establishes -whether a concrete physical candidate meets them using deployment models. -Physical planning assigns materialization and execution timing. +sub-DAGs and queries. Materialization and execution +placement are decided in later stages. #### Pass 1: Local candidate generation For each eligible sub-DAG, Pass 1 identifies its computation semantics, applies -registered rewrite rules, and retains alternatives whose semantic contracts -are established and whose accuracy is not provably infeasible. +rewrite rules, and generates every candidate that is not provably unable to +meet its accuracy requirement. -Illustrative families, admitted only with registered semantic contracts: +Example for summary candidates: | Original computation | Local candidates | |---|---| @@ -194,7 +164,7 @@ Each candidate records its input expression, filter, grouping, window, supported estimates and accuracy requirement. Pass 2 uses these to decide whether candidates can share a summary node. -Summary-based candidates distinguish these operations: +A summary-based candidate uses three kinds of summary nodes: * A **summary build node** builds and maintains a summary from input data, for example a KLL sketch over `latency_ms`. @@ -203,15 +173,17 @@ Summary-based candidates distinguish these operations: summaries into a coarser one. * A **summary estimation node** computes an answer from a summary, for example the p99 estimate from a KLL, or the entropy estimate from a UnivMon. -* **Summary subtract** and **summary delete** remain TODO. Merge is also a - capability-gated operation: #511 reserves its payload pending a complete - family-specific contract; the examples here do not establish runtime support. +* **summary subtract node** and **summary delete node** design is TODO. + +One summary build node can feed several estimation nodes, which is what Pass 2 +exploits. #### Pass 2: ASAP-aware common-subexpression elimination -ASAP-aware CSE adds summary-capability and window-composition rules to -identical-expression reuse. Reuse must preserve evaluation context, volatility, -series identity and missing-data behavior, not merely match expression text. +ASAP-aware CSE extends traditional CSE with summary-specific sharing rules. +Computations can share work when they use identical expressions, when one +summary build node supports several estimates, or when one window summary can answer +their overlapping windows. The rules compare computations by their **summary input data**: what a summary for that computation would ingest, namely the data source, the filters, and the key @@ -219,18 +191,38 @@ or value being summarized together with its grouping. The summary input data doe not include the window; the window-composition rule compares windows separately. -A **window** is the time range read by one query evaluation. A **window summary** -organizes state across windows so overlapping evaluations can reuse work: - -| Window organization | Update and readout | Required contract | -|---|---|---| -| Sliding | Keep one summary per active window; update every window containing an arriving sample. Read the completed window without merging. A 5-min window refreshed every minute has 5 active summaries. | Valid per-window updates and bounded state lifetime; mergeability is unnecessary. | -| Tumbling | Keep nonoverlapping fixed-length buckets; merge those covering each query window. A 5-min window refreshed every minute can merge 5 one-minute buckets. | Merge semantics for the complete state. Bucket width divides window length and refresh interval, and bucket origin aligns with evaluation boundaries. | -| Exponential Histogram (EH) | Keep nonoverlapping buckets of varying size, merging adjacent buckets as they age; read an interval from its covering buckets. | A defined boundary-bucket rule and error bound, including both boundaries of historical intervals. Coarser old buckets may include data outside the requested interval. | - -The EH+KLL case is conditional: KLL mergeability alone does not bound interval -boundary error. Missing boundary semantics must be defined before candidate -generation; empirical accuracy assessment cannot supply an undefined operation. +The window-composition rule distinguishes the window a query reads from the +window summary that answers it: + +* A **window** is the time range one query evaluation reads, for example the + last 5 min. Consecutive evaluations of a repeating query read overlapping + windows. Most summaries cannot remove old data, so one summary cannot simply + slide forward with the window. +* A **window summary** keeps summaries so that many windows can be answered. + Three window summaries are considered for now: + * **Sliding window:** one summary per active window. Each arriving sample is + inserted into every active window that contains it, and each evaluation + reads the window that has just completed, with no merge. For a 5-min + window evaluated every 1 min, 5 windows are active and each sample updates + all 5. It works for any summary, including ones that cannot be merged, at + the cost of more ingestion work and memory. + * **Tumbling window:** back-to-back, non-overlapping windows of one fixed + length, each with one summary. A longer query window is answered by + merging the tumbling windows it covers. The tumbling length must divide + both the query window length and the evaluation interval, so that every + query window starts and ends on a tumbling boundary: a 5-min window + evaluated every 1 min uses 1-min tumbling windows and merges exactly 5 of + them. It needs a mergeable summary. + * **Exponential Histogram (EH):** a sequence of EH buckets that covers a + long history. A query window is answered by merging the EH buckets it + covers. Few EH buckets cover a long history, at the cost that an old + query-window boundary may fall inside an EH bucket and is then + approximate. + * An **EH bucket** is one non-overlapping time range of the history with + one summary of the data in it. Unlike tumbling windows, EH buckets are + not all the same length: they grow with age, so recent data sits in + short EH buckets and older data in longer ones. Adjacent EH buckets are + merged into a longer one as they age. | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| @@ -238,68 +230,96 @@ generation; empirical accuracy assessment cannot supply an undefined operation. | Summary-capability rule | The computations have the same summary input data and the same window, and one summary supports all requested computations and their accuracy requirements. | One summary build node feeding several estimation nodes, e.g. UnivMon → distinct count, entropy, L2 norm. | | Window-composition rule | The computations have the same summary input data, and one window summary can answer the requested windows within their accuracy requirements. | One window summary feeding per-query merge (where needed) and estimation nodes, e.g. a sliding-window or tumbling-window KLL, or an Exponential Histogram with a KLL per EH bucket. | -Examples 2 and 3 illustrate summary-capability and window-composition sharing. -A shared configuration must satisfy **each** consumer's metric and confidence -requirement; ε values for different statistics cannot be ordered as one common -precision target. A second quantile may reuse a KLL if the existing configuration -meets its requirement, but its estimation work still contributes to cost. -Sharing adds alternatives; independent plans remain available to selection. +The examples behind these rules: + +* **Summary-capability rule (Example 2).** One UnivMon over `src_ip` from + `flows` in the last minute serves three queries refreshed every 10 s: + `COUNT(DISTINCT src_ip)`, the entropy of the `src_ip` distribution, and the + L2 norm of per-`src_ip` counts. Each flow record updates the UnivMon once; a + distinct-count, an entropy and an L2 estimation node each compute their + statistic from it. The UnivMon is sized for the strictest of the three accuracy + requirements. +* **Window-composition rule, sliding or tumbling window (Example 3, + Pattern B).** For `quantile_over_time(0.99, latency_ms[5m])` repeated every + minute, one window summary serves every evaluation. With a sliding-window + KLL, each sample updates the 5 active windows, and each evaluation reads the + one that has just completed. With 1-min tumbling-window KLLs, each sample + updates one window, and each evaluation merges the latest 5 with a merge + node; consecutive evaluations share 4 of them. +* **Window-composition rule, Exponential Histogram (Example 3, Pattern A).** + One Exponential Histogram over the last 5 years, with a KLL per EH bucket, + serves the p99 + queries over `[5y]`, `[1y]`, `[1y] offset 1y`, `[1y] offset 2y` and + `[3y] offset 2y`. Each query's merge node merges the EH buckets covering + its interval, and its estimation node computes p99 from the merged KLL. +* **Other quantiles share for free.** One KLL answers every quantile, so adding + `quantile_over_time(0.5, latency_ms[5m])` to the sliding-window dashboard + adds only a p50 estimation node next to the p99 one, reading the same KLL, + with no new summary. + +Rules are defined by each summary family's capabilities and semantic +requirements. A shared summary must meet the strictest accuracy requirement +among its consumers. Applying a rule adds a shared candidate and keeps the +independent candidates, so selection can compare both. ### 2. Physical ASAP-aware optimization -Physical optimization gives each logical candidate concrete implementations, -parameters, materialization and execution timing. Its output is structurally -complete, but has not yet passed deployment capability, accuracy and latency -checks. Selection performs those checks. The two decisions detailed below are -materialization and operator implementation; parallelism, partitioning and resource -management remain physical-planning responsibilities with contracts still TODO. +Physical optimization turns each logical candidate into executable candidates. +It makes two ASAP-specific decisions, described below. Parallelism, partitioning +and resource management are TODO. #### Materialization -Materialization records which intermediate outputs are retained, their lifetime, -and when they are computed: +Materialization decides, for each sub-DAG, whether its output is kept (persistent to disk or kept in memory) across +(batch) query executions, and if so, when it is computed and how long it is +stored. Materialization does not imply ingestion time; a sub-DAG has three +options: -| Choice | Computation and reuse | -|---|---| -| Ingestion-time materialization | Build/update as data arrives; keep output for later query evaluations. | -| Query-time materialization | Build when first needed; retain for other consumers in the batch or subsequent evaluations. | -| No materialization | Execute at query time without a retained result for later reuse. A producer may still feed multiple consumers during that execution. | - -Logical producer identity exposes possible reuse; it does not force storage. -Physical planning may duplicate a producer only when repeated evaluation preserves -semantics, including volatility and randomized-summary guarantees. Such copies -are explicit in the physical graph and are costed separately. A producer executed -once is costed once with all consumer demand. Example 4 contrasts these choices. - -Retain output until its last planned consumer has finished. An ingestion-time node -cannot depend on query-time work; its upstream computation must be available in -the ingestion phase. Every retained-state plan specifies initialization/backfill, -retention, and the handling of missed evaluations or late data. Incremental costs -in the examples describe steady state after initialization. Ingestion-time -maintenance requires data arrival and sufficient advance notice; an ad hoc query -cannot assume that years of maintenance have already occurred. +* **Materialized at ingestion time:** the sub-DAG runs as data arrives, and its + output is stored before any query asks for it. For example, the 1-min + tumbling-window KLLs in Example 4, Pattern B. +* **Materialized at query time:** the sub-DAG runs when a query first needs + it, and its output is stored so that later executions, or other queries in + the same batch, reuse it instead of recomputing it. For example, an + Exponential Histogram built when the batch in Example 4, Pattern A runs and + read by all five of its queries. +* **Not materialized:** the sub-DAG runs at query time for each execution, and + its output is discarded afterward. + +Whichever option is chosen for each sub-DAG, the plan must also satisfy these +constraints: + +* A materialized output is stored for as long as any of its consumers still + needs it. +* Every node upstream of an ingestion-time node also runs at ingestion time. + +The decision depends on the workload's `recurrence` and `predictability` and on +the `DataWorkload`. Typical outcomes: + +* Read by repeated queries while data keeps arriving: materialize at ingestion + time. +* Read by several queries in one batch, or over data at rest: materialize at + query time. +* Read once by an ad hoc query: do not materialize. + +A shared summary is materialized once for all its consumers. See Example 4. #### Physical operator implementation Physical operator implementation lowers every node to physical operators, for -example choosing a sorting implementation for an exact top-k plan or a backend -implementation for each KLL build, merge and readout. Logical build/merge/readout -semantics are already explicit; physical lowering implements those operations. +example TopK as a sort followed by a limit, or a KLL node as summary build, +merge and quantile estimation operators. ### 3. Plan selection Selection rejects every candidate that misses an accuracy target or a latency bound, or that needs a capability the deployment lacks, and then picks the cheapest remaining plan. It is the only stage that uses the deployment's cost -and accuracy models, and the only stage that discards valid candidates. Its -accuracy evidence must use the requested metric; -a nominal family bound or an empirical point estimate alone is not a confidence -guarantee. Cost is evaluated for the whole workload rather than per query, +and accuracy models, and the only stage that discards valid candidates. Accuracy is +estimated by the deployment's accuracy model, not assumed from a summary's +nominal bound. Cost is evaluated for the whole workload rather than per query, which is what lets one shared summary beat several cheaper independent ones: -a physically shared summary is costed once, with the demand of all its consumers. -Cost includes startup/backfill, maintenance, storage and query execution over the -workload horizon. If no candidate satisfies every requirement, return reasons -rather than a best-effort plan that violates the workload contract. +a shared summary is costed once, with the demand of all its consumers. ### 4. Execution @@ -308,40 +328,26 @@ given: it does not choose among summaries or decide what to materialize. ## End-to-end examples -Field names in code font refer to -[`workload.rs`](../../../crates/types/src/workload.rs). `TimeSelection.lookback` -is the window length and `as_of` its end; `as_of: None` uses the evaluation time. -Endpoint inclusion follows the source language: PromQL ranges use -`(as_of − lookback, as_of]`; SQL follows its predicates. PromQL `1y` is 365 days, -not a calendar year. The diagrams use years only as schematic labels. - -The examples reuse `AccuracyTarget::EpsilonDelta` and existing -[`ErrorMetric` / `ResultGuarantee`](../../../crates/types/src/post_asap/guarantee.rs). -Each target needs its result's metric, normalization and failure-probability scope: - -| Result | Metric used in these examples | -|---|---| -| Quantile | `Rank`: normalized rank error, not error in the returned latency value. | -| Distinct count | `Cardinality`: relative count error. | -| Entropy | `AbsoluteValue`: absolute error in nats, matching `LN`. | -| L2 norm | `RelativeValue`: relative norm error; the model must cover the zero case. | -| Top-k | A per-key `Frequency` bound does not establish `TopKMembership`. Selection needs membership evidence as well as score bounds, or a separately specified approximate-membership contract. | - -Here δ applies per query result per evaluation; simultaneous guarantees across -queries or time require an explicit joint failure budget. No independence is -assumed for shared summaries. A model must cover update, merge, boundary selection -and readout errors before a candidate is selectable. +Each example's workload is shown as tables. Field names in code font are the +fields of +[`workload.rs`](https://github.com/ProjectASAP/ASAPPlanner/blob/main/crates/types/src/workload.rs). +Each query reads the event-time window [`as_of` − `lookback`, `as_of`] (fields +of `TimeSelection`). `as_of` is the window's end; `lookback` is its length. An +`as_of` of "evaluation time" means `as_of: None`: the window ends whenever the +query runs, so it moves forward with each evaluation. A fixed `as_of`, such as +T − 1 y, pins the window to a historical interval. Approximate accuracy targets are `EpsilonDelta`: the answer's error +is at most ε with probability at least 1 − δ. | Example | Shows | |---|---| -| 1. A workload through every stage | Candidate generation, physical alternatives and selection | +| 1. Aggregation over dimensions | How the candidate set grows through every stage | | 2. One summary for several computations | Pass 2: summary-capability rule | | 3. Aggregation over windows | Pass 1 and Pass 2: window-composition rule | | 4. Materialization of window summaries | Stage 2: materialization, decoupled from stage 1 | -In the colored diagrams, grey cylinders are inputs, blue boxes build summaries, -and green nodes read estimates. White merge nodes do not imply exact estimates; -readout guarantees include the merge contract. +In the diagrams below, grey cylinders are input data, white boxes are exact +operations and summary merges, blue boxes are summary build nodes, and green +rounded boxes are summary estimation nodes. ### Shared data workload @@ -350,83 +356,234 @@ Unless an example says otherwise, every example uses this data workload: | `DataWorkload` field | Meaning | Value | |---|---|---| | `arrival` | Whether the data is at rest, still arriving, or both | `continuously_ingesting` | -| `data_ingestion_interval` | Sampling cadence, analogous to the Prometheus scrape interval; distinct from instant-selector lookback. | 15 s | +| `data_ingestion_interval` | How often each series delivers one sample (the scrape interval in Prometheus). PromQL uses it as the look-back horizon of instant selectors. | 15 s | | `ingestion_volume` | Total amount of ingested data | unknown | | `ingestion_rate` | Samples arriving per second across all series | about 66,667 samples/s | | `input_cardinality` | Number of distinct series (or keys) | 1,000,000 series | | `distribution` | How samples are spread over keys | `zipf` | With 1,000,000 series each sampled every 15 s, the ingestion rate is -1,000,000 / 15 ≈ 66,667 samples/s. Instant-selector lookback is a separate -query setting (Prometheus defaults to 5 min), not derived from that cadence. -See the [Prometheus selector semantics](https://prometheus.io/docs/prometheus/latest/querying/basics/#staleness). +1,000,000 / 15 ≈ 66,667 samples/s. -### Example 1: A workload through every stage +### Example 1: Aggregation over dimensions — the candidate set through every stage -Two PromQL panels refresh every 10 s over the same 1-min input range: +**Query workload.** Two PromQL dashboard panels over the last minute. The +first needs an exact total; the second tolerates error. -| Query | Accuracy requirement | Latency requirement | -|---|---|---| -| Q1: `sum by (job) (rate(http_requests_total[1m]))` | Exact | None | -| Q2: `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | Score ε = 0.01, δ = 0.001 under `Frequency`, plus `TopKMembership` evidence | ≤ 100 ms | +| Query | Repeats | `lookback` | `as_of` | Accuracy requirement | Latency requirement | +|---|---|---|---|---|---| +| Q1: `sum by (job) (rate(http_requests_total[1m]))` | every 10 s | 1 m | evaluation time | exact | none | +| Q2: `topk by (job) (10, sum_over_time(http_requests_total[1m]))` | every 10 s | 1 m | evaluation time | ε = 0.01, δ = 0.001 | ≤ 100 ms | + +This example follows the workload's candidate set through every stage. Each +candidate covers both queries. + +**Stage 0: 1 candidate.** The frontend lowers both queries into one workload +`LogicalDAG` with no summaries: + +```mermaid +flowchart LR + subgraph Q1["Q1 · sum by (job) (rate(http_requests_total[1m]))"] + direction LR + x1[("http_requests_total")]:::data --> x2["range 1m"]:::exact --> x3["rate"]:::exact --> x4["sum by (job)"]:::exact + end + subgraph Q2["Q2 · topk by (job) (10, sum_over_time(http_requests_total[1m]))"] + direction LR + y1[("http_requests_total")]:::data --> y2["range 1m"]:::exact --> y3["sum_over_time"]:::exact --> y4["topk by (job) (10)"]:::exact + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; +``` -Assume nonnegative finite float samples for Q2. Its sum ranks sampled counter -values; it is not a request-rate ranking. Retain the complete series labels. -The membership requirement is additional to score error; unknown membership -evidence prevents selection of a sketch plan. +**Stage 1, Pass 1: 3 candidates.** Pass 1 finds local options for each query: -**Stage 0.** One workload graph contains both query results: +* **Q1** has one option, the exact per-series rate and per-`job` sum. Its + accuracy requirement is exact, so no summary qualifies. +* **Q2** has three options. The exact one keeps a per-series sum and sorts + within each `job`, which is costly at one million Zipf-distributed series. + The `EpsilonDelta` target also admits a **Count-Min Sketch with a top-*k* + heap per `job`**, and **Hydra over the whole `job` column**, where one sketch + covers every (`job`, series) key. ```mermaid flowchart LR - input["http_requests_total: range 1m"] --> rate["rate per series"] --> sum["sum by job: Q1"] - input --> values["sum_over_time per series"] --> top["topk by job: Q2"] + subgraph C["Q2's three local options"] + direction TB + subgraph E["Exact"] + direction LR + e1[("input")]:::data --> e2["sum_over_time
per series"]:::exact --> e3["sort + limit 10
per job"]:::exact + end + subgraph CM["Count-Min + heap per job"] + direction LR + c1[("input")]:::data --> c2["Count-Min Sketch +
top-10 heap, one per job"]:::summary --> c3(["top 10
per job"]):::estimate + end + subgraph H["Hydra"] + direction LR + h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::estimate + end + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -The common input is a possible sharing point, shown once for readability; -frontend lowering need not perform CSE. - -**Stage 1, Pass 1.** Q1 keeps its exact evaluation. Q2 has an exact sort/limit -option and possible Count-Min+heap or Hydra alternatives. The sketch alternatives -require registered update/readout contracts for these weighted series values and -membership evidence at selection. Family names alone do not qualify a rewrite. -If all three alternatives have semantic contracts, there are initially -`1 × 3 = 3` workload candidates, not three selected plans. - -**Stage 1, Pass 2.** Add an alternative sharing the raw range input. Window -composition adds further alternatives only for operations with valid update and -merge contracts. Q1 stays on raw samples in this example: -[`rate()`](https://prometheus.io/docs/prometheus/latest/querying/functions/#rate) -handles counter resets and extrapolation at the full window boundaries; six -10-s rates cannot simply be merged into one 1-min rate. A future exact-state -rewrite must retain enough ordered boundary/reset information and prove the same -result before becoming an alternative. - -For Q2, compare a direct 1-min computation with active sliding windows or aligned -10-s tumbling windows. A mergeable frequency sketch does not automatically make -its candidate-key heap mergeable or preserve top-k membership. A tumbling option -requires a contract for the complete state and readout, including labels, absent -series and numerical behavior. Undefined merges are excluded; empirical accuracy -assessment is not a substitute for defined semantics. - -**Stage 2.** Each admitted Q2 window form has these physical alternatives: - -| Logical form | Physical alternatives | +Combining them gives 1 × 3 = 3 workload candidates: + +| Logical candidate | Q1 | Q2 | +|---|---|---| +| Exact | exact | exact | +| Count-Min | exact | Count-Min + heap | +| Hydra | exact | Hydra | + +**Stage 1, Pass 2: 54 candidates.** Pass 2 applies two ASAP-aware CSE rules +to each of the 3 candidates, and keeps every original: + +* **Identical-expression rule.** Both queries read the same range selector, + `http_requests_total[1m]`, so Pass 2 adds a variant in which Q1 and Q2 share + one input node. No summary is shared, because Q1 must be exact and no + summary supports both queries. +* **Window-composition rule.** Each query reads a 1-min window every 10 s, so + consecutive evaluations overlap by 50 s. For each query, Pass 2 adds two + variants: a **sliding window**, where each sample updates the 6 active 1-min + windows, and **10-s tumbling windows**, merged 6 at a time at every refresh. + Tumbling windows need a mergeable summary. Q1's rates and sums, Q2's exact + sums and Hydra merge exactly; Count-Min sketches do too, but their top-10 + heaps merge only approximately, so that candidate is kept and the accuracy + model judges it in stage 3. + +Each Pass 1 candidate therefore has 2 input choices (separate or shared) × +3 window forms for Q1 (none, sliding, tumbling) × 3 for Q2, so stage 1 outputs +3 × 2 × 3 × 3 = 54 logical candidates. + +The window-composition rule applied to Q2 in the Hydra candidate: + +```mermaid +flowchart LR + subgraph NO["Hydra · no window summary"] + direction LR + a1[("http_requests_total")]:::data --> a2["range 1m"]:::exact --> a3["Hydra"]:::summary --> a4(["top 10"]):::estimate + end + subgraph SL["Hydra · sliding window"] + direction LR + b1[("http_requests_total")]:::data --> b2["insert each sample into
6 active 1-min windows"]:::summary --> b3(["top 10 from the
completed window"]):::estimate + end + subgraph TU["Hydra · 10-s tumbling windows"] + direction LR + c1[("http_requests_total")]:::data --> c2["Hydra per
10-s window"]:::summary --> c3["merge latest 6"]:::exact --> c4(["top 10"]):::estimate + end + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; +``` + +**Stage 2: 156 candidates.** Stage 2 picks, for each query, when its state is +computed and whether it is kept. What it can choose depends on the query's +window form from Pass 2: + +| Window form | Physical options | |---|---| -| Direct 1-min computation | Rebuild at each refresh; no cross-refresh state reuse. | -| Active sliding windows | Retain and update at ingestion time, or at query time from newly available raw samples. | -| Aligned tumbling windows | Retain at ingestion time, retain at query time, or rebuild the required buckets per evaluation. | +| None | **Raw:** rebuild the state from the last 1 min of raw samples at every refresh. It cannot be materialized: the window slides every 10 s, and most summaries cannot drop old data. | +| Sliding window | **Sliding, ingestion time:** insert each sample into the 6 active windows as it arrives (materialized at ingestion time). **Sliding, query time:** at each refresh, insert the last 10 s of raw samples into the 6 kept active windows (materialized at query time). Building the window from raw samples at query time without keeping it is the same plan as Raw. | +| 10-s tumbling windows | **Tumbling, ingestion time:** build each tumbling window as samples arrive and keep the last 6. **Tumbling, query time:** at each refresh, build only the newest tumbling window from the last 10 s of raw samples and reuse the 5 kept ones. **Tumbling, rebuilt:** at each refresh, build all 6 tumbling windows from the last 1 min of raw samples, merge them, and discard them (not materialized). | -If all forms qualify, that is `1 + 2 + 3 = 6` options for a Q2 family before -combining with input-sharing and parameter choices. This is illustrative -branching, not a fixed total: shared inputs must have compatible evaluation times, -phases and fetched intervals, and duplicate physical plans count once. +```mermaid +flowchart LR + in[("http_requests_total
samples")]:::data -**Stage 3.** Evaluate each complete workload plan. Reject Q2 candidates without -both score and membership evidence or exceeding 100 ms. Compare total costs of -the remaining raw, maintained and rebuilt variants; Q1's exact computation remains -part of every cost. Return the cheapest qualifying plan, or the no-feasible-plan -outcome. No winner is implied without deployment model results. + subgraph R["Raw"] + direction LR + subgraph RQ["Query time, every 10 s"] + r1["last 1 min"]:::exact --> r2["build state"]:::summary --> r3(["answer"]):::estimate + end + end + + subgraph SI["Sliding, ingestion time"] + direction LR + subgraph SII["Ingestion time"] + s1["update 6 active
windows per sample"]:::summary + end + subgraph SIQ["Query time, every 10 s"] + s2(["answer from the
completed window"]):::estimate + end + s1 --> s2 + end + + subgraph SQ["Sliding, query time"] + direction LR + subgraph SQQ["Query time, every 10 s"] + t1["last 10 s"]:::exact --> t2["update 6 kept
active windows"]:::summary --> t3(["answer from the
completed window"]):::estimate + end + end + + subgraph TI["Tumbling, ingestion time"] + direction LR + subgraph TII["Ingestion time"] + u1["build 10-s window
keep last 6"]:::summary + end + subgraph TIQ["Query time, every 10 s"] + u2["merge 6"]:::exact --> u3(["answer"]):::estimate + end + u1 --> u2 + end + + subgraph TQ["Tumbling, query time"] + direction LR + subgraph TQQ["Query time, every 10 s"] + v1["last 10 s"]:::exact --> v2["build newest
10-s window"]:::summary --> v3["merge with
5 kept"]:::exact --> v4(["answer"]):::estimate + end + end + + subgraph TR["Tumbling, rebuilt"] + direction LR + subgraph TRQ["Query time, every 10 s"] + w1["last 1 min"]:::exact --> w2["build 6
10-s windows"]:::summary --> w3["merge 6"]:::exact --> w4(["answer"]):::estimate + end + end + + in --> r1 + in --> s1 + in --> t1 + in --> u1 + in --> v1 + in --> w1 + classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; + classDef exact fill:#fff,stroke:#5f6368,color:#000; + classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; + classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; +``` + +Across the three window forms, each query has 1 + 2 + 3 = 6 physical options, +so each Q2 option gives 6 × 6 = 36 combinations. The shared-input variant only +changes the plan when both queries read raw samples at query time, which 4 of +the 6 options do (all except the two ingestion-time ones). That adds +4 × 4 = 16 shared-input plans, for 52 plans per Q2 option and +3 × 52 = 156 physical candidates. + +**Stage 3: 1 plan.** Selection first rejects invalid candidates. The +deployment's accuracy model checks the summary candidates against ε = 0.01, +δ = 0.001, including the approximate heap merge of Count-Min with tumbling +windows. Its cost model estimates Q2's latency against the 100 ms bound; for +example, an exact top 10 rebuilt from one million series at every refresh may +miss it. Among the rest, selection picks the cheapest plan for the whole +workload. Typical winners: + +* **Hydra for Q2, with both queries on tumbling windows at ingestion time**, + when there are many small jobs: each sample updates one window, and each + refresh merges 6 small summaries. +* **Count-Min + heap for Q2 on a sliding window at ingestion time**, with Q1 + on tumbling windows, when there are a few large jobs: the heaps do not merge + cleanly, so updating 6 active windows per sample is worth it. +* **Query-time variants** of either, when ingestion-time work is expensive. +* **Raw with a shared input** for both queries, when storage is expensive and + raw data is available at query time. + +The other candidates, such as tumbling windows rebuilt at every refresh, stay +in the candidate set because they are valid, and selection rules them out on +cost. ### Example 2: One summary for several computations — the summary-capability rule in Pass 2 @@ -437,14 +594,14 @@ source IPs over the last minute. -- Q1: Distinct(src_ip) SELECT COUNT(DISTINCT src_ip) FROM flows -WHERE ts > now() - INTERVAL '1 minute' AND ts <= now(); +WHERE ts >= now() - INTERVAL '1 minute'; -- Q2: Entropy(src_ip) SELECT -SUM(p * LN(p)) FROM ( SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows - WHERE ts > now() - INTERVAL '1 minute' AND ts <= now() + WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip ); @@ -453,7 +610,7 @@ SELECT SQRT(SUM(c * c)) FROM ( SELECT src_ip, COUNT(*) AS c FROM flows - WHERE ts > now() - INTERVAL '1 minute' AND ts <= now() + WHERE ts >= now() - INTERVAL '1 minute' GROUP BY src_ip ); ``` @@ -471,24 +628,21 @@ The data workload differs from the shared one in two fields: | `input_cardinality` | 10,000,000 distinct source IPs | | `data_ingestion_interval` | not needed for SQL | -Assume non-null `src_ip`, a nonempty window and arithmetic that does not overflow. -All queries use the same bound evaluation time. Empty-input and NULL cases require -their own semantics-preserving lowering; this example does not infer them. - -**Pass 1.** Given registered contracts for these computations, each gets its +**Pass 1.** Rewrite rules recognize the three computations, and each gets its local candidates from the Pass 1 table: exact, a specialized summary, or UnivMon. Combined, that is 3 × 3 × 3 = 27 workload candidates. **Pass 2.** All three computations have the same summary input data (`src_ip` from `flows`, no other filter) and the same 1-min window. UnivMon supports all three estimates, so the summary-capability rule adds a shared candidate: **one -UnivMon build node feeding three estimation nodes**. Physical configuration must -satisfy cardinality, entropy and L2 requirements separately; choosing the smallest -ε is not sufficient. Pass 2 also adds candidates where two queries share a -UnivMon and the third keeps any of its own 3 options (3 pairs × 3 = 9), so stage 1 outputs -27 + 1 + 9 = 37 structural candidates, before parameter choices and feasibility -assessment, assuming all illustrated contracts are available. The figure shows -the all-three case; window alternatives are omitted to isolate this sharing rule. +UnivMon build node feeding three estimation nodes**. It must be sized for the strictest +requirement, ε = 0.01. The independent candidates are kept as well. Pass 2 +also adds candidates where only two of the three share a UnivMon and the third +keeps any of its own 3 options (3 pairs × 3 = 9), so stage 1 outputs +27 + 1 + 9 = 37 candidates. The figure shows the all-three case. The +window-composition rule would also add sliding-window and 10-s tumbling-window +variants, exactly as in Example 1; they are left out here to keep the focus on +the summary-capability rule. ```mermaid flowchart LR @@ -525,7 +679,7 @@ flowchart LR subgraph P2["Stage 1, Pass 2 · summary-capability rule adds a shared candidate"] direction LR - u["one UnivMon
three metric requirements"]:::summary + u["one UnivMon
sized for ε = 0.01"]:::summary u --> rd(["distinct count"]):::estimate u --> re(["entropy"]):::estimate u --> rl(["L2 norm"]):::estimate @@ -548,22 +702,25 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -**Stage 2.** Enumerate retained or rebuilt implementations as in Example 1. -Tumbling-window options require a registered merge contract, including compatible -configuration and randomness; merging state does not make estimates exact. +The three UnivMon options from Pass 1 (dashed arrows) are merged by Pass 2 into +one shared UnivMon. The independent candidates are kept, so stage 1 outputs +both kinds. + +**Stage 2.** The same options as in Example 1 apply: a summary without a +window summary is rebuilt from raw samples at every refresh, while a +sliding-window or tumbling-window UnivMon can be kept from ingestion time or +from query time (and tumbling windows can also be rebuilt). Because the +dashboard repeats over arriving data, ingestion-time window summaries are +usually cheapest. UnivMon merges exactly, so tumbling windows suit it. -**Stage 3.** Compare shared configurations that meet all three requirements with -independent summaries sized for their own requirements. Updating one structure -may save work, but its required size, readout costs and maintenance determine -whether sharing wins. +**Stage 3.** Selection compares one UnivMon sized for ε = 0.01 against three +separate summaries, each sized for its own requirement. The shared candidate +usually wins because each flow record updates one summary instead of three. ### Example 3: Aggregation over windows — the window-composition rule in Pass 2 -The two patterns below share window state across queries or evaluations. -`quantile_over_time` operates per series: each depicted KLL or window structure -is instantiated per series, preserving its labels; it never mixes different -series into one quantile. Assume finite float samples and nonempty windows for -the illustrated readouts. +This example has two workload patterns that both lead to a shared window +summary. **Pattern A: a batch of sub-interval queries over historical data.** An analyst submits a batch of p99 latency reports over different historical intervals, @@ -592,16 +749,15 @@ of data at rest, plus data still arriving. mergeable. The window-composition rule adds a shared candidate: **one Exponential Histogram over [T − 5 y, T], with a KLL per EH bucket**, with one merge and estimation node per query that merges the EH buckets covering - its interval. This is conditional on a defined two-boundary interval contract - and composed rank-error bound; without them, only the aligned tumbling option - below is justified. + its interval. The independent candidates are kept. -Tumbling buckets of 365 days, aligned to T, cover these intervals without -boundary approximation. EH is an additional proposal for nonaligned intervals, -not an automatic consequence of KLL mergeability. Pass 1 gives each query +Yearly tumbling windows would also work here, since every interval is a whole +number of years; the Exponential Histogram is shown because it also handles +intervals that are not. Counted at the workload level, Pass 1 gives each query 2 options (exact or KLL), so 2⁵ = 32 candidates, and Pass 2 adds one candidate -for each admitted grouping of two or more queries onto shared window summaries. -The total depends on admitted window contracts and parameter choices. +for every way of grouping two or more queries onto shared window summaries. +That quickly reaches hundreds of candidates, which is why candidate sets may be +enumerated lazily. The five query intervals overlap, and all lie inside the last five years: @@ -618,7 +774,8 @@ gantt q5 · [3y] offset 2y :q5, 2021, 2024 ``` -The conditional EH alternative has a merge and estimation node per query: +The shared candidate replaces five KLL sketches with one Exponential Histogram +and a merge and estimation node per query: ```mermaid flowchart LR @@ -641,14 +798,13 @@ over the last 5 min, refreshed every minute. |---|---|---|---|---|---| | `quantile_over_time(0.99, latency_ms[5m])` | every 1 min | 5 m | evaluation time | ε = 0.01, δ = 0.01 | ≤ 200 ms | -* **Pass 1.** Exact quantile or one KLL over 5 min for each evaluation. +* **Pass 1.** One KLL over 5 min for each evaluation. * **Pass 2.** Consecutive evaluations overlap by 4 of their 5 minutes. The window-composition rule adds two shared candidates for each Pass 1 option: a **sliding window**, where each sample updates the 5 active 5-min windows, and **1-min tumbling windows**, where each evaluation merges the latest 5. - With a valid merge contract for both exact state (for example, retaining - samples) and KLL, this illustrates 2 × 3 = 6 structural alternatives before - parameter choices. A compact exact-quantile accumulator is not assumed. + With exact and KLL from Pass 1, each in 3 window forms (none, sliding, + tumbling), stage 1 outputs 2 × 3 = 6 candidates. With 1-min tumbling windows, each evaluation merges five of them, and consecutive evaluations share four: @@ -679,8 +835,8 @@ Example 4 shows how stage 2 decides whether to store these window summaries. Stage 2 takes the shared window summaries from Example 3 and decides whether to materialize them. That choice is driven by the workload's `recurrence`, `predictability` and `data_workload.arrival`. This example shows only the -physical alternatives of the shared logical candidate. The conditional EH -case assumes the interval/error contract required in Example 3 has been supplied. +physical candidates of the shared logical candidate; every other logical +candidate from Example 3 gets its own physical candidates the same way. **Pattern A (sub-interval batch).** @@ -688,7 +844,7 @@ case assumes the interval/error contract required in Example 3 has been supplied |---|---|---| | A1 | The Exponential Histogram, at query time | At query time, when the batch runs at T; read by all five queries, then discarded | | A2 | The Exponential Histogram, at ingestion time | At ingestion time, with each new sample; old data backfilled once | -| A3 | Nothing | Physical planning explicitly duplicates the producer: each query rebuilds and discards its own state, if reevaluation preserves its contract | +| A3 | Nothing | At query time, once per query: each of the five queries rebuilds it for itself and discards it | ```mermaid flowchart LR @@ -721,12 +877,15 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -For the one-off ad hoc batch, A1 avoids repeated builds; selection still compares -its storage and readout costs against A3. A2 is eligible only if advance notice or -existing state and a feasible backfill schedule make it available by T. It cannot -retroactively maintain five years of history. A recurring predictable workload -may amortize that startup cost. For data entirely at rest there is no ongoing -ingestion-maintenance option; query-time builds remain available. +Which candidate wins depends on the workload: + +* **As given** (`invocations: 1`, `ad_hoc`): selection picks A1. A2 would + maintain the histogram for years only to serve one batch, and A3 builds it + five times instead of once. +* **Repeated monthly and `Predictable { known_at }`:** A2 can win, because its + maintenance cost is shared by many batches. +* **Data `"at_rest"`:** A2 is not generated, because there is no ingestion to + maintain the histogram. **Pattern B (overlapping windows, repeating).** @@ -767,10 +926,22 @@ flowchart LR classDef estimate fill:#e6f4ea,stroke:#188038,color:#000; ``` -B1 shifts bucket construction out of query latency; B2 trades retention for -repeated scans; B3 adds only the newest bucket build in steady state. B3's first -execution must initialize all five buckets, and a missed evaluation may require -several new buckets. Retention must cover the last consumer, not expire a bucket -before its final read. Selection compares these costs under the same latency and -accuracy requirements. Sliding-window variants use the materialization choices -already listed in Example 1. +This table covers the tumbling-window candidate. The query repeats every +minute and the data is continuously ingesting, so B1 builds each tumbling KLL +once and reuses it in five evaluations, while B2 rescans raw data every time. +B3 also builds each tumbling KLL once, but at query time, so it needs raw data +at query time and adds the newest build to each evaluation's latency. +Selection usually picks B1. B3 can win when ingestion-time work is expensive, +and B2 only when storage is expensive and raw data is available at query time. + +The sliding-window candidate from Example 3 gets its own physical candidates +the same way: kept from ingestion time (each sample updates the 5 active +windows) or from query time (each evaluation inserts the last minute of raw +samples into the 5 kept windows). It does more ingestion work than B1 but +needs no merge, so it wins only for a summary that merges poorly. + +**What this shows.** The same logical candidate (one shared Exponential +Histogram, or 1-min tumbling KLL windows) yields different physical plans depending +only on recurrence, predictability and data arrival. This is why window-summary +replacement happens in logical planning, while materialization is decided +separately in physical planning. From 118b4166c38196f4e933b44097e974a4471e2e49 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:15:07 -0400 Subject: [PATCH 54/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 49 ++++++++++++++----- 1 file changed, 38 insertions(+), 11 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index c2de34c7..a20dbb86 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -107,9 +107,14 @@ The planner receives three groups of inputs: Stages 0 to 2 each output a candidate set holding every semantically equivalent and legal candidate DAG of that stage; stage 3 is the only step that chooses one candidate DAG as output. Candidate sets are internal to ASAPPlanner and may -be shared or enumerated lazily. A stage may prune a candidate early only when -it is provably invalid (for example, a summary family that cannot meet the -query's accuracy target), and every rejected candidate carries a reason. +be shared or enumerated lazily. Two kinds of removal are kept apart: + +* **Pruning** removes an **invalid** candidate. Any stage may prune, but only + when the candidate is provably invalid (for example, a summary family that + cannot meet the query's accuracy target), and every pruned candidate carries + a reason. +* **Selection** discards **valid** candidates. Only stage 3 does this, when it + picks the cheapest one. A candidate is a DAG for the **whole workload**, not for one query. Each stage combines its choices for every sub-DAG with the candidates it receives (the × @@ -224,6 +229,15 @@ window summary that answers it: short EH buckets and older data in longer ones. Adjacent EH buckets are merged into a longer one as they age. + TODO: evaluate other sliding-window frameworks for sketches as further + window summaries, for example: + * [Smooth Histograms for Sliding Windows](https://web.cs.ucla.edu/~rafail/PUBLIC/82.pdf) + (Braverman and Ostrovsky, FOCS 2007), an alternative to EH. + * [MicroscopeSketch: Accurate Sliding Estimation Using Adaptive Zooming](https://yangtonghome.github.io/uploads/MicroscopeSketch_SIGKDD_23_final_paper.pdf) + (Wu et al., KDD 2023). + * [Sliding Sketches: A Framework using Time Zones for Data Stream Processing in Sliding Windows](https://dl.acm.org/doi/10.1145/3394486.3403144) + (Gou et al., KDD 2020). + | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| | Identical-expression rule | The input and computation semantics are identical. | One common computation node serving multiple consumers. | @@ -270,10 +284,18 @@ and resource management are TODO. #### Materialization -Materialization decides, for each sub-DAG, whether its output is kept (persistent to disk or kept in memory) across -(batch) query executions, and if so, when it is computed and how long it is -stored. Materialization does not imply ingestion time; a sub-DAG has three -options: +Materialization decides, for each sub-DAG, whether its output is kept across +(batch) query executions. Stage 2 makes this decision in three steps: + +1. **Materialize or not.** A sub-DAG that is not materialized always runs at + query time and keeps nothing. +2. **If materialized, ingestion time or query time.** This is when the output + is computed. A sub-DAG that runs at ingestion time is therefore always + materialized, because its output must be kept until a query reads it. +3. **Where and how long to store it:** on disk or in memory, and for how long. + The deployment's cost model prices each choice. + +This gives each sub-DAG three options: * **Materialized at ingestion time:** the sub-DAG runs as data arrives, and its output is stored before any query asks for it. For example, the 1-min @@ -315,7 +337,8 @@ merge and quantile estimation operators. Selection rejects every candidate that misses an accuracy target or a latency bound, or that needs a capability the deployment lacks, and then picks the cheapest remaining plan. It is the only stage that uses the deployment's cost -and accuracy models, and the only stage that discards valid candidates. Accuracy is +and accuracy models, and the only stage that discards valid candidates; +earlier stages only prune provably invalid ones. Accuracy is estimated by the deployment's accuracy model, not assumed from a summary's nominal bound. Cost is evaluated for the whole workload rather than per query, which is what lets one shared summary beat several cheaper independent ones: @@ -382,7 +405,7 @@ candidate covers both queries. `LogicalDAG` with no summaries: ```mermaid -flowchart LR +flowchart TB subgraph Q1["Q1 · sum by (job) (rate(http_requests_total[1m]))"] direction LR x1[("http_requests_total")]:::data --> x2["range 1m"]:::exact --> x3["rate"]:::exact --> x4["sum by (job)"]:::exact @@ -391,6 +414,7 @@ flowchart LR direction LR y1[("http_requests_total")]:::data --> y2["range 1m"]:::exact --> y3["sum_over_time"]:::exact --> y4["topk by (job) (10)"]:::exact end + Q1 ~~~ Q2 classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; classDef exact fill:#fff,stroke:#5f6368,color:#000; classDef summary fill:#e8f0fe,stroke:#1a73e8,stroke-width:2px,color:#000; @@ -417,11 +441,14 @@ flowchart LR end subgraph CM["Count-Min + heap per job"] direction LR - c1[("input")]:::data --> c2["Count-Min Sketch +
top-10 heap, one per job"]:::summary --> c3(["top 10
per job"]):::estimate + c1[("input")]:::data --> cg["group by job"]:::exact + cg --> ca["Count-Min + heap
job A"]:::summary --> ra(["top 10
job A"]):::estimate + cg --> cb["Count-Min + heap
job B"]:::summary --> rb(["top 10
job B"]):::estimate + cg --> cn["… one instance
per job"]:::summary --> rn(["top 10
per job"]):::estimate end subgraph H["Hydra"] direction LR - h1[("input")]:::data --> h2["Hydra over
(job, series)"]:::summary --> h3(["top 10
for each job"]):::estimate + h1[("input")]:::data --> h2["one Hydra over all
(job, series) keys"]:::summary --> h3(["top 10
for each job"]):::estimate end end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; From 9db3656c1593582b7f6562774101471c92e09ea9 Mon Sep 17 00:00:00 2001 From: Zeying Zhu <50204836+zzylol@users.noreply.github.com> Date: Thu, 1 Oct 2026 20:20:48 -0400 Subject: [PATCH 55/56] Update planner-layering.md --- .../design_docs/proposals/planner-layering.md | 46 ++++++++++++------- 1 file changed, 30 insertions(+), 16 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index a20dbb86..cf4d1101 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -205,14 +205,25 @@ window summary that answers it: slide forward with the window. * A **window summary** keeps summaries so that many windows can be answered. Three window summaries are considered for now: - * **Sliding window:** one summary per active window. Each arriving sample is - inserted into every active window that contains it, and each evaluation - reads the window that has just completed, with no merge. For a 5-min - window evaluated every 1 min, 5 windows are active and each sample updates - all 5. It works for any summary, including ones that cannot be merged, at - the cost of more ingestion work and memory. + * **Sliding window:** summaries over windows of a fixed length L that start + every s (the slide), so several windows are active at once. Each arriving + sample is inserted into every active window that contains it, at the cost + of more ingestion work and memory. A query window of length W is answered + from completed windows: + * **L = W:** each evaluation reads one completed window, with no merge. + For a 5-min window evaluated every 1 min, L = 5 min and s = 1 min, so 5 + windows are active and each sample updates all 5. This works even for + summaries that cannot be merged. + * **L shorter than W:** the query window is covered by W / L + non-overlapping completed windows, which are merged[^sliding-merge]. For + example, a 10-min window evaluated every 1 min merges two 5-min windows + with a 1-min slide. This needs a mergeable summary. + + L must divide W, and s must divide both L and the evaluation interval, so + the windows a query needs have always just completed. * **Tumbling window:** back-to-back, non-overlapping windows of one fixed - length, each with one summary. A longer query window is answered by + length, each with one summary; a sliding window whose slide equals its + length. A longer query window is answered by merging the tumbling windows it covers. The tumbling length must divide both the query window length and the evaluation interval, so that every query window starts and ends on a tumbling boundary: a 5-min window @@ -230,13 +241,8 @@ window summary that answers it: merged into a longer one as they age. TODO: evaluate other sliding-window frameworks for sketches as further - window summaries, for example: - * [Smooth Histograms for Sliding Windows](https://web.cs.ucla.edu/~rafail/PUBLIC/82.pdf) - (Braverman and Ostrovsky, FOCS 2007), an alternative to EH. - * [MicroscopeSketch: Accurate Sliding Estimation Using Adaptive Zooming](https://yangtonghome.github.io/uploads/MicroscopeSketch_SIGKDD_23_final_paper.pdf) - (Wu et al., KDD 2023). - * [Sliding Sketches: A Framework using Time Zones for Data Stream Processing in Sliding Windows](https://dl.acm.org/doi/10.1145/3394486.3403144) - (Gou et al., KDD 2020). + window summaries, such as Smooth Histograms[^smooth-histograms], + MicroscopeSketch[^microscope-sketch] and Sliding Sketches[^sliding-sketches]. | ASAP-aware CSE rule | Sharing condition | Shared computation | |---|---|---| @@ -474,8 +480,11 @@ to each of the 3 candidates, and keeps every original: summary supports both queries. * **Window-composition rule.** Each query reads a 1-min window every 10 s, so consecutive evaluations overlap by 50 s. For each query, Pass 2 adds two - variants: a **sliding window**, where each sample updates the 6 active 1-min - windows, and **10-s tumbling windows**, merged 6 at a time at every refresh. + variants: a **sliding window** with L = 1 min and a 10-s slide, where each + sample updates the 6 active windows, and **10-s tumbling windows**, merged 6 + at a time at every refresh. (Shorter sliding windows that are merged, such + as 30-s windows with a 10-s slide, would add more candidates; this example + leaves them out.) Tumbling windows need a mergeable summary. Q1's rates and sums, Q2's exact sums and Hydra merge exactly; Count-Min sketches do too, but their top-10 heaps merge only approximately, so that candidate is kept and the accuracy @@ -972,3 +981,8 @@ Histogram, or 1-min tumbling KLL windows) yields different physical plans depend only on recurrence, predictability and data arrival. This is why window-summary replacement happens in logical planning, while materialization is decided separately in physical planning. + +[^smooth-histograms]: V. Braverman and R. Ostrovsky. [Smooth Histograms for Sliding Windows](https://web.cs.ucla.edu/~rafail/PUBLIC/82.pdf). FOCS 2007. An alternative to EH. +[^microscope-sketch]: Y. Wu et al. [MicroscopeSketch: Accurate Sliding Estimation Using Adaptive Zooming](https://yangtonghome.github.io/uploads/MicroscopeSketch_SIGKDD_23_final_paper.pdf). KDD 2023. +[^sliding-sketches]: X. Gou et al. [Sliding Sketches: A Framework using Time Zones for Data Stream Processing in Sliding Windows](https://dl.acm.org/doi/10.1145/3394486.3403144). KDD 2020. +[^sliding-merge]: [ACM DOI 10.1145/1055558.1055598](https://dl.acm.org/doi/10.1145/1055558.1055598). TODO: add authors, title and venue. From dd4635fbb1aaf7df10c15160a48c0fca84ae1c7e Mon Sep 17 00:00:00 2001 From: zzylol <50204836+zzylol@users.noreply.github.com> Date: Fri, 2 Oct 2026 01:04:24 +0000 Subject: [PATCH 56/56] docs: address remaining planner-layering review comments Replace "executable", "lower", "price" and "costed" wording, note that materialized output may live on disk or in memory, keep Hydra's per-job heap explicit, mark SQL entropy/L2 recognition as frontend TODO, and fill in the sliding-window merge citation. Co-Authored-By: Claude Opus 5.5 --- .../design_docs/proposals/planner-layering.md | 28 +++++++++++-------- 1 file changed, 16 insertions(+), 12 deletions(-) diff --git a/docs/design_docs/proposals/planner-layering.md b/docs/design_docs/proposals/planner-layering.md index cf4d1101..b4b02160 100644 --- a/docs/design_docs/proposals/planner-layering.md +++ b/docs/design_docs/proposals/planner-layering.md @@ -62,7 +62,7 @@ In the diagram, × means the Cartesian product: each stage combines every option │ Physical planning — how to compute │ │ │ │ 2. Physical ASAP-aware optimization │ -│ Explore executable implementations of each logical candidate: │ +│ Explore physical implementations of each logical candidate: │ │ │ │ materialization decisions │ │ × physical operator implementations │ @@ -123,7 +123,7 @@ selection picks one. Example 1 traces this growth step by step. | Stage | Input | Decides | Output | |---|---|---|---| -| 0. Frontends | `query`, `language` | Parse and lower to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | +| 0. Frontends | `query`, `language` | Parse and convert to a common logical form; reject what cannot be represented | `CandidateLogicalDAGs` | | 1. Logical ASAP-aware optimization | Logical DAGs; accuracy requirements, `time_selection`, repetition interval | Summary replacement (Pass 1); ASAP-aware CSE (Pass 2) | `CandidateLogicalASAPDAGs` | | 2. Physical ASAP-aware optimization | Logical ASAP DAGs; `recurrence`, `predictability`, `DataWorkload` | Materialization; physical operators; parallelism and resources (TODO) | `CandidatePhysicalASAPDAGs` | | 3. Plan selection | Physical candidates; `requirements`; cost model, accuracy model, capabilities | Reject invalid candidates; pick the cheapest plan for the whole workload | One `PhysicalASAPDAG` | @@ -284,14 +284,14 @@ independent candidates, so selection can compare both. ### 2. Physical ASAP-aware optimization -Physical optimization turns each logical candidate into executable candidates. +Physical optimization turns each logical candidate into physical candidates. It makes two ASAP-specific decisions, described below. Parallelism, partitioning and resource management are TODO. #### Materialization -Materialization decides, for each sub-DAG, whether its output is kept across -(batch) query executions. Stage 2 makes this decision in three steps: +Materialization decides, for each sub-DAG, whether its output is stored, on disk +or in memory, across (batch) query executions. Stage 2 makes this decision in three steps: 1. **Materialize or not.** A sub-DAG that is not materialized always runs at query time and keeps nothing. @@ -299,7 +299,7 @@ Materialization decides, for each sub-DAG, whether its output is kept across is computed. A sub-DAG that runs at ingestion time is therefore always materialized, because its output must be kept until a query reads it. 3. **Where and how long to store it:** on disk or in memory, and for how long. - The deployment's cost model prices each choice. + The deployment's cost model estimates the cost of each choice. This gives each sub-DAG three options: @@ -334,7 +334,7 @@ A shared summary is materialized once for all its consumers. See Example 4. #### Physical operator implementation -Physical operator implementation lowers every node to physical operators, for +Physical operator implementation converts every node to physical operators, for example TopK as a sort followed by a limit, or a KLL node as summary build, merge and quantile estimation operators. @@ -348,7 +348,8 @@ earlier stages only prune provably invalid ones. Accuracy is estimated by the deployment's accuracy model, not assumed from a summary's nominal bound. Cost is evaluated for the whole workload rather than per query, which is what lets one shared summary beat several cheaper independent ones: -a shared summary is costed once, with the demand of all its consumers. +the cost of a shared summary is estimated once, with the demand of all its +consumers. ### 4. Execution @@ -407,7 +408,7 @@ first needs an exact total; the second tolerates error. This example follows the workload's candidate set through every stage. Each candidate covers both queries. -**Stage 0: 1 candidate.** The frontend lowers both queries into one workload +**Stage 0: 1 candidate.** The frontend converts both queries into one workload `LogicalDAG` with no summaries: ```mermaid @@ -435,7 +436,7 @@ flowchart TB within each `job`, which is costly at one million Zipf-distributed series. The `EpsilonDelta` target also admits a **Count-Min Sketch with a top-*k* heap per `job`**, and **Hydra over the whole `job` column**, where one sketch - covers every (`job`, series) key. + covers every (`job`, series) key and a top-*k* heap is still kept per `job`. ```mermaid flowchart LR @@ -454,7 +455,7 @@ flowchart LR end subgraph H["Hydra"] direction LR - h1[("input")]:::data --> h2["one Hydra over all
(job, series) keys"]:::summary --> h3(["top 10
for each job"]):::estimate + h1[("input")]:::data --> h2["one Hydra over all
(job, series) keys
+ heap per job"]:::summary --> h3(["top 10
for each job"]):::estimate end end classDef data fill:#f1f3f4,stroke:#5f6368,color:#000; @@ -651,6 +652,9 @@ FROM ( ); ``` +The SQL frontend does not yet recognize the Q2 and Q3 forms as `Entropy` and +`L2`. TODO: add this recognition to the per-language frontends. + | Query | Computation | Repeats | `lookback` | `as_of` | Accuracy requirement | |---|---|---|---|---|---| | Q1 | `Distinct(src_ip)` | every 10 s | 1 m | evaluation time | ε = 0.02, δ = 0.01 | @@ -985,4 +989,4 @@ separately in physical planning. [^smooth-histograms]: V. Braverman and R. Ostrovsky. [Smooth Histograms for Sliding Windows](https://web.cs.ucla.edu/~rafail/PUBLIC/82.pdf). FOCS 2007. An alternative to EH. [^microscope-sketch]: Y. Wu et al. [MicroscopeSketch: Accurate Sliding Estimation Using Adaptive Zooming](https://yangtonghome.github.io/uploads/MicroscopeSketch_SIGKDD_23_final_paper.pdf). KDD 2023. [^sliding-sketches]: X. Gou et al. [Sliding Sketches: A Framework using Time Zones for Data Stream Processing in Sliding Windows](https://dl.acm.org/doi/10.1145/3394486.3403144). KDD 2020. -[^sliding-merge]: [ACM DOI 10.1145/1055558.1055598](https://dl.acm.org/doi/10.1145/1055558.1055598). TODO: add authors, title and venue. +[^sliding-merge]: A. Arasu and G. S. Manku. [Approximate Counts and Quantiles over Sliding Windows](https://dl.acm.org/doi/10.1145/1055558.1055598). PODS 2004.