Conversation
Add the layering proposal, its proposals-index entry, and an Output layers section in the input/output/workflow design. Every Planner layer outputs all legal candidates; the deployment keeps every summary-family candidate and selects with its own costs. Design DAG names are marked as the target API with the current main type alongside. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Split physical planning into physical design (summary materialization, workload-level, like materialized-view selection) and physical implementation (per-node lowering and cutting). The annotated graph stays a Post-ASAP DAG; materialization is decided in design and realized by the cut. State node-by-node correspondence and the Fallback exception that the operator-flattening proposal removes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Describe current Planner behaviour (no family pruning) and the backend gap, and replace the Binary 'exception' with the general rule that a timing- sensitive node needs one compilation per distinct timing, with an example. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ng proposal Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Planner takes the query and data workloads plus the deployment's cost model, accuracy requirements and capabilities, and returns one optimal PhysicalDAG; candidate sets stay internal. Name the annotated stage MaterializedPostASAPDAG, tabulate what each DAG encodes, fold timing and selection into the layer descriptions, rename physical design to summary materialization, and make the example trace one query through every DAG. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…g row A lifecycle fixes materialization, timing, maintenance, retention and window framework together, so the stage and its DAG are named after the lifecycle (LifecyclePostASAPDAG) rather than one of those aspects. The deployment row no longer mentions ranking or selection. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Describe the two plain candidate collections, the view-based compile, the caller-typed physical candidate errors, lifecycle_guarantee, and PhysicalExecution as an execution handle. Remove APIs the docs said were removed but never existed (compile_timed_candidates, PhysicalDAGCandidate), the agent instructions in the alignment proposal's baseline, and the ingestion-time Binary exception from design docs, where it is an implementation detail (the developer migration guide keeps it). Restore #485's statement that candidates do not choose placement and #508's CandidatePostASAPDAGs<Id> names in input-output-workflow.md, rejoin the split test table in physical-planning-and-deployment.md, and take planner-backend-layering.md verbatim from #509. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The layering proposal (#509) names a single annotated DAG, LifecyclePostASAPDAG: a PostASAPDAG plus its lifecycle assignment. The code kept a SummaryMaintenanceLifecyclePlan beside the DAG and named the timed collection CandidatePostASAPDAGsWithTiming. - Rename SummaryMaintenanceLifecyclePlan to LifecyclePostASAPDAG and its error to LifecyclePostASAPDAGError. The type already held the root and each state's lifecycle, retention and window framework; per-node timing stays derived by execution_assignment as the existing PostASAPDAGAssignment overlay, so no timed graph is stored beside it and no new type is added. - Rename CandidatePostASAPDAGsWithTiming to CandidateLifecyclePostASAPDAGs; CandidatePostASAPDAGs stays the logical collection. - Call layer 2 "summary lifecycle planning" and reword comments that had the deployment choose or rank; selection is Planner's, over the deployment's cost model. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Several docs still said the deployment or "downstream" selects, ranks or chooses the plan, and the glossary said physical plans are owned by downstream systems. Per the layering proposal (#509), Planner compiles and selects the physical plan using the deployment's cost model; the deployment supplies prices and executes the selected plan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| Each stage adds decisions to the DAG it receives. The table shows which | ||
| decisions each DAG carries. | ||
|
|
||
| | | `PreASAPDAG` | `PostASAPDAG` | `LifecyclePostASAPDAG` | `PhysicalDAG` | |
There was a problem hiding this comment.
Other parts lgtm! The current naming is still a little bit weird.
PreASAPDAGandPostASAPDAGlooks like a pair of input and output, but they actually are not- The preamble
PostASAPappears in part of the DAG names in the layering but disappear in the end, so what doesPostASAPmean then?
That said, I don't want to make the naming discussion to go back and forth forever. If you still feel like to stick to your opinion, I will not push back more.
But I suggest you discuss the naming and the semantic of each layer of DAG with Milind before finalizing.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every stage after the frontend is a Post-ASAP DAG; the prefix says how far planning has gone: Logical, Lifecycle, Physical. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PostASAPDAG -> LogicalPostASAPDAG, CandidatePostASAPDAGs -> CandidateLogicalPostASAPDAGs, PhysicalDAG -> PhysicalPostASAPDAG, CandidatePhysicalDAGs -> CandidatePhysicalPostASAPDAGs, plus the PostASAPDAG* compounds and InvalidPostASAPDAG. Serialized forms are unchanged. Re-copies planner-layering.md from #509 (41076ca). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| |---|---|---| | ||
| | `PreASAPDAG` | `Rc<QueryExpr>` | `PreASAPDAG` | | ||
| | `LogicalPostASAPDAG` | `Rc<SummaryNode>` tree; exported as `PostAsapDag` | `LogicalPostASAPDAG` | | ||
| | `LifecyclePostASAPDAG` | `SummaryMaintenanceLifecyclePlan`, one selected assignment beside the DAG | `LifecyclePostASAPDAG`; `SummaryMaintenanceLifecyclePlan` is merged into it | |
There was a problem hiding this comment.
In another PR, the lifecycle layer should be reconsidered the design by the following points.
-
Decisions: Whether the output of a sub-DAG is materialized or not.
- whether a sub-DAG is computed at ingestion time or query time, or collection time (once subDAG is computed at ingestion/collection time, subDAG is materialized)
- How long is the materialized state maintained and stored.
-
Question: How can / Should we model streaming vs batch data as Data Workload and how does the data workload affect the lifecycle decisions? As the background, DataFusion has "bounded" vs "unbounded" data input for DF nodes, what is the relationship between this question to DF design?
-
The "lifecycle" in this layer / ASAPPlanner refers to the query workload patterns related to time, e.g., queries repeating over time, historical queries with longer latency tolerance. But the term "lifecycle" is reused against Data Lifecycle in ASAP (referring to the stages of data collection, transmission, storage, and analytics). What should be a better term and the scope for this "lifecycle" layer in ASAPPlanner?
-
Also, we should find the minimal decision space for "lifecycle" layer by analyzing concrete use case deployments and their query workloads and data workloads.
| ```text | ||
| Query workload (PromQL / SQL / MetricsQL, query repeating pattern, etc.) | ||
| + data workload | ||
| + deployment inputs: cost model, accuracy requirements, capabilities |
There was a problem hiding this comment.
- Accuracy requirement should be part of query workload
- I believe there is a separate model for estimating Accuracy. This can be part of "deployment input"?
| │ -> CandidatePhysicalDAGs │ | ||
| │ │ │ | ||
| │ 4. Selection │ | ||
| │ Cost every candidate with the deployment's cost model; │ |
There was a problem hiding this comment.
This should also include pruning/filtering based on accuracy requirements?
| | | `PreASAPDAG` | `PostASAPDAG` | `LifecyclePostASAPDAG` | `PhysicalDAG` | | ||
| |---|---|---|---|---| | ||
| | Produced by | 0. Frontends | 1. Logical optimization | 2. Summary lifecycle planning | 3. Compilation; 4. selects one | | ||
| | Node | Query operation | Logical operation, including summary operations | Same, plus annotations | Physical operator | |
There was a problem hiding this comment.
I think "Query operation" should be replaced by "Logical operations without any summary operations"?
There was a problem hiding this comment.
Will be useful to link the enum here
| | Node | Query operation | Logical operation, including summary operations | Same, plus annotations | Physical operator | | ||
| | Logical optimization (summary family, rewrites) | No | Yes | Yes | Yes | | ||
| | Materialization decided (which summary states persist) | No | No | Yes | Yes | | ||
| | Data lifecycle (how each state is maintained: `Ephemeral`, `Prepared`, `Shared`, `ContinuouslyMaintained`) | No | No | Yes | Yes | |
There was a problem hiding this comment.
We are using both terms "data lifecycle" and "summary lifecycle". Are they same or differnet?
| | Design name | Current main | Target API (open PRs #508, #480) | | ||
| |---|---|---| | ||
| | `PreASAPDAG` | `Rc<QueryExpr>` | `PreASAPDAG` | | ||
| | `PostASAPDAG` | `Rc<SummaryNode>` tree; exported as `PostAsapDag` | `PostASAPDAG` | |
There was a problem hiding this comment.
current is tree or DAG?
| | 2. Summary lifecycle planning | For each unique summary state and maintained population, the admissible lifecycle assignments and their timing, window framework and retention. | Cost values; operator implementation | | ||
| | 3. Compilation | All computation: value operations, aggregation, PromQL functions and subqueries, vector matching, comparisons and set operators, `histogram_quantile`, summary build, merge and estimate, sort, limit, joins. | Raw ingestion, pane construction, storage formats, decoding persisted state, scheduling | | ||
| | 4. Selection | Costing every candidate with the deployment's cost model and returning the cheapest admissible `PhysicalDAG` for the whole workload that meets the accuracy requirements. A state shared by several queries is costed once with all consumers' demand (only when compilation installs one shared output: same window layout, evaluation interval and phase). Unknown cost stays unknown and such a candidate is not selected. | The cost values | | ||
| | 5. Deployment | Inputs: the cost model (build, per-update maintenance, read, store price per byte-second, retirement, query-time raw processing; optionally whole-plan quotes), accuracy requirements and capabilities (for example whether query-time raw data is available). Execution: ingestion and routing, panes and completeness, lateness and revisions, storage and codecs over Planner kernel states, reading stored state into typed inputs, query-time raw sources, the exact-engine fallback. Sampled or delta edge frames are rejected. | Any computation algorithm | |
There was a problem hiding this comment.
The way it's written it seems like these are the inputs to the deployment, not the inputs provided by the deployment.
There was a problem hiding this comment.
I think we can mention the planning and execution phase separately. Everything in the layers above is about planning. This layer provides information about planning but its role is actually executing the plans on data.
| The deployment passes its inputs to ASAPPlanner and receives one optimal | ||
| `PhysicalDAG`, which contains: | ||
|
|
||
| * a **precompute DAG**, whose inputs are raw-sample contracts (rows carrying |
There was a problem hiding this comment.
Confused by what "contract" means
| ## Example | ||
|
|
||
| This traces `sum by (job) (rate(m[1m]))`, evaluated every 10 s, through the | ||
| four DAGs, and shows that only the deployment's store price changes the plan. |
There was a problem hiding this comment.
"price" is also a weird english word here. We can just say cost or cost model
| (a) both retained, (b) Rate retained and Sum `Ephemeral`, (c) both | ||
| `Ephemeral`. | ||
| * `PhysicalDAG`: one compilation, cut three ways. (a) Precompute builds Rate | ||
| and Sum per pane; the query only reads Sum. (b) Precompute keeps Rate; the |
There was a problem hiding this comment.
do these a,b,c choices depend on the a,b,c choices from the previous bullet point? I think we can start with a concrete e2e example. Then given one example say, "oh if the cost of storage is higher, then plan B is more optimal than the example we showed"
| (a) both retained, (b) Rate retained and Sum `Ephemeral`, (c) both | ||
| `Ephemeral`. | ||
| * `PhysicalDAG`: one compilation, cut three ways. (a) Precompute builds Rate | ||
| and Sum per pane; the query only reads Sum. (b) Precompute keeps Rate; the |
There was a problem hiding this comment.
"pane" is undefined
| │ Physical planning (how) │ | ||
| │ 2. Summary lifecycle planning │ | ||
| │ Per summary state: Ephemeral | Prepared | Shared | │ | ||
| │ ContinuouslyMaintained -> node timing, window framework, │ |
There was a problem hiding this comment.
pls give examples of window frameworks
Why
Planner and ASAPQuery-backend work is split across many open PRs that assume a layering contract no merged doc states. This lands that contract on main now, with the owner decisions applied, so reviewers of the code stacks have one reference.
What
docs/design_docs/proposals/planner-layering.md(new proposal, designer/architect audience): layers, what each DAG encodes, name mapping to code, responsibilities, boundary, and a worked example.docs/design_docs/proposals/README.md: index entry.Owner decisions reflected:
PhysicalPostASAPDAG. Candidate sets are internal; the deployment never ranks or selects.CandidatePreASAPDAGs,CandidateLogicalPostASAPDAGs,CandidateLifecyclePostASAPDAGs,CandidatePhysicalPostASAPDAGs); selection (layer 4) is the only step that chooses, including among summary families.LifecyclePostASAPDAG;SummaryMaintenanceLifecyclePlanmerges into it.LogicalPostASAPDAG,LifecyclePostASAPDAGandPhysicalPostASAPDAGname the stages; the new table states, per DAG, whether logical optimization, materialization, lifecycle, retention, execution time and physical optimization are encoded.Before this PR
Main describes ASAPPlanner as a logical-only library whose output is
PlanSpace; nothing documents physical compilation, lifecycle timing, or where selection belongs.After this PR
Not in scope
CandidatePostASAPDAGsWithTiming→CandidateLifecyclePostASAPDAGs.🤖 Generated with Claude Code