Conversation
This was referenced Sep 30, 2026
Open
Start `asap-physical-operators` with thin summary kernels over `asap_sketchlib`, the kernel capability checks and the typed value model. The boundary is: - `asap_sketchlib` owns sketch algorithms and their state encodings. - Kernels hold one population's in-memory state. They expose `merge`, a typed sketch readout (`estimate(&SketchQuery)`) and memory accounting. Exact states answer a typed `ExactReadout`; empty MIN/MAX read as `None`. - Group-by belongs to physical operators. - Deployments own wire decoding, delta frames, edge sampling and storage statistics. So wire decoding, `SerializableToSink`, `AggregationType`, `aux_stats`, `reset_to_empty` and the keyed/sum/min/max kernels are not carried over from ASAPQuery-backend. The `asap_sketch_codec` crate is not carried over either; it moves to `asap_sketchlib`. Hydra KLL remains as the Hydra shared-grouping kernel. HLL uses sketchlib's classic estimator. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Rebased onto main: the Cargo.lock conflict with #478 (new asap-planner crate) was resolved by regenerating workspace entries with `cargo update -w --offline`.
Add the execution layer of `asap-physical-operators`: - `plan`, `runtime` and `sources`: physical DAGs, run context, bounded backpressure, cancellation, memory reservations and raw-source scans. - `expressions` and `operators`: relational/scalar operators, windows, temporal panes, current-series snapshots, and summary build, merge and readout. A readout is typed: a `SketchQuery`, or an `ExactReadout` whose counter lookback resolves to the run's evaluation range. Deserialized operators are validated before use. - `readout`: readouts over merged exact summary states. There is no physical planner yet; operators are built directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add `physical_planner`: reader-independent compilation of selected logical candidates into precompute/query physical DAGs with typed materialization frontiers, bounded frontier enumeration, workload cost selection, temporal KLL pane compilation and PromQL row/value lowering. `dag` remains a compatibility re-export. Tests cover physical DAGs, plan recovery, precompute candidates and populations, PromQL values and binaries, weighted TopK binding and current-series heaps, plus Planner-to-execution integration tests. The design doc describes the ownership boundary with deployments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pane geometry, pane population checks and per-pane scheduling are deployment concerns. The physical layer keeps the computation the deployment binds: per-input summary build, union, shared merge and readout. - Remove `physical_planner::compile_temporal_pane_candidate` and its `TemporalPaneMaintenance`/`TemporalPaneCandidate`/`TemporalEntityIdentity` contract. - Remove the `PaneInput` operator; `ScopeTimestamp` remains as a general operator in `operators/scope_timestamp`. - Drop the `selected_temporal_lifecycle_compiles_panes_and_executes` E2E test and update the crate README and design doc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Planner search needs the physical row representation to propose whole-root physical alternatives, and asap-aware-mapping cannot depend on the physical runtime crate. promql_rows::with_series_identity keeps its signature and delegates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The backend built one workload forest per per-query physical alternative by calling SketchAlgorithmStrategy's proposal methods outside PlanSpace. Planner now owns them: search_workload_with_targets asks each strategy's new ReplacementStrategy::propose_for_root once per targeted root. SketchAlgorithmStrategy resolves the series identity and proposes current-series TopK, fixed-window and query-time Rate aggregation, and realizations over per-series Rate state, finalized and deduplicated. The candidates carry ReplacementProvenance::RootPhysicalRealization: DAG assembly uses them verbatim, and global selection never commits them. Existing candidates and their order are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PlanSpace decides what to compute; the summary maintenance lifecycle decides node timing and the physical compiler reads it. The fixed-window and query-time Rate aggregation candidates only rewrote a finalization node's timing, so they are placement variants and no longer go through propose_for_root. The backend still calls those methods directly. SketchAlgorithmStrategy::propose_for_root now proposes only current-series TopK heaps, which rank identity-carrying rows the logical root lacks. The per-series Rate filter and the verbatim-assembly branch are removed: the latter is unnecessary because non-Aggregate roots are already assembled verbatim. RootPhysicalRealization stays so global selection never commits an unvalidated identity-carrying readout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nding Split plan_summary_maintenance_lifecycles into enumeration and selection so a deployment can price every lifecycle alternative per unique summary state and bind its own choice. Planner selection is unchanged: it now enumerates and then selects the cheapest complete combination through the same path. SummaryMaintenanceLifecycleCandidates::select validates that each choice is an alternative Planner could select, enforces schedule compatibility, and obtains window frameworks and totals from the same complete-candidate hook. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A chosen summary-maintenance lifecycle did not reach the DAG: timing was fixed by realization strategies. SummaryMaintenanceLifecyclePlan:: execution_timed_dag assigns every node's timing from the selection: retained (non-Ephemeral) states and their inputs at ingestion time, everything else at query time. States without a selected lifecycle, and maintained populations that lifecycle enumeration does not cover, are refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the open question about placement ownership with the agreed four-layer contract: logical Post-ASAP decides what to compute, the chosen lifecycle decides timing, physical compilation partitions by timing, and the backend prices lifecycle choices including store cost. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`compile_candidate` re-lowered the whole Post-ASAP DAG for every materialization frontier, and `enumerate_frontiers` compiled it once more. Number helper operators from their Planner node (`u64::MAX - node_id`, at most one helper per node) so every boundary choice is a subgraph of one lowering. `cut_candidate` partitions a `compile` result for one frontier and `enumerate_compiled_frontiers` enumerates over it; both produce candidates byte-identical to per-frontier recompilation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The #462 split no longer exposes PhysicalCandidate::encode; its serde form gives the same byte-for-byte comparison. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lifecycle layer assigns each node's timing; physical compilation now reads it. `frontier_from_timing` returns the ingestion-time nodes read by query-time nodes (or an ingestion-time root) and rejects a query-time node feeding an ingestion-time one, so each lifecycle assignment is a `cut_candidate` of one `compile` result. `enumerate_compiled_frontiers` is private: placement comes from timing, and its only caller is `enumerate_frontiers`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
For the KLL quantile and grouped Rate->Sum fixtures, ContinuouslyMaintained and Ephemeral timed DAGs cut one compilation into exactly the candidates `compile_candidate` builds. The hand-written timing frontier in the chosen lifecycle test now uses `frontier_from_timing`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Number helpers as u64::MAX - (node << 16) - index so a node that lowers to an operator chain keeps deterministic, traversal-independent helper IDs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With several retained states, only those read by a query-time node or forming the root are cut points. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DAG assembly replaced any selected outer Sum over an inner aggregate with a query-time exact Sum, even when the selected summary realizes the inner Rate itself. The grouped Sum state therefore never reached the inventory, and no lifecycle choice could move grouped Sum into precompute. Assembly now keeps such a selected summary; the query-time residual still applies when the outer summary would hide its inner aggregate in KeepPreAsap. Default selection for sum by(job)(rate(...)) now yields Rate -> grouped Sum state; the physical frontier test reads its query root accordingly. Conflicts with earlier stack changes resolved to the integration tree: - crates/integration-tests/tests/summary_maintenance_lifecycle_e2e.rs: c98281a Merge remote-tracking branch 'origin/feat/compile-once-cuts' into integration/planner-for-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SketchAlgorithmStrategy::fixed_window_rate_candidates and query_time_rate_aggregation_candidates returned the same logical DAG as the ordinary heap or grouped Sum candidate with Rate finalization flipped between ingestion and query time. Placement now comes only from a chosen lifecycle via SummaryMaintenanceLifecyclePlan::execution_timed_dag. compile_fixed_window_rate_aggregation takes that lifecycle-timed PostAsapDag instead of a SummaryNode with baked-in timing. The fixed-window heap test binds continuously maintained lifecycles; the grouped Sum placement pair is covered by the lifecycle end-to-end test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
execution_timed_dag refused any plan with a MaintainPopulation node because lifecycle enumeration covered SummaryAgg states only, so population timing stayed fixed by the realization strategy. Enumeration now also emits one deployment per unique maintained population that does not feed a SummaryAgg, with the usual alternatives costed through the caller's lifecycle hooks (unknown stays unknown). select and execution_timed_dag treat it like summary state: retained at ingestion with its raw input, Ephemeral rebuilt from raw input at query time. A population feeding a SummaryAgg is that state's input and follows its timing. A plan whose population deployment was removed is still refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The execution-data-state validator required MaintainPopulation at ingestion time, and MaintainedPopulationStrategy hard-coded it there. Timing is now the lifecycle's choice: the validator accepts either timing and keeps the structural contracts (the population reads its matching raw input; its readout is query-time over a population that supports it). The strategy writes query time as the initial layout, as other realization strategies do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rebase note: keeps the #479 compile-once documentation and test row beside the population text, as the integration branch does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A population read directly and also consumed by a SummaryAgg became its own deployment or not depending on traversal order, and its own lifecycle could disagree with the retained state built from it. Any population reachable from a SummaryAgg is now that state's input and never a separate deployment, so the state's lifecycle times it in either order. Also clarify review-noted wording: retained-state docs, the uniform-pricing consequence for cost models, and stale population-timing notes in the design proposals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ReadPopulation Sum/Count/Average/Quantile now compile to a grouped aggregate over the population snapshot, so deployments no longer evaluate these readouts in their current-series store. Quantile uses a new exact Reduction::Quantile with PromQL rank interpolation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A query-time Binary with a PromQL scalar-literal operand folds the literal into a projection, which also covers unary negation. Two grouped row inputs match one-to-one on equal label columns through an inner equi-join before the operator is applied. Comparisons and per-series rows still fail at compile time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Exact count readout yields Int64, but PromQL declares a Float64 sample, so compile rejected count finalization. Convert exactly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A global population aggregate with no members emitted one row; PromQL returns an empty vector. Quantile now orders NaN samples first, as Prometheus does. Found in independent review. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Evaluates instant selection and range functions over left-open windows at the query time or on a subquery step grid. Conflicts with earlier stack changes resolved to the integration tree: - crates/asap-physical-operators/src/operators/mod.rs: a7ff3ae Merge remote-tracking branch 'origin/feat/physical-compile-promql-fallback' into integration/planner-for-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts with earlier stack changes resolved to the integration tree: - crates/asap-physical-operators/src/physical_planner/promql_rows.rs: a7ff3ae Merge remote-tracking branch 'origin/feat/physical-compile-promql-fallback' into integration/planner-for-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tile over time Per-series windows evaluate these PromQL range functions with Prometheus semantics over fresh (non-stale) samples. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, without and @ Each Fallback selector reads its own raw-series slot. Vector-vector arithmetic uses PromQL one-to-one matching with on/ignoring via new series_labels and series_binary operators; without grouping rewrites the series identity; @ <timestamp> fixes selector and subquery evaluation. Conflicts with earlier stack changes resolved to the integration tree: - crates/asap-physical-operators/src/operators/mod.rs: 9a13ae4 integrate: extend asap-types series identity with #486/#487 shapes - crates/asap-physical-operators/src/physical_planner/promql_rows.rs: 9a13ae4 integrate: extend asap-types series identity with #486/#487 shapes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…antile IR needs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#477 moved PromQL series-identity resolution into asap-types. The fallback shapes compiled by #486 and #487 (time shifts, subqueries, scalar bridges, and arithmetic between series) also need identity realization there. Taken from integration commit 9a13ae4. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A precompute boundary at a raw time-series scan binds as raw sample rows ($population label map, $timestamp, value). The label map is the complete source identity, so per-series and grouped SummaryAgg lower over it with constant, column or unit-frequency (HLL) updates, and keyed heaps resolve their items from labels, the sample value or the canonical label identity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Raw sample rows drop empty label values so one series has one population. Heap items resolve names against the scan (the value column, the series identity, labels; other scan columns are rejected), and identity items may exclude labels, as `topk by` emits. Unit-frequency HLL updates are raw-only and never apply to keyed families. Tests cover Planner-generated heaps, missing and empty labels, and `without` grouping. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
State the canonical label-set obligation and assert per-query coverage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
force-pushed
the
feat/physical-compile-coverage-3
branch
from
September 30, 2026 18:27
eb34e4f to
d4375b3
Compare
zzylol
force-pushed
the
feat/precompute-raw-sample-input
branch
from
September 30, 2026 18:27
4b6dffd to
f574c00
Compare
zzylol
force-pushed
the
feat/physical-compile-coverage-3
branch
3 times, most recently
from
October 1, 2026 22:58
148353e to
b2c0aa6
Compare
zzylol
added a commit
that referenced
this pull request
Oct 1, 2026
…ssors (#488, #495) Raw-sample precompute boundary, plus format-agnostic kernel accessors (serde with validation, sketch()/from_sketch(), as_any_mut). The WeightedFrequency byte codec from #495 is intentionally left out: byte formats belong to the deployment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #487.
Why
The backend still evaluates raw-ingest summary updates itself (value/item expressions, per-series and grouped update), because Planner's precompute compiler only accepts stored summary states as boundaries. For a raw time-series scan it failed with
stored population requires one typed summary state; the general compiler failed per-series summaries withper-entity summary requires complete source identity. Planner should own this computation; a deployment should only supply rows and panes.What
precompute::compileaccepts a raw time-series scan (FallbackoverScan{TimeSeries}, optionally underTimeRange) as a boundary, bound as raw sample rows[$population: label map, $timestamp, value](precompute::raw_sample_schema, rows viaraw_sample_row). The canonical label map is the complete source identity, so per-series summaries need no identity-typed plan.SummaryAggover that boundary lowers per-series and grouped (by/without), with constant or column weights; HLL unit-frequency updates observe the sample value; CMS/CountSketch heaps resolve items from labels, the sample value, or the canonical label identity (optionally excluding labels, astopk byemits) via newExpression::Label/Expression::LabelIdentity.Before this PR
Backend probe over 16 PromQL queries (exact, Epsilon, EpsilonDelta; 75 raw stored outputs): 11 compile, 64 fail (
per-entity summary requires complete source identity); no raw boundary is accepted byprecompute::compile.After this PR
The same probe: 75/75 raw outputs compile through
precompute::compile(dag, &[raw_source], &[summary]). New acceptance testprecompute_raw_samplesruns every Planner raw-input candidate for 18 queries (Sum, Count, Min, Max, Rate, Increase, KLL, DDSketch, HLL, CountSketchWithHeap from Planner; CmsWithHeap hand-built) and matches each population's estimates against its kernel fed sample by sample. Families without a native state (plain CMS/CountSketch, Kmv, Theta, UnivMon) still fail to compile, as before.Behaviour differences
SummaryAggfragment is unchanged (items over finalized readouts remain rejected).FiniteFloat64); stale markers are not samples.raw_sample_rowguarantees it.Remaining
topk by) have no readout-side decoder in this crate yet; a reader must add group labels back from$population.sum without (instance) (sum_over_time(m[5m]))candidates hit a pre-existingcompile_post_asap_dagSummaryFamilySchemaMismatch;withoutis covered by an edited DAG.Validation
cargo fmt --all -- --check,cargo clippy --workspace --all-targets --all-features --locked -- -D warnings,cargo test --workspace --no-fail-fast(pass). The acceptance test fails on the base (stored population requires one typed summary state) and passes here. Reviewed by a separate reviewer agent; its findings are addressed in the follow-up commits.🤖 Generated with Claude Code