Conversation
Start `asap-physical-operators` with thin summary kernels over `asap_sketchlib`, the kernel capability checks and the typed value model. The boundary is: - `asap_sketchlib` owns sketch algorithms and their state encodings. - Kernels hold one population's in-memory state. They expose `merge`, a typed sketch readout (`estimate(&SketchQuery)`) and memory accounting. Exact states answer a typed `ExactReadout`; empty MIN/MAX read as `None`. - Group-by belongs to physical operators. - Deployments own wire decoding, delta frames, edge sampling and storage statistics. So wire decoding, `SerializableToSink`, `AggregationType`, `aux_stats`, `reset_to_empty` and the keyed/sum/min/max kernels are not carried over from ASAPQuery-backend. The `asap_sketch_codec` crate is not carried over either; it moves to `asap_sketchlib`. Hydra KLL remains as the Hydra shared-grouping kernel. HLL uses sketchlib's classic estimator. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Rebased onto main: the Cargo.lock conflict with #478 (new asap-planner crate) was resolved by regenerating workspace entries with `cargo update -w --offline`.
Add the execution layer of `asap-physical-operators`: - `plan`, `runtime` and `sources`: physical DAGs, run context, bounded backpressure, cancellation, memory reservations and raw-source scans. - `expressions` and `operators`: relational/scalar operators, windows, temporal panes, current-series snapshots, and summary build, merge and readout. A readout is typed: a `SketchQuery`, or an `ExactReadout` whose counter lookback resolves to the run's evaluation range. Deserialized operators are validated before use. - `readout`: readouts over merged exact summary states. There is no physical planner yet; operators are built directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add `physical_planner`: reader-independent compilation of selected logical candidates into precompute/query physical DAGs with typed materialization frontiers, bounded frontier enumeration, workload cost selection, temporal KLL pane compilation and PromQL row/value lowering. `dag` remains a compatibility re-export. Tests cover physical DAGs, plan recovery, precompute candidates and populations, PromQL values and binaries, weighted TopK binding and current-series heaps, plus Planner-to-execution integration tests. The design doc describes the ownership boundary with deployments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pane geometry, pane population checks and per-pane scheduling are deployment concerns. The physical layer keeps the computation the deployment binds: per-input summary build, union, shared merge and readout. - Remove `physical_planner::compile_temporal_pane_candidate` and its `TemporalPaneMaintenance`/`TemporalPaneCandidate`/`TemporalEntityIdentity` contract. - Remove the `PaneInput` operator; `ScopeTimestamp` remains as a general operator in `operators/scope_timestamp`. - Drop the `selected_temporal_lifecycle_compiles_panes_and_executes` E2E test and update the crate README and design doc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Planner search needs the physical row representation to propose whole-root physical alternatives, and asap-aware-mapping cannot depend on the physical runtime crate. promql_rows::with_series_identity keeps its signature and delegates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The backend built one workload forest per per-query physical alternative by calling SketchAlgorithmStrategy's proposal methods outside PlanSpace. Planner now owns them: search_workload_with_targets asks each strategy's new ReplacementStrategy::propose_for_root once per targeted root. SketchAlgorithmStrategy resolves the series identity and proposes current-series TopK, fixed-window and query-time Rate aggregation, and realizations over per-series Rate state, finalized and deduplicated. The candidates carry ReplacementProvenance::RootPhysicalRealization: DAG assembly uses them verbatim, and global selection never commits them. Existing candidates and their order are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PlanSpace decides what to compute; the summary maintenance lifecycle decides node timing and the physical compiler reads it. The fixed-window and query-time Rate aggregation candidates only rewrote a finalization node's timing, so they are placement variants and no longer go through propose_for_root. The backend still calls those methods directly. SketchAlgorithmStrategy::propose_for_root now proposes only current-series TopK heaps, which rank identity-carrying rows the logical root lacks. The per-series Rate filter and the verbatim-assembly branch are removed: the latter is unnecessary because non-Aggregate roots are already assembled verbatim. RootPhysicalRealization stays so global selection never commits an unvalidated identity-carrying readout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nding Split plan_summary_maintenance_lifecycles into enumeration and selection so a deployment can price every lifecycle alternative per unique summary state and bind its own choice. Planner selection is unchanged: it now enumerates and then selects the cheapest complete combination through the same path. SummaryMaintenanceLifecycleCandidates::select validates that each choice is an alternative Planner could select, enforces schedule compatibility, and obtains window frameworks and totals from the same complete-candidate hook. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A chosen summary-maintenance lifecycle did not reach the DAG: timing was fixed by realization strategies. SummaryMaintenanceLifecyclePlan:: execution_timed_dag assigns every node's timing from the selection: retained (non-Ephemeral) states and their inputs at ingestion time, everything else at query time. States without a selected lifecycle, and maintained populations that lifecycle enumeration does not cover, are refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the open question about placement ownership with the agreed four-layer contract: logical Post-ASAP decides what to compute, the chosen lifecycle decides timing, physical compilation partitions by timing, and the backend prices lifecycle choices including store cost. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`compile_candidate` re-lowered the whole Post-ASAP DAG for every materialization frontier, and `enumerate_frontiers` compiled it once more. Number helper operators from their Planner node (`u64::MAX - node_id`, at most one helper per node) so every boundary choice is a subgraph of one lowering. `cut_candidate` partitions a `compile` result for one frontier and `enumerate_compiled_frontiers` enumerates over it; both produce candidates byte-identical to per-frontier recompilation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The #462 split no longer exposes PhysicalCandidate::encode; its serde form gives the same byte-for-byte comparison. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lifecycle layer assigns each node's timing; physical compilation now reads it. `frontier_from_timing` returns the ingestion-time nodes read by query-time nodes (or an ingestion-time root) and rejects a query-time node feeding an ingestion-time one, so each lifecycle assignment is a `cut_candidate` of one `compile` result. `enumerate_compiled_frontiers` is private: placement comes from timing, and its only caller is `enumerate_frontiers`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
For the KLL quantile and grouped Rate->Sum fixtures, ContinuouslyMaintained and Ephemeral timed DAGs cut one compilation into exactly the candidates `compile_candidate` builds. The hand-written timing frontier in the chosen lifecycle test now uses `frontier_from_timing`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Number helpers as u64::MAX - (node << 16) - index so a node that lowers to an operator chain keeps deterministic, traversal-independent helper IDs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With several retained states, only those read by a query-time node or forming the root are cut points. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DAG assembly replaced any selected outer Sum over an inner aggregate with a query-time exact Sum, even when the selected summary realizes the inner Rate itself. The grouped Sum state therefore never reached the inventory, and no lifecycle choice could move grouped Sum into precompute. Assembly now keeps such a selected summary; the query-time residual still applies when the outer summary would hide its inner aggregate in KeepPreAsap. Default selection for sum by(job)(rate(...)) now yields Rate -> grouped Sum state; the physical frontier test reads its query root accordingly. Conflicts with earlier stack changes resolved to the integration tree: - crates/integration-tests/tests/summary_maintenance_lifecycle_e2e.rs: c98281a Merge remote-tracking branch 'origin/feat/compile-once-cuts' into integration/planner-for-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SketchAlgorithmStrategy::fixed_window_rate_candidates and query_time_rate_aggregation_candidates returned the same logical DAG as the ordinary heap or grouped Sum candidate with Rate finalization flipped between ingestion and query time. Placement now comes only from a chosen lifecycle via SummaryMaintenanceLifecyclePlan::execution_timed_dag. compile_fixed_window_rate_aggregation takes that lifecycle-timed PostAsapDag instead of a SummaryNode with baked-in timing. The fixed-window heap test binds continuously maintained lifecycles; the grouped Sum placement pair is covered by the lifecycle end-to-end test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
execution_timed_dag refused any plan with a MaintainPopulation node because lifecycle enumeration covered SummaryAgg states only, so population timing stayed fixed by the realization strategy. Enumeration now also emits one deployment per unique maintained population that does not feed a SummaryAgg, with the usual alternatives costed through the caller's lifecycle hooks (unknown stays unknown). select and execution_timed_dag treat it like summary state: retained at ingestion with its raw input, Ephemeral rebuilt from raw input at query time. A population feeding a SummaryAgg is that state's input and follows its timing. A plan whose population deployment was removed is still refused. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The execution-data-state validator required MaintainPopulation at ingestion time, and MaintainedPopulationStrategy hard-coded it there. Timing is now the lifecycle's choice: the validator accepts either timing and keeps the structural contracts (the population reads its matching raw input; its readout is query-time over a population that supports it). The strategy writes query time as the initial layout, as other realization strategies do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rebase note: keeps the #479 compile-once documentation and test row beside the population text, as the integration branch does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A without grouping's keys are the excluded labels, so the state column index was the excluded-key count, which could fall past the kept labels. The family was then never written into the schema and exporting the candidate panicked with SummaryFamilySchemaMismatch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…outs Query-time Binary nodes over rows with a series identity, such as avg_over_time as stored sum/count or rate(a)/rate(b), now compile to series_labels and series_binary: one-to-one matching on labels without the metric name, honoring on/ignoring when the IR carries them. A literal operand drops the metric name and applies to each value. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Grouped sum/avg (including current-series readouts) and sum/avg_over_time now use Kahan-Neumaier summation. An average switches to Prometheus' incremental mean once the running sum would overflow, so it no longer becomes +Inf where Prometheus returns a finite mean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y docs Addresses review: pin Inf/-Inf sum and average cases, show that literal arithmetic drops __name__ from a stored readout, correct the zero-start comment, and note one-to-one-only coverage and the SQL SUM/AVG change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PromQL arithmetic with a literal drops `__name__`; if two series then
share a label set, Prometheus errors ("vector cannot contain metrics with
the same labelset"). The per-series literal path returned both rows, and
the Fallback literal path kept `__name__` altogether.
Add `Operator::series_without_name`, a `SeriesLabels` rewrite that errors
on equal resulting label sets, and use it for both literal paths. Matching
and `without` relabeling still tolerate repeats: `series_binary` already
rejects right-hand and matched left-hand duplicates, and aggregation
merges them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The frontend dropped the bool modifier, so `a > bool 1` lowered like the filter `a > 1`. A separate variant keeps bool off non-comparisons and reaches both the query-level BinaryOp and the post-ASAP Binary payload. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…modifiers series_binary now implements Prometheus' VectorBinop, VectorAnd/Or/Unless and vector-scalar semantics for Fallback subtrees and query-time Binary nodes, including scalar() operands. Range functions drop __name__ in the Fallback and reject equal label sets. Conflicts with earlier stack changes resolved to the integration tree: - crates/asap-physical-operators/src/physical_planner/promql_rows.rs: a9b8fdd Merge #492 comparison and set coverage with shared series identity typing Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The comparisons, set operators, and group modifiers compiled here share series identity through asap-types (moved there by #477), while promql_rows.rs no longer keeps its own copy. Accept every BinaryOp kind and group modifier in `with_promql_series_identity`, excluding only operands that read the evaluation timestamp. Taken from integration merge a9b8fdd. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
VectorMatch has no fill field, so lowering silently dropped fill, fill_left, and fill_right, changing query results. Reject them with a fill-specific UnsupportedFeature error so the query falls back to exact execution, and lower the testdata corpus floor to the measured 1485. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A without aggregate over a nested aggregate kept the renamed value column (e.g. `sum`) as a label. Identify the value like SampleValue resolution does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HistogramQuantile now names its bucket-bound column, and the PromQL and SQL frontends group it without (le) instead of by (). The output keeps every input label except le; the compiler can now recover each histogram. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Divide before scaling when interpolating, and use almost.Equal for the small-delta fix so an infinite count is not merged into a finite one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A without (le) HistogramQuantile aggregate compiles to a series operator that groups rows by every label except le, applies bucketQuantile, and drops le and __name__, rejecting equal result label sets. Conflicts with earlier stack changes resolved to the integration tree: - crates/asap-physical-operators/src/operators/mod.rs: 2bea2f6 Merge #493 classic histogram quantile coverage - series_labels.rs and tests/promql_fallback.rs: placement from 2bea2f6 without the later c39156e/15ab72b changes of this PR Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts with earlier stack changes resolved to the integration tree: - docs/develop_docs/physical-compile-coverage.md: 2bea2f6 Merge #493 classic histogram quantile coverage, without the paragraph that 6d895b7 in this PR adds later. The comparison section from #492 stays, the Remaining list drops the entries both PRs now cover, and the totals count both PRs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SQL marker is the projection's only column, so grouping without (le) would add undeclared output columns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This PR's histogram_quantile branch reads a function-level `left_layout`, but #492 restructured `series_labels::execute` to bind it inside each binary branch. Bind it from the operator input in the histogram branch, and format the widened `Kind` match arm. Taken from integration commit 0d132b9. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A deployment reading a stored Count-Min state's total had to call a
kernel accessor itself, so the computation lived outside the compiled
plan. The Planner already expresses that readout as
SketchQuery::PointCount { value: None }; it failed to compile because
plain Count-Min was not a native state.
Plain Count-Min is now a native stored state for merge and bare-count
readout (not for building from rows, which still fails closed). The
kernel answers the bare count through AggregateCore::estimate with one
row's mass, so colliding items keep their weight. The readout is Int64,
matching the Planner's unit-weight count output, and rejects a
non-integral total.
Count-Min with heap stays without a compiled bare count: its native
state is WeightedFrequency, which carries no total.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
force-pushed
the
feat/histogram-quantile
branch
from
September 30, 2026 20:24
fbeddda to
9b53ea7
Compare
zzylol
force-pushed
the
feat/cms-scaled-total
branch
from
September 30, 2026 20:24
cc766fe to
112148d
Compare
zzylol
force-pushed
the
feat/histogram-quantile
branch
2 times, most recently
from
October 1, 2026 23:14
50c9d6d to
ae5f517
Compare
zzylol
added a commit
that referenced
this pull request
Oct 1, 2026
#507) Count-Min totals, typed series identity, time() operands, non-finite values, compensated exact Sum, label_replace, @ start()/end(), stored-series filters and set operators, deriv/predict_linear, and typed maintenance arithmetic. The compensated Sum keeps its plain Sum name: there is no persisted format to version yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #493.
Supersedes #496, which carried the same change against
integration/planner-for-backend; GitHub marked #496 merged when its head was incorporated into that branch, but it never reached main. Review this PR instead; #496 stays closed.Why
Backend bare-count readouts need a sampling-aware Count-Min total without reproducing kernel arithmetic.
What
Expose
CountMinSketchAccumulator::total()and the unsampled heap variant, using one row's mass scaled by the edge sampling probability. CountSketch has signed rows and intentionally has no total accessor.Before this PR
A sketch containing ten sampled updates at p=0.25 required backend arithmetic to return the estimated total 40.
After this PR
The backend asks the Planner kernel for that total directly; collisions and merges preserve the mass.
Validation
Focused total tests pass (including evicted heap entries) for colliding keys, p=0.25, merged states, and an empty state.
🤖 Generated with Claude Code