Conversation
zzylol
added a commit
that referenced
this pull request
Oct 4, 2026
RootDemand gained latency_ms in #604; the Example 2 coverage test's demand sets it to None (no latency bound). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
added a commit
that referenced
this pull request
Oct 4, 2026
#604 gave RootDemand a latency_ms and a logical candidate several physical candidates. The Hydra count test sets no latency bound and checks the all-query-time physical candidate, which comes first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 4, 2026
…ummary #509 Stage 2 materialization (#580 step 5, decisions S1-S6, Q44): - Eligibility (S4): a SummaryAgg may run at ingestion time when the data is continuously ingesting or mixed, every root reaching it repeats and is predictable, and it is a tumbling pane or reads one fixed window per evaluation (FixedIntervalAt, window no longer than the interval). The panes of one merge are one unit. - Stage 2 enumerates the down-closed sets of units (all query time first), each built with a workload-wide split_shared_by_phase that copies a node shared by an ingestion-time and a query-time consumer per phase and re-keys the assignment. Above 16 sets it searches greedily by Stage 3's cost and flags the selection as not guaranteed optimal. - Stage 3 takes the cheapest valid physical candidate of each logical one, in the DP and in enumeration. Ingestion-time panes are one pane built as rows arrive: the newest pane pays the build and retains N + 1 panes; the older panes and their inputs cost nothing. - Latency check (S6): a query whose query-time work per evaluation exceeds its latency_ms is rejected with the estimate and the bound (Stage3Calibration::latency_ms_per_cost_unit, 1 CPU-ms per ms). - RootDemand carries latency_ms. Example 1: 112 physical candidates (24 maintain Q2's sum panes), same selection P82 at 4.620; maintained panes cost 47.4 (42.0 memory); 40 candidates over Q2's 100 ms bound. Example 3B: P2 KLL still selected at 2.083; tumbling panes at query time miss the 200 ms bound, maintained ones cost 768.6 (768.0 memory). The executor runs Example 3B's maintained plan per pane and returns the whole-window rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
force-pushed
the
stack/509-x8a-hydra-kernel
branch
from
October 5, 2026 06:21
09fa09e to
f97037a
Compare
zzylol
force-pushed
the
stack/509-x5-stage2-materialization
branch
from
October 5, 2026 06:21
ec82a21 to
78cf83f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: Wave 2 chain: #599 → #601 → #600 → #604 → #606 → #603 → #605
Rebased on main d4869a7 (DF 54).
Stacked on #600. Part of #580, step 5; #509 Stage 2 "Materialization". Implements decisions S1–S6 and Q44 (pricing #594, panes #601).
Why
Stage 2 ran every summary at query time, so Stage 3 never compared maintaining state at ingestion time with rebuilding it at each evaluation. Example 4's B1 (tumbling KLLs at ingestion time) was not generated. Query latency bounds were not checked.
What
Eligibility (S4). A
SummaryAggmay run at ingestion time when all of these hold:DataWorkload.arrivalisContinuouslyIngestingorMixed.Unknowncounts as not ingesting.Predictable.FixedIntervalAt, with a window no longer than the interval.A summary over a merge of panes reads a sliding window, so it stays at query time. The panes of one merge form one unit.
Down-closed sets. Stage 2 returns one physical candidate per set of units in which everything below an ingestion-time unit is also at ingestion time. The all-query-time candidate comes first. Ids are
P<n>for all query time andP<n>-m<k>otherwise; labels end in "· ingestion time: Kll ×5 panes". Above 16 sets, Stage 2 adds units greedily by Stage 3's cost and flags the selection as not guaranteed optimal.Cross-root phase split.
split_shared_by_phase(&[roots], assignment)copies a node per phase when one consumer reads it at ingestion time and another at query time, across queries. It returns the assignment re-keyed to the copies.Selection. The tree DP's
evaluate, enumeration andfinishtake the cheapest valid physical candidate of each logical candidate. DP-versus-exhaustive tests run with materialization on.Pricing panes. At ingestion time, a pane chain is one pane, built as rows arrive and kept. The newest pane pays the λ-rate build and the memory of
lookback/width + 1panes (S3: in memory only). Older panes, and the shifts and ranges that feed only them, cost 0.Latency check (S6). A candidate is rejected when a query's query-time work for one evaluation is above its
latency_ms. The work counts every query-time node the query reaches, shared nodes included. NewStage3Calibration::latency_ms_per_cost_unit= 1 (one CPU-ms on one core). Example reason:q1: query-time work takes 310.0 ms per evaluation, over the 200 ms latency bound.RootDemandnow carrieslatency_ms.Docs.
stage3-cost-model.mdcovers pane retention, the latency check and the calibration field, and updates Example 1's worked example.Before / After (built-in models, constants untuned)
sum_over_timepanesIngestion time does not win any example under the built-in models. With 1M series, memory for retained per-series state is much larger than the CPU it saves. In a unit test with 10 series at 100k rows/s, maintained panes win, and the DP selects them too.
Tests
timing.rs: a cross-root split keeps the assignment, and a node reached at one timing stays shared.materialization.rs: panes go to ingestion time and the merge stays at query time. Arrival, ad hoc and one-off rules. A whole window needs a phase and must not overlap. Sets are down-closed: an exact sum plus a CMS gives 3 sets, and over panes only the panes count. Greedy above the cap.plan-selection: the newest pane pays the build plus (N+1)× memory, and the older panes pay 0. The latency bound rejects slow candidates with a reason. Maintained panes win with few series, and the DP agrees with exhaustive selection.runtime_b_maintained_panes_match_the_whole_windowcompiles the B1 candidate once. It cuts the candidate atfrontier_from_timingwithcut_candidate, runs the precompute once per pane in that pane'sScope::Ingestionwindow, then runs the query over the 5 kept panes. It returns the whole-window rows (p99 19 and 119). Example 1's selected plan has nothing eligible, so there is no end-to-end case for it.Gate:
cargo fmt --all --check,cargo clippy --workspace --all-targets --all-features -- -D warnings,cargo test --workspace: 1,630 passed / 12 ignored (#600: 1,619 / 12). Viewer: 29 OK (6 skipped).Not in this PR
Links: #509, #580, #594 (Stage 3 cost), #601 (panes), #600.
🤖 Generated with Claude Code