Skip to content

Stage 2 materialization: ingestion vs query time per summary - #604

Draft
zzylol wants to merge 1 commit into
stack/509-x8a-hydra-kernelfrom
stack/509-x5-stage2-materialization
Draft

zzylol wants to merge 1 commit into
stack/509-x8a-hydra-kernelfrom
stack/509-x5-stage2-materialization

Conversation

@zzylol

@zzylol zzylol commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Stack: Wave 2 chain: #599 → #601 → #600 → #604 → #606 → #603 → #605

Rebased on main d4869a7 (DF 54).

Stacked on #600. Part of #580, step 5; #509 Stage 2 "Materialization". Implements decisions S1–S6 and Q44 (pricing #594, panes #601).

Why

Stage 2 ran every summary at query time, so Stage 3 never compared maintaining state at ingestion time with rebuilding it at each evaluation. Example 4's B1 (tumbling KLLs at ingestion time) was not generated. Query latency bounds were not checked.

What

  1. Eligibility (S4). A SummaryAgg may run at ingestion time when all of these hold:

    • DataWorkload.arrival is ContinuouslyIngesting or Mixed. Unknown counts as not ingesting.
    • Every root that reaches it repeats and is Predictable.
    • It is a tumbling pane, or it reads one fixed window per evaluation: FixedIntervalAt, with a window no longer than the interval.

    A summary over a merge of panes reads a sliding window, so it stays at query time. The panes of one merge form one unit.

  2. Down-closed sets. Stage 2 returns one physical candidate per set of units in which everything below an ingestion-time unit is also at ingestion time. The all-query-time candidate comes first. Ids are P<n> for all query time and P<n>-m<k> otherwise; labels end in "· ingestion time: Kll ×5 panes". Above 16 sets, Stage 2 adds units greedily by Stage 3's cost and flags the selection as not guaranteed optimal.

  3. Cross-root phase split. split_shared_by_phase(&[roots], assignment) copies a node per phase when one consumer reads it at ingestion time and another at query time, across queries. It returns the assignment re-keyed to the copies.

  4. Selection. The tree DP's evaluate, enumeration and finish take the cheapest valid physical candidate of each logical candidate. DP-versus-exhaustive tests run with materialization on.

  5. Pricing panes. At ingestion time, a pane chain is one pane, built as rows arrive and kept. The newest pane pays the λ-rate build and the memory of lookback/width + 1 panes (S3: in memory only). Older panes, and the shifts and ranges that feed only them, cost 0.

  6. Latency check (S6). A candidate is rejected when a query's query-time work for one evaluation is above its latency_ms. The work counts every query-time node the query reaches, shared nodes included. New Stage3Calibration::latency_ms_per_cost_unit = 1 (one CPU-ms on one core). Example reason: q1: query-time work takes 310.0 ms per evaluation, over the 200 ms latency bound. RootDemand now carries latency_ms.

  7. Docs. stage3-cost-model.md covers pane retention, the latency check and the calibration field, and updates Example 1's worked example.

Before / After (built-in models, constants untuned)

Before (#600) After
Example 1 physical candidates 88 112: 24 also maintain Q2's 10-s sum_over_time panes
Example 1 selected P82, 4.620 cost/s P82, 4.620 (unchanged)
Example 1 cheapest maintained — P39-m1, 47.407: scan 0.387, shift + range 0.133, newest-pane build 0.067, memory 42.0 (7 panes × 1M sums = 336 MB), plus Q1 and the query-time merge, finalize and CMS
Example 1 latency not checked 40 candidates over Q2's 100 ms bound, every Count-Sketch + heap among them
Example 3B candidates 5 7: KLL and DDSketch panes at ingestion time
Example 3B selected P2 KLL, 2.083 P2 KLL, 2.083 (unchanged)
Example 3B tumbling at query time 5.167, valid rejected: 310 ms, over the 200 ms bound
Example 3B tumbling at ingestion time (B1) not generated P4-m1, 768.58: build 0.067, scan/shift/range 0.41, merge + estimate 0.10, memory 768.0 (6 panes × 1M series × 1 KiB KLL = 6.1 GB)
Example 3A / Example 4 Pattern A P365 unchanged. Run once ad hoc, or monthly without a phase and with windows longer than the interval, or at rest: nothing is eligible

Ingestion time does not win any example under the built-in models. With 1M series, memory for retained per-series state is much larger than the CPU it saves. In a unit test with 10 series at 100k rows/s, maintained panes win, and the DP selects them too.

Tests

  • timing.rs: a cross-root split keeps the assignment, and a node reached at one timing stays shared.
  • materialization.rs: panes go to ingestion time and the merge stays at query time. Arrival, ad hoc and one-off rules. A whole window needs a phase and must not overlap. Sets are down-closed: an exact sum plus a CMS gives 3 sets, and over panes only the panes count. Greedy above the cap.
  • plan-selection: the newest pane pays the build plus (N+1)× memory, and the older panes pay 0. The latency bound rejects slow candidates with a reason. Maintained panes win with few series, and the DP agrees with exhaustive selection.
  • Example 1: counts are 88 → 112. Only the 6 sum panes and their inputs run at ingestion time, and the scan is split from Q1's. New latency test. Ranking and sharing tests cover all-query-time candidates.
  • Example 3: query-time tumbling misses the latency bound, and the maintained form is valid. End to end: runtime_b_maintained_panes_match_the_whole_window compiles the B1 candidate once. It cuts the candidate at frontier_from_timing with cut_candidate, runs the precompute once per pane in that pane's Scope::Ingestion window, then runs the query over the 5 kept panes. It returns the whole-window rows (p99 19 and 119). Example 1's selected plan has nothing eligible, so there is no end-to-end case for it.
  • Example 1 fixture regenerated: 112 physical candidates, 12 MB (was 9.3 MB).

Gate: cargo fmt --all --check, cargo clippy --workspace --all-targets --all-features -- -D warnings, cargo test --workspace: 1,630 passed / 12 ignored (#600: 1,619 / 12). Viewer: 29 OK (6 skipped).

Not in this PR

  • A3 / Q44: "not materialized" for a sub-DAG with several consumers, generated by duplicating it per consumer. Follow-up.
  • B3: a pane built at query time and kept for later evaluations. The newest pane would need to be marked.
  • Sliding windows, storage tier and retention (S3), backfill (S5).
  • A maintained state does not record its retention in the IR yet. Pricing derives it from the merge.

Links: #509, #580, #594 (Stage 3 cost), #601 (panes), #600.

🤖 Generated with Claude Code

zzylol added a commit that referenced this pull request Oct 4, 2026
RootDemand gained latency_ms in #604; the Example 2 coverage test's
demand sets it to None (no latency bound).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Oct 4, 2026
#604 gave RootDemand a latency_ms and a logical candidate several
physical candidates. The Hydra count test sets no latency bound and checks
the all-query-time physical candidate, which comes first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ummary

#509 Stage 2 materialization (#580 step 5, decisions S1-S6, Q44):

- Eligibility (S4): a SummaryAgg may run at ingestion time when the data
  is continuously ingesting or mixed, every root reaching it repeats and is
  predictable, and it is a tumbling pane or reads one fixed window per
  evaluation (FixedIntervalAt, window no longer than the interval). The
  panes of one merge are one unit.
- Stage 2 enumerates the down-closed sets of units (all query time first),
  each built with a workload-wide split_shared_by_phase that copies a node
  shared by an ingestion-time and a query-time consumer per phase and
  re-keys the assignment. Above 16 sets it searches greedily by Stage 3's
  cost and flags the selection as not guaranteed optimal.
- Stage 3 takes the cheapest valid physical candidate of each logical one,
  in the DP and in enumeration. Ingestion-time panes are one pane built as
  rows arrive: the newest pane pays the build and retains N + 1 panes; the
  older panes and their inputs cost nothing.
- Latency check (S6): a query whose query-time work per evaluation exceeds
  its latency_ms is rejected with the estimate and the bound
  (Stage3Calibration::latency_ms_per_cost_unit, 1 CPU-ms per ms).
- RootDemand carries latency_ms.

Example 1: 112 physical candidates (24 maintain Q2's sum panes), same
selection P82 at 4.620; maintained panes cost 47.4 (42.0 memory); 40
candidates over Q2's 100 ms bound. Example 3B: P2 KLL still selected at
2.083; tumbling panes at query time miss the 200 ms bound, maintained ones
cost 768.6 (768.0 memory). The executor runs Example 3B's maintained plan
per pane and returns the whole-window rows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zzylol
zzylol force-pushed the stack/509-x8a-hydra-kernel branch from 09fa09e to f97037a Compare October 5, 2026 06:21
@zzylol
zzylol force-pushed the stack/509-x5-stage2-materialization branch from ec82a21 to 78cf83f Compare October 5, 2026 06:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant