Skip to content

feat(pass2): shared year-aligned window segments (Q60) - #628

Draft
zzylol wants to merge 1 commit into
stack/windows-1-b3from
stack/windows-2-segments
Draft

zzylol wants to merge 1 commit into
stack/windows-1-b3from
stack/windows-2-segments

Conversation

@zzylol

@zzylol zzylol commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Stack: #574 → #620 → #618 → #621 → #627 → #625 → #628 → #632 → #634 → #616 → #617 → #622 → #624 → #629 → #630 → #631 → #633 → #635 → #636 → #637

Problem

Example 3's Pattern A (five p99 reports over [5y], [1y], [1y offset 1y], [1y offset 2y] and [3y offset 2y]) and Example 4a (the same, repeated monthly) had no plan that shares window summaries across the reports. The spec's Exponential Histogram tests were ignored.

Changes

Stacked on #625. Decisions Q60–Q62.

  • Pass 2 shared-segment rule. share_window_segments in window_composition.rs groups approximate single-estimate targets over windows of one scan: same estimate up to accuracy, same grouping, no filters, offset ≥ 0. The segment width is the gcd of every window's lookback and offset, so every window boundary lies on the grid. For Pattern A that gives five 1-year segments.
    • Each query merges the segments its range covers, built by the existing tumbling-pane construction (WindowForm::Segments). Coverage is proved by SummaryCoverage::merge_disjoint, as for panes.
    • Composition builds identical segments, which the identical-expression rule merges, so each segment is one node read by every covering query's merge.
    • Merging disjoint KLLs is exact, so no ε is split. Each segment is sized for the strictest consumer, reusing the summary-capability rule's strictest and with_accuracy.
  • All-shared only (Q62). The rule's output is its own Stage 1 variant, Sharing::WindowSegments, labeled · shared segments. In it each grouped target has the segment form as its only alternative, so Pattern A gains exactly one logical candidate (486 → 487). There are no partial or pairwise groupings.
  • Skips. A group is skipped when its windows are all identical (the summary-capability rule's case), when the union has more than MAX_PANES (64) segments, or when its queries recur at different cadences.
  • Materialization (Q61). Unchanged. 3a, a one-off ad hoc batch, gets no ingestion-time option. 4a, monthly and predictable, gets one ingestion-time option for the segments through Stage 2's existing eligibility, so 4a has 488 plans. No kept option is offered, because the 1-year segment width is not the monthly cadence.
  • Display. duration_label gains days, so segments read 365d segments.
  • Viewer stories. The 3a and 4a stories are updated.

Tests

  • Rewritten and un-ignored:
    • stage1_a_window_composition_adds_one_eh_for_all_five becomes stage1_a_shared_segments_serve_all_five: five 1-year KLL segments together read by all five queries, each query covering as many segments as its range has years, and five p99 estimates each reading a merge.
    • stage1_a_keeps_independent_and_shared_window_summaries (Example 3).
    • stage2_a_materialized_eh_is_built_once_for_all_consumers becomes stage2_a_segments_are_built_once_for_all_consumers: five builds read by all five queries through five merges, charged once. Its options are {not materialized} for the one-off batch and {ingestion time, not materialized} for monthly.
  • Still ignored, with new reasons:
    • stage1_a_window_composition_groups_every_pair: no partial or pairwise groupings (Q62).
    • stage2_a_shared_eh_has_three_materialization_options, stage2_a_at_rest_drops_the_ingestion_time_option, stage3_a_rebuilding_per_query_costs_more_than_building_once, stage3_a_once_adhoc_prefers_the_query_time_eh and stage3_a_monthly_amortizes_ingestion_time_maintenance. These expect A1 (kept for one batch), A3 (Q44) or a one-off ingestion-time option (Q61), none of which is generated.
    • No remaining test is EH-only, so none uses the "true EH later" reason.
  • New:
    • pattern_a_shares_five_one_year_segments and segments_need_one_cadence_and_different_windows, unit tests in Pass 2.
    • shares_segments in the test adapter.
    • Example 3's raw-input-only and shared-scan tests skip the segment candidate.
    • The devtool test example4a_repeats_monthly_with_nothing_maintainable becomes ..._with_only_segments_maintainable: 488 plans, the only ingestion-time one being the segments.
  • Example 1 fixture: unchanged. Example 1 has no group.

Finding: segments do not win under today's cost model

The 3a segment plan costs 44 384/s against 39 128/s for the selected shared-input KLLs. In 4a it is 61.64/s at query time and 1 536.9/s at ingestion time, against 54.34/s.

The plan builds 5 years of rows into KLLs where the independent KLLs build 11 years. It loses on pass-through work instead: Stage 3 charges each query-time TimeShift and TimeRange per row of its input, and every segment's shift sees the whole 5-year scan. That is 10 such nodes against 8 in the selected plan, at 2 920 each.

Pricing a query-time shift as free, and a range on the rows it keeps (a time-ordered scan seeks), would make segments cheapest in 3a: roughly 18 100 against 25 100. That pricing change is not made here.

Test plan

  • cargo fmt --all --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo test --workspace: 1689 passed, 0 failed, 13 ignored (re-run after the rebase on main d4869a7)
  • python3 -m unittest test_viewer (tools/dag-viewer): 25 OK
  • Regenerated the six example documents. Every selection is unchanged: Example 1 P82, Example 2 P86, 3a P365, 3b P2, 4a P365, 4b P4-m1.

🤖 Generated with Claude Code

Pass 2's window-composition rule gains shared segments: approximate
single-estimate targets over windows of one scan (same estimate up to
accuracy, same grouping, no filters) share one summary per segment of a
grid whose width is the gcd of every window's lookback and offset. Each
query merges the segments its range covers; composition builds identical
segments, which the identical-expression rule merges. Merging disjoint
KLLs is exact, so each segment is sized for the strictest consumer and no
accuracy is split.

Only the all-shared form is offered (Q62), as its own Stage 1 variant
(Sharing::WindowSegments, "· shared segments") in which each grouped
target has the segment form as its only alternative. A group is skipped
when its windows are identical, the union has more than MAX_PANES
segments, or its queries recur at different cadences. Example 3a gets one
candidate with five 1-year KLL segments; 4a, repeated monthly, can also
maintain them at ingestion time (Q61).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant