Version / build tested against
origin/main @ e235fe5 (2026-09-29)
Deployment mode
Origin — single node (local)
Engine(s) involved
Timeseries
Summary
Continuous aggregation advances its high watermark and scans only [high_watermark, now()). Late-arriving samples below the watermark in valid columnar partitions are skipped, and no re-aggregation path exists (clear_o3 is dead code) — aggregates stay permanently skewed.
Steps to reproduce
Reproduction path (runtime reproduction pending): write a sample behind the watermark (O3 window), run the maintenance tick, observe the aggregate never updates. rg -n "clear_o3|o3_min" nodedb/src/engine/timeseries/continuous_agg.
Expected behavior
Out-of-order writes below the watermark are detected and the affected older windows re-aggregate on a maintenance tick (e.g. dirty-window tracking), without moving the monotonic ingestion watermark.
Actual behavior
Watermark advances to max(sample.ts); late samples are silently excluded from rolling aggregation forever.
What actually happened? (severity facts)
Proposed severity
SEV-2 — High: silently-wrong aggregate results while stored data is intact.
Reproducibility
Always — every attempt (given a late sample)
Last known-good version / commit (if a regression)
(unknown / not a regression)
Environment & logs
Linux x86_64; Verified by static code reading at the pin above; runtime reproduction pending.
Code references:
- (see prior-art line below)
Before submitting
Additional evidence (origin/main @ e235fe55c)
- What: The continuous-aggregate engine detects out-of-order (O3) samples below the watermark and records the oldest into
WatermarkState.o3_watermark_ts, but nothing ever reads that field to re-aggregate the affected buckets. WatermarkState::clear_o3 has zero callers.
- Where:
nodedb/src/engine/timeseries/continuous_agg/watermark.rs:14-19 (o3_watermark_ts), :58-61 (clear_o3, no callers); .../continuous_agg/refresh.rs:82-89 (O3 detection); .../continuous_agg/manager.rs:146-148,192-194 (record_o3 only); nodedb/src/engine/timeseries/merge/o3.rs:21 (merge_o3_into_partition, no caller).
- Evidence:
rg -n 'clear_o3' → definition only. rg -n 'o3_watermark_ts' → writes + one test read (manager.rs:543); no consumer. The field doc says "On next refresh, re-aggregate buckets between o3_watermark and watermark", but no code does this.
- Impact: Latent. At HEAD late rows that reach the columnar memtable are merged by the flush-driven refresh, so materialized buckets stay correct for the drain path. The gap bites if/when the O3 buffer merge path (
merge_o3_into_partition, per-partition) is wired to ingest and lands rows directly in sealed partitions: those rows would never reach refresh_from_drain and the old buckets would not be recomputed. The declared-but-unimplemented re-aggregation is a correctness trap for that wiring.
- Fix: On a maintenance tick, when
o3_watermark_ts is Some(w), re-scan sealed partitions in [w, watermark_ts], recompute the affected buckets, then clear_o3(). Alternatively delete o3_watermark_ts/clear_o3 until the re-aggregation consumer lands, so the intent is not implied by dead state.
- Note: The originally filed claim ("driver scans only
[high_watermark, now()) and skips late samples") does not match HEAD: refresh_from_drain has no watermark scan window and merges every drained row. File this narrower issue instead.
- Prior-art: #333 (open) covers ROLLUP/CUBE/HAVING/GROUP-BY-expr gaps — a different scope; O3 late-sample re-aggregation is uncovered. No other open issue matches.
Why: late-arriving samples below the watermark are dropped from the rolling scan and no re-aggregation path exists (clear_o3 is dead code), so aggregates stay permanently skewed.
Steps to verify: write a late sample behind the watermark, run the maintenance tick, observe the aggregated value never updates; rg -n "clear_o3|o3_min" nodedb/src/engine/timeseries/continuous_agg.
Version / build tested against
origin/main @ e235fe5 (2026-09-29)
Deployment mode
Origin — single node (local)
Engine(s) involved
Timeseries
Summary
Continuous aggregation advances its high watermark and scans only
[high_watermark, now()). Late-arriving samples below the watermark in valid columnar partitions are skipped, and no re-aggregation path exists (clear_o3is dead code) — aggregates stay permanently skewed.Steps to reproduce
Reproduction path (runtime reproduction pending): write a sample behind the watermark (O3 window), run the maintenance tick, observe the aggregate never updates.
rg -n "clear_o3|o3_min" nodedb/src/engine/timeseries/continuous_agg.Expected behavior
Out-of-order writes below the watermark are detected and the affected older windows re-aggregate on a maintenance tick (e.g. dirty-window tracking), without moving the monotonic ingestion watermark.
Actual behavior
Watermark advances to max(sample.ts); late samples are silently excluded from rolling aggregation forever.
What actually happened? (severity facts)
Proposed severity
SEV-2 — High: silently-wrong aggregate results while stored data is intact.
Reproducibility
Always — every attempt (given a late sample)
Last known-good version / commit (if a regression)
(unknown / not a regression)
Environment & logs
Linux x86_64; Verified by static code reading at the pin above; runtime reproduction pending.
Code references:
Before submitting
mainbuild (pending) — code path verified ate235fe55c.Additional evidence (origin/main @
e235fe55c)WatermarkState.o3_watermark_ts, but nothing ever reads that field to re-aggregate the affected buckets.WatermarkState::clear_o3has zero callers.nodedb/src/engine/timeseries/continuous_agg/watermark.rs:14-19(o3_watermark_ts),:58-61(clear_o3, no callers);.../continuous_agg/refresh.rs:82-89(O3 detection);.../continuous_agg/manager.rs:146-148,192-194(record_o3only);nodedb/src/engine/timeseries/merge/o3.rs:21(merge_o3_into_partition, no caller).rg -n 'clear_o3'→ definition only.rg -n 'o3_watermark_ts'→ writes + one test read (manager.rs:543); no consumer. The field doc says "On next refresh, re-aggregate buckets between o3_watermark and watermark", but no code does this.merge_o3_into_partition, per-partition) is wired to ingest and lands rows directly in sealed partitions: those rows would never reachrefresh_from_drainand the old buckets would not be recomputed. The declared-but-unimplemented re-aggregation is a correctness trap for that wiring.o3_watermark_tsisSome(w), re-scan sealed partitions in[w, watermark_ts], recompute the affected buckets, thenclear_o3(). Alternatively deleteo3_watermark_ts/clear_o3until the re-aggregation consumer lands, so the intent is not implied by dead state.[high_watermark, now())and skips late samples") does not match HEAD:refresh_from_drainhas no watermark scan window and merges every drained row. File this narrower issue instead.Why: late-arriving samples below the watermark are dropped from the rolling scan and no re-aggregation path exists (
clear_o3is dead code), so aggregates stay permanently skewed.Steps to verify: write a late sample behind the watermark, run the maintenance tick, observe the aggregated value never updates;
rg -n "clear_o3|o3_min" nodedb/src/engine/timeseries/continuous_agg.