Skip to content

perf: inspect runtime overhead against sketch and raw exact computation - #766

Open
zzylol wants to merge 64 commits into
issue-752from
perf/issue-758
Open

zzylol wants to merge 64 commits into
issue-752from
perf/issue-758

Conversation

@zzylol

@zzylol zzylol commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Problem and behavior

Before this PR: runtime/thread controls and a reproducible comparison of direct computation, installed-plan execution and HTTP overhead were missing.

After this PR: runtime controls and an inspection executable compare four paths over identical deterministic inputs: direct sketch, raw exact, backend bound-query execution and HTTP. Results must agree before measurements count as successful. Reports retain latency, throughput, request failures/drops, CPU, memory, threads, effective controls and cgroup limits.

Each cell now includes installed-plan.json, six separately measured setup stages, and an explicit list of included/excluded architecture stages. Setup time is separate from request latency. The fixture installs imported state: it does not benchmark Planner search/physical compilation, backend deployment compilation, precompute execution or durable recovery. Listed request stages are not individually timed. See docs/developer_docs/performance/overhead-inspection.md.

The inspection guide also explains how to retain the selected query-time or precomputed Rate/Sum/heap placement, including window width and stride. Runtime controls do not change those Planner decisions.

Stack

Based on #761. #766 and #756 are independent follow-ups; the acceptance stack does not depend on either. This PR does not require diagnostic logging to benchmark the production query path.

Validation and limits

  • Strict all-target Clippy passes after merging the current feat(control-plane): scope evidence and compute ERP/analytical workload costs #761 base. The earlier four inspection tests and four-layer smoke run below remain the fixture evidence; the new native Rate/Sum paths are not benchmarked by that synthetic fixture.
  • Four-layer smoke run: 10 successful queries per layer, one worker/concurrency, 1,000 samples; each cell retains the installed plan and six setup timings.
  • Debug-build checks validate behavior, not a release speedup. The fixture is a DDSketch median over one series and one complete pane.
  • Existing Level 3 performance failures are not resolved by this tooling PR.

Refs #758.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant