Skip to content

refactor(tools): organize post_experiment scripts; total vs query CPU; latency units - #742

Merged
milindsrivastava1997 merged 6 commits into
mainfrom
reorganize-post-experiment-scripts
Sep 24, 2026
Merged

milindsrivastava1997 merged 6 commits into
mainfrom
reorganize-post-experiment-scripts

Conversation

@milindsrivastava1997

@milindsrivastava1997 milindsrivastava1997 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

What changed and why

Organized the post-experiment scripts by what they analyze.
post_experiment/ was a flat directory of ~20 scripts, which made it hard to tell per-experiment analysis apart from cross-experiment figure generation. It is now split into:

  • single_experiment/: analyzes one experiment (cost, latency, fidelity, throughput)
  • multi_experiment/: sweeps many experiments to produce comparison plots
  • lib/: shared loaders
  • debug/: inspection tools

Imports, script paths and doc references were updated to match, and a short README explains the layout.

Made each cost number say which cost it is.
Figures could mix up two different CPU costs:

  • total CPU: ingest plus query, summed over all monitored processes
  • query CPU: estimated, and for Prometheus it depends on a guess at ingest cost

Cost labels now say "Total CPU" or "Query CPU", and the cardinality plots gain a total_cost option.

Fixed cost values that were silently wrong or missing.
The multi-experiment scripts used regexes to scrape numbers from the text output of other scripts. The scale scripts' default "cost p95" never matched for Prometheus (returned nothing), and for ASAPQuery matched only the query engine, leaving out Prometheus's own CPU. These scripts now read the JSON output of compare_costs.py / compare_latencies.py, so a format change raises an error instead of quietly dropping data. The default cost is now total CPU p95 for both systems.

Fixed latency units on the latency–cost tradeoff plot.
Latency is measured in seconds end to end, but the plot labeled it milliseconds. Absolute values were therefore shown 1000× too small; ratios were unaffected. Labels now say seconds.

Removed default argument values that could hide mistakes.
Many functions defaulted to baseline mode, p95 or query CPU even though every caller passes these values explicitly. A future call that forgot one would silently get the wrong numbers; it now raises an error. Parameters that were never varied were folded into the function body.

Kept one version of plot_scale_vs_metrics.py, the newer one that adds an --experiment_mode option. Also added the previously untracked plotting scripts to git.

Figures to regenerate

  • Anything from plot_scale_vs_metrics.py / plot_scale_vs_benefits.py using the default cost: it is now total CPU, where before it was missing or engine-only.
  • Latency–cost tradeoff plots: the axis unit is now correct.

How it was checked

  • Ran the old (regex) and new (JSON) code side by side on 7 local experiments. Latency and query-CPU values are identical. The only differences are the intended cost-p95 fix, and v1 of the cardinality plot now passing through infinite ratios the way v2 already did.
  • Confirmed the seconds unit: stored latencies match the query client's own log (median 0.376 s over 5,775 queries).
  • The throughput analyzer produces identical output before and after the default-argument cleanup.
  • Every moved script still runs from its new location. The pre-commit hooks pass.

🤖 Generated with Claude Code

Split post_experiment/ into single_experiment/ (per-experiment cost,
latency, fidelity, throughput analysis), multi_experiment/ (cross-
experiment comparison plots), lib/ (results_loader) and debug/.
Replace plot_scale_vs_metrics.py with its v2 (adds --experiment_mode)
and track plot_latency_cost_tradeoff.py, plot_scale_vs_benefits.py and
plot_cardinality_vs_benefit_v2.py. Fix sys.path, imports and sibling
script paths for the new layout; update doc references.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… query CPU

Replace regex parsing of compare_costs.py / run_compare_latencies.sh
text output in the multi_experiment scripts with --machine-readable
JSON. plot_scale_vs_metrics' default cost is now total CPU p95 (it
previously never matched for baseline and matched only the query
engine for sketchdb). Add --benefit-type total_cost to the
cardinality plots and name total vs query CPU in all cost labels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Query latency is recorded as time.time() differences (seconds) and is
never converted, but plot_latency_cost_tradeoff.py labeled it [ms].

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… values

- plot_latency_cost_tradeoff: only require the CPU stats for the chosen
  --cpu_type; drop unsupported --cost_metric mean.
- plot_scale_vs_metrics: return None with a warning when compare_costs
  omits query CPU instead of raising KeyError.
- plot_cardinality_vs_benefit(_v2): drop non-finite benefit ratios before
  plotting; inf made the y-tick loop unbounded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every caller already passes these arguments; removing the defaults makes
a forgotten argument (e.g. experiment_mode, metric, benefit_type, total)
an error instead of silently selecting baseline/p95/query CPU. Inline
verify_scale (always True), analyze_throughput's label_filter (always
the float-only filter) and calculate_stable_throughput's num_windows
(always self.num_windows).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@milindsrivastava1997
milindsrivastava1997 merged commit 1d825eb into main Sep 24, 2026
11 checks passed
@milindsrivastava1997
milindsrivastava1997 deleted the reorganize-post-experiment-scripts branch September 24, 2026 23:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant