Skip to content

Benchmark sustained OTLP streaming and real network resource use - #19

Open
zzylol wants to merge 22 commits into
mainfrom
bench/scale-traffic-kll-resources
Open

zzylol wants to merge 22 commits into
mainfrom
bench/scale-traffic-kll-resources

Conversation

@zzylol

@zzylol zzylol commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR adds an OTAP-based benchmark to compare raw forwarding, exact quantile aggregation, and KLL-based approximate aggregation using production KLLWrapper with k=400.

The benchmark uses OTAP's upstream urn:otel:receiver:traffic_generator and independently sweeps:

  • Traffic rate: 10k, 25k, 50k, 100k, 200k, 300k, 400k signals/s per source
  • Window size: 16,384, 65,536, 262,144 observations/source/window

Two source containers generate traffic from a shared configurable generator CPU pool. Branch A, branch B, merge, and estimate each receive one dedicated CPU core. The validating backend receives all remaining host cores.

System architecture

flowchart LR
  subgraph GP["Generator CPU pool: N cores"]
    GA["Generator A<br/>N OTAP workers"]
    GB["Generator B<br/>N OTAP workers"]
  end

  BA["Branch A<br/>1 exclusive core<br/>normalize + create"]
  BB["Branch B<br/>1 exclusive core<br/>normalize + create"]
  M["Merge<br/>1 exclusive core"]
  E["Estimate<br/>1 exclusive core"]
  V["Validating backend<br/>all remaining host cores"]

  GA -->|"OTLP/HTTP protobuf"| BA
  GB -->|"OTLP/HTTP protobuf"| BB
  BA -->|"raw / exact / KLL"| M
  BB -->|"raw / exact / KLL"| M
  M -->|"merged window"| E
  E -->|"raw / p50 + p99"| V
Loading

The three scenarios operate on the same deterministic input values:

  • Raw: forward observations without quantile computation.
  • Exact: sort branch values, merge the two sorted runs, and compute p50/p99.
  • KLL: incrementally create KLL sketches, merge them with KLLWrapper::merge, and estimate p50/p99.

All data-plane communication uses uncompressed OTLP/HTTP protobuf over persistent TCP connections. KLL windows use the self-describing ASAPv1 Msgpack representation.

A separate control network carries readiness, run/shutdown commands, and measurement snapshots. Control traffic is excluded from data-plane RX/TX measurements.

Default nightly configuration

Setting Value
Scenarios raw, exact, KLL
Offered rates 10k–400k signals/s per source
Window sizes 16,384; 65,536; 262,144 observations/source
Generator 2 containers, 2 OTAP workers/container by default
Evaluated processors 1 exclusive core each: branch A, branch B, merge, estimate
Validating backend All remaining host cores; excluded from evaluated resource budget
Batch size 1,024
Warm-up / observation 5 s / 30 s
Repetitions 3
Matrix 189 runs
Transport OTLP/HTTP protobuf, uncompressed, persistent
Memory Uncapped for every component; sampled and lifetime peak RSS reported

Each run records its CPU mapping, image ID, host/CPU information, workload configuration, interfaces, and measurement settings in config.json.

The benchmark reports per-component CPU, sampled and lifetime peak RSS, data-plane RX/TX, throughput, latency, and correctness. resources.svg summarizes CPU and memory usage across components and configurations.

Optional scoped CPU accounting and perf capture further separate pdata codec, sketch codec, computation, and processor bookkeeping costs.

Capacity evaluation

The primary result is based on maximum sustainable capacity, rather than a throughput ratio at one offered load.

A point is sustainable when the three-run median:

  • delivers at least 95% of offered traffic,
  • grows backlog by at most one paired window,
  • keeps p99 latency at or below 5 seconds, and
  • passes correctness in every repetition.

capacity(S, W) is the highest backend-validated throughput among sustainable points for scenario S and window size W.

The primary comparison is:

capacity(KLL, W) / capacity(Exact, W)

Raw forwarding is reported separately as the transport baseline. Results are written to capacity.csv and capacity.json.

Single-run Exact/KLL results: uncapped-memory rerun

Every aggregation processor has one exclusive core. Containers have no Docker memory cap. The validating backend uses the remaining host CPUs and is excluded from the Exact/KLL resource budget. KLL remains fixed at k=400.

Maximum observed throughput plateau

Observations/source/window Exact KLL KLL / Exact
16,384 59,978/s 4,905,493/s 81.79x
65,536 74,156/s 5,413,621/s 73.00x
262,144 69,799/s 5,601,596/s 80.25x

These are one-run maximum observed throughput results. Each peak has a higher offered-load point that fails to increase throughput. KLL branch CPU reaches approximately one full core at the selected peaks; four generator cores are therefore sufficient to expose the downstream plateau.

Matched workload resource efficiency

Both scenarios use 25,000 signals/s/source, identical values, transport, batch size, CPU placement, and observation rules. Pipeline resources sum branch A, branch B, merge, and estimate; generator and backend are excluded.

Window/source Scenario Validated throughput Pipeline CPU Core-s / M signals Pipeline sampled peak RSS
16,384 Exact 50,169/s 1.227 cores 24.46 141.8 MiB
16,384 KLL 50,164/s 0.057 cores 1.14 124.5 MiB
65,536 Exact 47,987/s 1.142 cores 23.79 179.2 MiB
65,536 KLL 47,982/s 0.050 cores 1.05 123.2 MiB
262,144 Exact 33,288/s 0.913 cores 27.43 305.8 MiB
262,144 KLL 49,932/s 0.046 cores 0.93 127.0 MiB

Per-component CPU at the throughput peaks

Values are average cores used during observation; 1.000 is one fully occupied core.

Scenario Window/source Offered/source generator_a generator_b branch_a branch_b merge estimate backend
Exact 16,384 200,000/s 0.019 0.018 0.277 0.279 0.654 0.275 0.002
KLL 16,384 4,000,000/s 0.624 0.586 0.987 0.982 0.360 0.256 0.085
Exact 65,536 200,000/s 0.021 0.020 0.307 0.307 0.802 0.269 0.001
KLL 65,536 2,800,000/s 0.634 0.664 0.989 0.987 0.108 0.073 0.026
Exact 262,144 400,000/s 0.024 0.025 0.324 0.325 0.814 0.246 0.001
KLL 262,144 2,800,000/s 0.680 0.665 0.991 0.990 0.028 0.021 0.008

Per-component sampled peak RSS at the throughput peaks

Scenario Window/source generator_a generator_b branch_a branch_b merge estimate backend
Exact 16,384 78.2 MiB 78.2 MiB 42.3 MiB 42.3 MiB 38.7 MiB 30.9 MiB 24.3 MiB
KLL 16,384 1065.4 MiB 1065.5 MiB 32.9 MiB 33.0 MiB 30.6 MiB 29.6 MiB 25.1 MiB
Exact 65,536 77.9 MiB 78.2 MiB 46.3 MiB 48.1 MiB 54.5 MiB 31.8 MiB 24.2 MiB
KLL 65,536 754.2 MiB 755.0 MiB 34.1 MiB 34.0 MiB 29.9 MiB 29.7 MiB 24.6 MiB
Exact 262,144 128.5 MiB 129.8 MiB 78.2 MiB 79.1 MiB 120.3 MiB 34.6 MiB 24.6 MiB
KLL 262,144 754.3 MiB 754.2 MiB 37.5 MiB 37.7 MiB 29.7 MiB 29.5 MiB 24.5 MiB

Per-component data-plane network at the throughput peaks

Cells are RX/TX Mbit/s; control-network traffic is excluded.

Scenario Window/source generator_a generator_b branch_a branch_b merge estimate backend
Exact 16,384 0.1/65.7 0.1/65.1 65.8/25.3 65.2/25.4 50.7/50.6 50.4/0.1 0.0/0.0
KLL 16,384 4.8/5360.5 4.8/5307.9 5360.9/15.3 5308.2/15.2 21.2/4.5 4.2/1.9 1.6/0.2
Exact 65,536 0.1/81.0 0.1/81.0 81.1/31.2 81.1/31.2 62.4/62.3 62.2/0.2 0.0/0.0
KLL 65,536 5.3/5867.7 5.3/5869.3 5867.8/7.2 5869.3/7.2 3.8/1.9 1.8/0.5 0.5/0.1
Exact 262,144 0.1/85.0 0.1/84.3 85.1/29.4 84.3/29.4 58.7/58.7 58.5/0.1 0.0/0.0
KLL 262,144 5.5/6072.4 5.5/6072.4 6072.6/6.0 6072.3/6.1 1.1/0.5 0.4/0.1 0.1/0.0

The repeated nightly evaluation is still required before treating these one-run measurements as final capacity claims. The sustainable classifier is especially sensitive to completed-window quantization at the largest Exact window, so the plateau comparison and matched-load resource table are reported separately.

Exact bottleneck diagnosis

A targeted --profile-cpu run reproduced the Exact plateau at 58.9k, 74.2k, and 69.8k signals/s for the three windows. Across those runs, the four evaluated processors consumed 22.23 core-seconds/M signals:

CPU category Core-s / M signals Share
pdata codec 16.64 74.87%
processor bookkeeping 3.32 14.94%
runtime/transport outside scopes 2.14 9.64%
Exact sort, merge, quantile computation 0.12 0.55%

The merge role is the largest CPU consumer at 10.67 core-s/M signals; 8.24 of that is pdata codec and only 0.02 is numeric merge computation. Exact is therefore dominated by materializing, decoding, encoding, and transporting full-data pdata rather than sorting or merging values.

Wall-time counters at 65,536 observations/source/window, batch=1,024, and 200k offered/source explain why no process reaches one full CPU core:

Role CPU cores Processor wall Output-send wait Accounted wall share
Branch A 0.308 8.15 s 20.54 s 95.5%
Branch B 0.305 8.09 s 20.58 s 95.4%
Merge 0.798 22.22 s 7.50 s 98.9%
Estimate 0.274 7.95 s 0.001 s 26.4%

The merge stage occupies essentially the full observation interval: about three quarters performing full-data processor/codec work and one quarter awaiting capacity in the processor-to-exporter path. Both branches spend about two thirds of wall time blocked on output because pressure propagates back from merge. Estimate mostly waits for input. The Exact bottleneck is therefore full-data pdata codec at merge combined with serialized output backpressure, rather than the Exact numeric algorithm.

Increasing the batch from 1,024 to 4,096 did not raise throughput (74.2k to 69.8k/s), and 16,384 reduced it further, so the limit is not merely requests per second. Removing intermediate full-data pdata encode/decode boundaries or introducing a native sorted-run representation is the next optimization target.

Scope

This benchmark measures same-host TCP traffic across Linux network namespaces; it does not claim physical WAN behavior.

Count-aligned windows ensure that all three scenarios operate on identical values. They do not implement production event-time rotation or late-event handling.

The full 7 rates × 3 windows × 3 scenarios × 3 repetitions = 189 runs evaluation provides the final sustainable-capacity comparison.

See the capacity experiment design, streaming benchmark reference, and profiling guide.

@zzylol zzylol changed the title Benchmark KLL scaling and per-component resource use Benchmark sustained OTLP streaming and real network resource use Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant