Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR adds an OTAP-based benchmark to compare raw forwarding, exact quantile aggregation, and KLL-based approximate aggregation using production
KLLWrapperwithk=400.The benchmark uses OTAP's upstream
urn:otel:receiver:traffic_generatorand independently sweeps:10k, 25k, 50k, 100k, 200k, 300k, 400ksignals/s per source16,384, 65,536, 262,144observations/source/windowTwo source containers generate traffic from a shared configurable generator CPU pool. Branch A, branch B, merge, and estimate each receive one dedicated CPU core. The validating backend receives all remaining host cores.
System architecture
flowchart LR subgraph GP["Generator CPU pool: N cores"] GA["Generator A<br/>N OTAP workers"] GB["Generator B<br/>N OTAP workers"] end BA["Branch A<br/>1 exclusive core<br/>normalize + create"] BB["Branch B<br/>1 exclusive core<br/>normalize + create"] M["Merge<br/>1 exclusive core"] E["Estimate<br/>1 exclusive core"] V["Validating backend<br/>all remaining host cores"] GA -->|"OTLP/HTTP protobuf"| BA GB -->|"OTLP/HTTP protobuf"| BB BA -->|"raw / exact / KLL"| M BB -->|"raw / exact / KLL"| M M -->|"merged window"| E E -->|"raw / p50 + p99"| VThe three scenarios operate on the same deterministic input values:
KLLWrapper::merge, and estimate p50/p99.All data-plane communication uses uncompressed OTLP/HTTP protobuf over persistent TCP connections. KLL windows use the self-describing ASAPv1 Msgpack representation.
A separate control network carries readiness, run/shutdown commands, and measurement snapshots. Control traffic is excluded from data-plane RX/TX measurements.
Default nightly configuration
Each run records its CPU mapping, image ID, host/CPU information, workload configuration, interfaces, and measurement settings in
config.json.The benchmark reports per-component CPU, sampled and lifetime peak RSS, data-plane RX/TX, throughput, latency, and correctness.
resources.svgsummarizes CPU and memory usage across components and configurations.Optional scoped CPU accounting and
perfcapture further separate pdata codec, sketch codec, computation, and processor bookkeeping costs.Capacity evaluation
The primary result is based on maximum sustainable capacity, rather than a throughput ratio at one offered load.
A point is sustainable when the three-run median:
capacity(S, W)is the highest backend-validated throughput among sustainable points for scenarioSand window sizeW.The primary comparison is:
capacity(KLL, W) / capacity(Exact, W)Raw forwarding is reported separately as the transport baseline. Results are written to
capacity.csvandcapacity.json.Single-run Exact/KLL results: uncapped-memory rerun
Every aggregation processor has one exclusive core. Containers have no Docker memory cap. The validating backend uses the remaining host CPUs and is excluded from the Exact/KLL resource budget. KLL remains fixed at
k=400.Maximum observed throughput plateau
These are one-run maximum observed throughput results. Each peak has a higher offered-load point that fails to increase throughput. KLL branch CPU reaches approximately one full core at the selected peaks; four generator cores are therefore sufficient to expose the downstream plateau.
Matched workload resource efficiency
Both scenarios use
25,000 signals/s/source, identical values, transport, batch size, CPU placement, and observation rules. Pipeline resources sum branch A, branch B, merge, and estimate; generator and backend are excluded.Per-component CPU at the throughput peaks
Values are average cores used during observation;
1.000is one fully occupied core.Per-component sampled peak RSS at the throughput peaks
Per-component data-plane network at the throughput peaks
Cells are
RX/TX Mbit/s; control-network traffic is excluded.The repeated nightly evaluation is still required before treating these one-run measurements as final capacity claims. The sustainable classifier is especially sensitive to completed-window quantization at the largest Exact window, so the plateau comparison and matched-load resource table are reported separately.
Exact bottleneck diagnosis
A targeted
--profile-cpurun reproduced the Exact plateau at58.9k,74.2k, and69.8k signals/sfor the three windows. Across those runs, the four evaluated processors consumed22.23 core-seconds/M signals:The merge role is the largest CPU consumer at
10.67 core-s/M signals;8.24of that is pdata codec and only0.02is numeric merge computation. Exact is therefore dominated by materializing, decoding, encoding, and transporting full-data pdata rather than sorting or merging values.Wall-time counters at
65,536 observations/source/window,batch=1,024, and200k offered/sourceexplain why no process reaches one full CPU core:The merge stage occupies essentially the full observation interval: about three quarters performing full-data processor/codec work and one quarter awaiting capacity in the processor-to-exporter path. Both branches spend about two thirds of wall time blocked on output because pressure propagates back from merge. Estimate mostly waits for input. The Exact bottleneck is therefore full-data pdata codec at merge combined with serialized output backpressure, rather than the Exact numeric algorithm.
Increasing the batch from
1,024to4,096did not raise throughput (74.2kto69.8k/s), and16,384reduced it further, so the limit is not merely requests per second. Removing intermediate full-data pdata encode/decode boundaries or introducing a native sorted-run representation is the next optimization target.Scope
This benchmark measures same-host TCP traffic across Linux network namespaces; it does not claim physical WAN behavior.
Count-aligned windows ensure that all three scenarios operate on identical values. They do not implement production event-time rotation or late-event handling.
The full
7 rates × 3 windows × 3 scenarios × 3 repetitions = 189 runsevaluation provides the final sustainable-capacity comparison.See the capacity experiment design, streaming benchmark reference, and profiling guide.