Skip to content

perf: hardware SHA-256, single classification per distinct line, fat LTO - #8

Merged
maxgfr merged 2 commits into
mainfrom
perf-engine
Sep 11, 2026
Merged

maxgfr merged 2 commits into
mainfrom
perf-engine

Conversation

@maxgfr

@maxgfr maxgfr commented Sep 11, 2026

Copy link
Copy Markdown
Owner

Summary

Profiling the 32 MiB capture on 0.5.1 gave: parse 5 ms, view 32 ms, blob 63 ms, artifact 70 ms. Three changes, all output-identical:

  • Hardware SHA-256. sha2 only uses ARMv8 / SHA-NI instructions with its asm feature. The 32 MiB digest drops from 55 ms to 10 ms; every capture hashes twice and every read verifies, so this is the largest single cost.
  • One classification per distinct line. The compact-v3 pass ran two vocabulary regexes on every line and built a template for every non-diagnostic line. Identical text is now classified once and repetitions copy the result by index. The template builder scans bytes and only decodes non-ASCII characters; a unit test proves it equal to the previous character-by-character version on Unicode whitespace, control bytes, digit and hex runs. Line texts are hashed with foldhash instead of SipHash.
  • Release profile. Fat LTO, one codegen unit, abort on panic: a third smaller binary and a few percent per call.

Also: bench/publish.py measures one binary and publishes the only figures the repository keeps (bench/results/content.json, bench/results/performance.json, and the README speed table between its markers, stamped with the version measured). Older figures are deleted rather than archived. The dated competitor inspection in docs/comparison.md is removed with them.

Measurements (macOS arm64, cold cache, 30 repetitions, outputs byte-identical on every case)

Case 0.5.1 This branch
Compress a 32 MiB stream 177 ms 62 ms
Compress a 3 MB log 27 ms 15 ms
Compress a 12,000-row JSONL table 29 ms 20 ms
Paged recovery from a stored original 11 ms 7 ms
Compress a 136 KB log 5 ms 5 ms
Small output (process startup) 3.6 ms 3.6 ms

Byte equality was also checked on a 19-input differential corpus (ANSI colours, carriage-return progress, vertical tab and form feed, no-break and ideographic spaces, hex identifiers, stack frames, JSON, JSONL, malformed lines).

Considered and not taken

  • mimalloc: 13% on the allocation-heavy aggregate query, neutral elsewhere; not worth a second C dependency.
  • sonic-rs / simd-json for the artifact: artifact identities are the SHA-256 of the exact serde_json serialization, so a different serializer would have to be byte-identical; not attempted.
  • Skipping the artifact copy of the input would change the recovery contract; not attempted.

Verification

cargo fmt --check, cargo clippy --all-targets -- -D warnings, cargo test (new template equivalence test), cargo run -- bench, scripts/check_skill.py, bench unit tests (17), npm test (10/10 with dependencies installed), scripts/check_readme.py, scripts/check_adapters.py, bench/publish.py end to end.

Profiling a 32 MiB capture put 55 ms of 175 in SHA-256: the sha2 crate only
uses the ARMv8 and SHA-NI instructions with its asm feature, which brings the
same digest to 10 ms. Every capture hashes its input twice (blob and
artifact) and every read verifies, so this is the largest single cost.

The compact-v3 line pass ran two vocabulary regexes on every line and built a
template for every non-diagnostic line. Identical text is now classified once
and repetitions copy the result by index; the template builder scans bytes and
only decodes non-ASCII characters, with a test proving it equal to the
character-by-character version on Unicode whitespace, control bytes, digits
and hex runs. Line texts are hashed with foldhash instead of SipHash.

The release profile uses fat LTO, one codegen unit and abort on panic: a
third smaller binary and a few percent per call.

Outputs are byte-identical to 0.5.1 on the content fixtures, the performance
fixtures and a 19-input differential corpus (ANSI, CR progress, vertical tab,
no-break space, hex ids, stack frames, JSON, JSONL). Cold medians on macOS
arm64: 32 MiB stream 177 -> 62 ms, 3 MB log 27 -> 15 ms, 12,000-row table
29 -> 20 ms, paged recovery 11 -> 7 ms.
bench/publish.py measures one binary with both probes, replaces everything
under bench/results/ with content.json and performance.json, and rewrites the
README speed table between its markers with the version it measured. Older
figures are deleted, not archived: a benchmark describes the version that
ships. The dated competitor inspection in docs/comparison.md goes with them.
@maxgfr
maxgfr merged commit 27d5a6f into main Sep 11, 2026
6 checks passed
@maxgfr
maxgfr deleted the perf-engine branch September 11, 2026 12:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant