perf: hardware SHA-256, single classification per distinct line, fat LTO - #8
Merged
Merged
Conversation
Profiling a 32 MiB capture put 55 ms of 175 in SHA-256: the sha2 crate only uses the ARMv8 and SHA-NI instructions with its asm feature, which brings the same digest to 10 ms. Every capture hashes its input twice (blob and artifact) and every read verifies, so this is the largest single cost. The compact-v3 line pass ran two vocabulary regexes on every line and built a template for every non-diagnostic line. Identical text is now classified once and repetitions copy the result by index; the template builder scans bytes and only decodes non-ASCII characters, with a test proving it equal to the character-by-character version on Unicode whitespace, control bytes, digits and hex runs. Line texts are hashed with foldhash instead of SipHash. The release profile uses fat LTO, one codegen unit and abort on panic: a third smaller binary and a few percent per call. Outputs are byte-identical to 0.5.1 on the content fixtures, the performance fixtures and a 19-input differential corpus (ANSI, CR progress, vertical tab, no-break space, hex ids, stack frames, JSON, JSONL). Cold medians on macOS arm64: 32 MiB stream 177 -> 62 ms, 3 MB log 27 -> 15 ms, 12,000-row table 29 -> 20 ms, paged recovery 11 -> 7 ms.
bench/publish.py measures one binary with both probes, replaces everything under bench/results/ with content.json and performance.json, and rewrites the README speed table between its markers with the version it measured. Older figures are deleted, not archived: a benchmark describes the version that ships. The dated competitor inspection in docs/comparison.md goes with them.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Profiling the 32 MiB capture on 0.5.1 gave: parse 5 ms, view 32 ms, blob 63 ms, artifact 70 ms. Three changes, all output-identical:
sha2only uses ARMv8 / SHA-NI instructions with itsasmfeature. The 32 MiB digest drops from 55 ms to 10 ms; every capture hashes twice and every read verifies, so this is the largest single cost.Also:
bench/publish.pymeasures one binary and publishes the only figures the repository keeps (bench/results/content.json,bench/results/performance.json, and the README speed table between its markers, stamped with the version measured). Older figures are deleted rather than archived. The dated competitor inspection indocs/comparison.mdis removed with them.Measurements (macOS arm64, cold cache, 30 repetitions, outputs byte-identical on every case)
Byte equality was also checked on a 19-input differential corpus (ANSI colours, carriage-return progress, vertical tab and form feed, no-break and ideographic spaces, hex identifiers, stack frames, JSON, JSONL, malformed lines).
Considered and not taken
Verification
cargo fmt --check,cargo clippy --all-targets -- -D warnings,cargo test(new template equivalence test),cargo run -- bench,scripts/check_skill.py, bench unit tests (17),npm test(10/10 with dependencies installed),scripts/check_readme.py,scripts/check_adapters.py,bench/publish.pyend to end.