Skip to content

perf: optimize filtered HNSW search and benchmark Cohere 100K - #98

Open
zhenghaoz wants to merge 19 commits into
mainfrom
perf/hnsw-filter-benchmarks
Open

zhenghaoz wants to merge 19 commits into
mainfrom
perf/hnsw-filter-benchmarks

Conversation

@zhenghaoz

@zhenghaoz zhenghaoz commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Integer filters repeatedly cloned postings and enumerated matching rows before vector search, while scalar-code HNSW allocated comparator-based heaps for every traversal. This PR aggregates integer ranges in 256-term blocks, freezes posting aggregates, defers ordinal enumeration, reuses snapshot visibility and selected-key scans, caches encoded FP16 cosine norms on compatible kernels, and now reuses specialized dual heaps in upper-layer and base-layer scalar-code searches.

The traversal keeps separate best-first candidate and worst-first accepted-result heaps. Comparisons can inline; heap updates move parents/children rather than swapping at every step. Accepted results drain into an independently owned sorted slice, so later queries cannot overwrite them. Candidate ties still use graph positions, base-result ties still use document keys, and rejected filter nodes remain available as traversal bridges. Admission, stopping, radius and scoring rules are preserved. Pooled heap buffers above 4,096 nodes are discarded, and index-key references are cleared on release.

No persisted format or graph changes. Encoded FP16 norms remain derived at build/open, with four bytes per vector. AVX-512 retains its original uncached cosine scoring. All targets use the generic prefetch helper; AMD64-specific prefetch files stay removed. .github/workflows/ci.yml matches the PR base. The integer/label Markdown benchmark reports and secondary integer CSVs stay deleted.

Controlled dual-heap comparison, 2026-09-30

Clean xvec source a0d0d15c9e1c92184162ea12d24afe194467b7e3 → 43ab0e93fd05114c613a2dcc20eee313b62cfc22, using the same original persisted collections. Both versions were freshly measured three times at 20% and 50% matching for both integer and label filtering: 24 successful timed runs, with alternating revision/rate order and no competing local test/build workloads during timing.

Filter Matching Before QPS After QPS Change Before P99 (ms) After P99 (ms) Recall@100 (%)
Int 20% 436.79 427.98 -2.02% 33.481 28.552 99.884
Int 50% 810.12 778.30 -3.93% 18.173 17.780 99.806
Label 20% 480.84 474.80 -1.26% 23.681 23.609 99.901
Label 50% 915.45 922.27 +0.74% 12.959 12.964 99.833

QPS and P99 are separate three-run medians. These end-to-end FP16 measurements do not demonstrate a consistent QPS improvement. Every run preserves its preceding aggregate Recall@100. Query binaries, datasets and xvec graph/vector/posting artifacts are hashed before and after; no collection was rebuilt or optimized.

Cohere 100K, 768 dimensions, cosine, FP16 HNSW, M=50, EFConstruction=500, EFSearch=300, K=100, no rotation/refinement, mmap, ID-only results. Each fresh process runs 100 warmup queries, eight workers for 30 seconds, a three-second cooldown and all 1,000 serial recall queries. Same e2-standard-8 / AMD EPYC 7B12 host, affinity 0–7, Go 1.27.1, CGO_ENABLED=0, GOMAXPROCS=8, GOMEMLIMIT=24GiB. No profiler or filesystem-cache flush. Integer truth is exhaustive float64 cosine over the original FP32 vectors; label truth is the published label-filtered neighbors.

The unchanged BenchmarkQuantizedHNSWInt8Batch/batch=false additionally compares three one-second runs per revision after the end-to-end tests. On its 2,000-vector, 128-dimensional INT8 index with an always-true filter, allocations fall 19 → 6 per query, allocated bytes 11,448 → 5,108, and median time 39.520 → 33.166 µs. These are separate microbenchmark results.

Published results

  • Integer-filter CSV is the only retained integer-filter CSV.
  • Label-filter CSV preserves the original collection-build measurements with their own build source and timestamp.

Each CSV keeps 26 individual observations across all nine matching rates. Its six high-match xvec rows now use 43ab0e93fd05114c613a2dcc20eee313b62cfc22; the other 20 rows keep the preceding paired rerun's source and timestamp. zvec was not rerun for this heap comparison; its retained rows use unchanged official native v0.7.0. The CSVs contain measurements from different runs/revisions, not a new simultaneous xvec/zvec comparison. Both default planners select candidate linear search up to 10% matching and filtered HNSW at 20% and 50%.

Validation

  • Local CGO_ENABLED=0 go test ./... passes, plus targeted noasm and CGO-enabled race checks for heap ordering/reuse, filtered scalar-code searches and FP16 cache equivalence.
  • Randomized heap operations match the original generic-heap comparators for all four metrics, frontier/upper-layer position ties and base-result key ties. Tests cover independent result ownership, cleared key references and bounded pooled buffers. Existing filtered/batch tests cover rejected bridges, duplicate neighbors, radius searches, zero vectors, score ties and scalar/batch result equivalence.
  • All 24 timed runs succeed with identical aggregate recall and unchanged persisted artifacts. Both CSVs are checked against raw reports; six rows per CSV are refreshed, other observations and original build measurements stay unchanged. Whitespace checks pass.
  • CI for the heap source passes lint, SIMD, macOS, Linux ARM and Windows ARM. Linux x64 and Windows x64 fail the same FP16 exact-score assertion previously seen on the before revision: key 516, expected 0 versus approximately 1.19e-7 in TestScalarHNSWBuildRecallAndPersistence. This remains unresolved; the heap change does not modify scoring, test tolerances or CI.
  • No files under docs/benchmark-runs are included in the PR. Raw reports, binaries, profiles, scripts and previous CSV archives stay local and ignored.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Deploying xvec with  Cloudflare Pages  Cloudflare Pages

Latest commit: b65371a
Status: ✅  Deploy successful!
Preview URL: https://0db0a58a.xvec.pages.dev
Branch Preview URL: https://perf-hnsw-filter-benchmarks.xvec.pages.dev

View logs

@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.59420% with 29 lines in your changes missing coverage. Please review.
✅ Project coverage is 81.37%. Comparing base (5d9b8f5) to head (b65371a).

Files with missing lines Patch % Lines
collection_filter.go 87.50% 11 Missing ⚠️
collection.go 85.29% 10 Missing ⚠️
internal/core/algorithm/quantized_flat_searcher.go 76.19% 5 Missing ⚠️
internal/ailego/math_batch/distance_fp16_amd64.go 0.00% 1 Missing ⚠️
internal/core/algorithm/hnsw_prefetch_generic.go 85.71% 1 Missing ⚠️
internal/core/algorithm/hnsw_quantized_searcher.go 97.67% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main      #98      +/-   ##
==========================================
+ Coverage   81.25%   81.37%   +0.11%     
==========================================
  Files         149      153       +4     
  Lines       30743    31000     +257     
==========================================
+ Hits        24980    25226     +246     
- Misses       5761     5772      +11     
  Partials        2        2              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant