Conversation
Deploying xvec with
|
| Latest commit: |
b65371a
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://0db0a58a.xvec.pages.dev |
| Branch Preview URL: | https://perf-hnsw-filter-benchmarks.xvec.pages.dev |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #98 +/- ##
==========================================
+ Coverage 81.25% 81.37% +0.11%
==========================================
Files 149 153 +4
Lines 30743 31000 +257
==========================================
+ Hits 24980 25226 +246
- Misses 5761 5772 +11
Partials 2 2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Integer filters repeatedly cloned postings and enumerated matching rows before vector search, while scalar-code HNSW allocated comparator-based heaps for every traversal. This PR aggregates integer ranges in 256-term blocks, freezes posting aggregates, defers ordinal enumeration, reuses snapshot visibility and selected-key scans, caches encoded FP16 cosine norms on compatible kernels, and now reuses specialized dual heaps in upper-layer and base-layer scalar-code searches.
The traversal keeps separate best-first candidate and worst-first accepted-result heaps. Comparisons can inline; heap updates move parents/children rather than swapping at every step. Accepted results drain into an independently owned sorted slice, so later queries cannot overwrite them. Candidate ties still use graph positions, base-result ties still use document keys, and rejected filter nodes remain available as traversal bridges. Admission, stopping, radius and scoring rules are preserved. Pooled heap buffers above 4,096 nodes are discarded, and index-key references are cleared on release.
No persisted format or graph changes. Encoded FP16 norms remain derived at build/open, with four bytes per vector. AVX-512 retains its original uncached cosine scoring. All targets use the generic prefetch helper; AMD64-specific prefetch files stay removed.
.github/workflows/ci.ymlmatches the PR base. The integer/label Markdown benchmark reports and secondary integer CSVs stay deleted.Controlled dual-heap comparison, 2026-09-30
Clean xvec source
a0d0d15c9e1c92184162ea12d24afe194467b7e3→43ab0e93fd05114c613a2dcc20eee313b62cfc22, using the same original persisted collections. Both versions were freshly measured three times at 20% and 50% matching for both integer and label filtering: 24 successful timed runs, with alternating revision/rate order and no competing local test/build workloads during timing.QPS and P99 are separate three-run medians. These end-to-end FP16 measurements do not demonstrate a consistent QPS improvement. Every run preserves its preceding aggregate Recall@100. Query binaries, datasets and xvec graph/vector/posting artifacts are hashed before and after; no collection was rebuilt or optimized.
Cohere 100K, 768 dimensions, cosine, FP16 HNSW, M=50, EFConstruction=500, EFSearch=300, K=100, no rotation/refinement, mmap, ID-only results. Each fresh process runs 100 warmup queries, eight workers for 30 seconds, a three-second cooldown and all 1,000 serial recall queries. Same e2-standard-8 / AMD EPYC 7B12 host, affinity 0–7, Go 1.27.1, CGO_ENABLED=0, GOMAXPROCS=8, GOMEMLIMIT=24GiB. No profiler or filesystem-cache flush. Integer truth is exhaustive float64 cosine over the original FP32 vectors; label truth is the published label-filtered neighbors.
The unchanged
BenchmarkQuantizedHNSWInt8Batch/batch=falseadditionally compares three one-second runs per revision after the end-to-end tests. On its 2,000-vector, 128-dimensional INT8 index with an always-true filter, allocations fall 19 → 6 per query, allocated bytes 11,448 → 5,108, and median time 39.520 → 33.166 µs. These are separate microbenchmark results.Published results
Each CSV keeps 26 individual observations across all nine matching rates. Its six high-match xvec rows now use
43ab0e93fd05114c613a2dcc20eee313b62cfc22; the other 20 rows keep the preceding paired rerun's source and timestamp. zvec was not rerun for this heap comparison; its retained rows use unchanged official native v0.7.0. The CSVs contain measurements from different runs/revisions, not a new simultaneous xvec/zvec comparison. Both default planners select candidate linear search up to 10% matching and filtered HNSW at 20% and 50%.Validation
CGO_ENABLED=0 go test ./...passes, plus targetednoasmand CGO-enabled race checks for heap ordering/reuse, filtered scalar-code searches and FP16 cache equivalence.TestScalarHNSWBuildRecallAndPersistence. This remains unresolved; the heap change does not modify scoring, test tolerances or CI.docs/benchmark-runsare included in the PR. Raw reports, binaries, profiles, scripts and previous CSV archives stay local and ignored.