perf: reduce Flat and HNSW vector memory - #91
Merged
Merged
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #91 +/- ##
==========================================
+ Coverage 81.37% 81.66% +0.29%
==========================================
Files 148 150 +2
Lines 30253 30534 +281
==========================================
+ Hits 24618 24937 +319
+ Misses 5633 5595 -38
Partials 2 2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Deploying xvec with
|
| Latest commit: |
116c783
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://7e10a460.xvec.pages.dev |
| Branch Preview URL: | https://perf-flat-memory.xvec.pages.dev |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Flat and HNSW retained decoded FP32 document vectors alongside additional index-owned originals, and quantized Flat allocated every candidate position/score per query. On 100K × 768D vectors, this pushed quantized Flat peak RSS to about 3 GiB and FP32 HNSW to about 2.3 GiB.
This follows zvec's stored HNSW vector access, mmap storage selection, reusable rotation buffer, and bounded result heaps.
Encoded document storage is limited to immutable segments in read-only handles. Writable handles and mutable WAL segments retain decoded documents. Query leases keep mappings alive until queries finish; projected vectors are independent copies. Existing owning core constructors retain copy semantics. No unsafe casts or file-format changes are introduced. Grouped exact refinement may still materialize an exact index on demand. FP32 Flat storage remains unchanged.
Validation
go test ./...CGO_ENABLED=1 go test -race ./ ./internal/core/algorithm -run 'Test.*(HNSW|ReadOnly|Encoded|Borrow|Immutable)' -count=1; prior Flat/rotation race coverage also passed.golangci-lint run(v2.13.1): 0 issues.HNSW full-lifetime benchmark
Same e2-standard-8 host, Go 1.27.1, CGO disabled, Cohere
Performance768D100K, 100K × 768D, cosine, M=50, construction EF=500, search EF=300, K=100, 8 construction/query workers, 30-second concurrency phase, 1,000 serial queries, GOMAXPROCS=8, GOMEMLIMIT=24GiB, CPU affinity 0–7, mmap enabled, no refiner, IDs-only payloads. INT4/INT8 retain rotation. Runs alternate previous/current revisions and include insertion, Optimize, close/reopen and queries. Peak RSS is measured over the entire process withwait4.Previous implementation:
4c7db6e(binary built at identical implementation88bdaf9); final HNSW implementation:a2766b5. The CSV uses the final full runs.Parallel HNSW construction can produce different graphs across runs. To isolate loading/search behavior, the final binary also reopened each unchanged baseline graph, used the same 1,000 benchmark queries, and matched baseline recall exactly for every precision. Fresh-process heap diagnostics used the same baseline artifacts, one warmup query, and GC while retaining the collection:
The full runs reduced peak RSS by 7.6% for FP32 and 20.9% for INT4, but FP16 and INT8 peak RSS rose by 5.7% and 4.9%. FP16 concurrent throughput also fell 36.6% in the full run; its same-graph query-only comparison was 1325 → 1290 QPS (−2.7%). Memory savings are established for read-only query state; whole-lifetime peaks and throughput depend on precision and run.
INT4's first same-graph comparison also decreased from 1901 to 1384 QPS. A reverse-order repeat on the same artifact measured final/previous at 2038/2028 concurrent QPS and 313.00/311.56 serial QPS; recall remained identical. The large decrease did not reproduce consistently. Both rounds are retained here, and the published CSV continues to use the full-lifetime run.
Live heap and whole-lifetime peak RSS are distinct measurements. The query path removes approximately 293 MiB of duplicate decoded originals per 100K × 768D field; insertion/building and mapped pages can still dominate process peak RSS. Explicit GC is only diagnostic, not a production or benchmark change. Single-run QPS is subject to host/process variation.
Flat full-lifetime benchmark
Same host and query/runtime settings, Flat index. Original baseline
8f85972(same implementation as PR base); final Flat measurements from88bdaf9. All recall values match exactly.Quantized Flat peak RSS falls by 64–65%. Separately, read-only live heap falls from 508.3 to 215.3 MiB for FP16 and 398.4 to 105.5 MiB for INT4 relative to the previous PR implementation. The Flat INT8 full run showed lower throughput than the previous PR implementation (381 → 317 concurrent QPS); same-collection query-only runs in final/previous/previous/final order measured 411/389/396/428 QPS. The CSV retains the full-run value; no universal throughput improvement is claimed.
Both
docs/benchmark-hnsw.csvanddocs/benchmark-flat.csvcontain measured results with the tested source revisions. Existing zvec measurements are retained.Reproduction
Build the desired revision with
CGO_ENABLED=0 go build -o /tmp/vector-db-bench ./cmd/vector-db-bench. Use a fresh collection for each precision; change quantization toint4,int8, orfp16, and omit it for FP32. For Flat, use--index-type flatand omit HNSW parameters.For same-graph queries, reuse the baseline collection path and add
--skip-drop-old --skip-loadwith the final binary.