Skip to content

perf: reduce Flat and HNSW vector memory - #91

Merged
zhenghaoz merged 7 commits into
mainfrom
perf/flat-memory
Sep 26, 2026
Merged

zhenghaoz merged 7 commits into
mainfrom
perf/flat-memory

Conversation

@zhenghaoz

@zhenghaoz zhenghaoz commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Flat and HNSW retained decoded FP32 document vectors alongside additional index-owned originals, and quantized Flat allocated every candidate position/score per query. On 100K × 768D vectors, this pushed quantized Flat peak RSS to about 3 GiB and FP32 HNSW to about 2.3 GiB.

  • Share immutable decoded originals with quantized indexes, defer exact fallback construction, and retain only top-k while scanning Flat codes.
  • Reuse rotation work buffers during scalar-code construction.
  • For immutable segments opened read-only, retain validated encoded FP32 document bytes for quantized Flat and all HNSW precisions. Quantized indexes decode originals on demand; FP32 HNSW retains one owned contiguous scoring array, preserving search locality while eliminating the duplicate decoded document array.
  • Verify HNSW artifacts against supplied originals. Temporary mmap loading avoids a whole-artifact heap buffer, and reconstructed codes share one source of originals with the graph and Flat view.

This follows zvec's stored HNSW vector access, mmap storage selection, reusable rotation buffer, and bounded result heaps.

Encoded document storage is limited to immutable segments in read-only handles. Writable handles and mutable WAL segments retain decoded documents. Query leases keep mappings alive until queries finish; projected vectors are independent copies. Existing owning core constructors retain copy semantics. No unsafe casts or file-format changes are introduced. Grouped exact refinement may still materialize an exact index on demand. FP32 Flat storage remains unchanged.

Validation

  • go test ./...
  • CGO_ENABLED=1 go test -race ./ ./internal/core/algorithm -run 'Test.*(HNSW|ReadOnly|Encoded|Borrow|Immutable)' -count=1; prior Flat/rotation race coverage also passed.
  • golangci-lint run (v2.13.1): 0 issues.
  • Tests compare persisted HNSW bytes, graph topology, scores and query results across owned/encoded storage; cover FP16/INT8/INT4 with L2/IP/cosine, FP32 contiguous storage, prefetch, refinement, filters, query by primary key, nullable vectors, mmap on/off, result ownership, snapshots and close during active queries. Rotation destination reuse is allocation-free.
  • All 18 docs tests, Astro check/build, and Chromium checks of charts and SVG exports passed.

HNSW full-lifetime benchmark

Same e2-standard-8 host, Go 1.27.1, CGO disabled, Cohere Performance768D100K, 100K × 768D, cosine, M=50, construction EF=500, search EF=300, K=100, 8 construction/query workers, 30-second concurrency phase, 1,000 serial queries, GOMAXPROCS=8, GOMEMLIMIT=24GiB, CPU affinity 0–7, mmap enabled, no refiner, IDs-only payloads. INT4/INT8 retain rotation. Runs alternate previous/current revisions and include insertion, Optimize, close/reopen and queries. Peak RSS is measured over the entire process with wait4.

Previous implementation: 4c7db6e (binary built at identical implementation 88bdaf9); final HNSW implementation: a2766b5. The CSV uses the final full runs.

Precision Peak RSS MiB (before → after) Reduction Concurrent QPS (before → after) Serial QPS (before → after) Recall@100 % (before → after)
FP32 2380.1 → 2200.3 7.6% 1054.81 → 1096.11 211.62 → 213.97 99.713 → 99.714
FP16 1707.2 → 1804.0 -5.7% 1381.51 → 876.35 232.88 → 127.94 99.681 → 99.681
INT8 1557.6 → 1634.4 -4.9% 2008.80 → 1875.31 318.56 → 290.17 98.956 → 98.960
INT4 1542.1 → 1220.5 20.9% 1903.82 → 1973.12 303.57 → 313.69 87.986 → 87.984

Parallel HNSW construction can produce different graphs across runs. To isolate loading/search behavior, the final binary also reopened each unchanged baseline graph, used the same 1,000 benchmark queries, and matched baseline recall exactly for every precision. Fresh-process heap diagnostics used the same baseline artifacts, one warmup query, and GC while retaining the collection:

Precision Live Go heap MiB (before → after) Same-graph concurrent QPS (before → after) Same-graph serial QPS (before → after) Recall@100 %
FP32 677.6 → 384.6 1037.77 → 1089.78 214.35 → 212.66 99.713
FP16 541.8 → 248.8 1325.42 → 1290.13 233.90 → 216.22 99.681
INT8 468.5 → 175.5 1864.36 → 1851.87 280.65 → 276.18 98.956
INT4 432.2 → 139.2 1901.04 → 1384.38 200.05 → 209.76 87.986

The full runs reduced peak RSS by 7.6% for FP32 and 20.9% for INT4, but FP16 and INT8 peak RSS rose by 5.7% and 4.9%. FP16 concurrent throughput also fell 36.6% in the full run; its same-graph query-only comparison was 1325 → 1290 QPS (−2.7%). Memory savings are established for read-only query state; whole-lifetime peaks and throughput depend on precision and run.

INT4's first same-graph comparison also decreased from 1901 to 1384 QPS. A reverse-order repeat on the same artifact measured final/previous at 2038/2028 concurrent QPS and 313.00/311.56 serial QPS; recall remained identical. The large decrease did not reproduce consistently. Both rounds are retained here, and the published CSV continues to use the full-lifetime run.

Live heap and whole-lifetime peak RSS are distinct measurements. The query path removes approximately 293 MiB of duplicate decoded originals per 100K × 768D field; insertion/building and mapped pages can still dominate process peak RSS. Explicit GC is only diagnostic, not a production or benchmark change. Single-run QPS is subject to host/process variation.

Flat full-lifetime benchmark

Same host and query/runtime settings, Flat index. Original baseline 8f85972 (same implementation as PR base); final Flat measurements from 88bdaf9. All recall values match exactly.

Precision Peak RSS MiB (before → after) Concurrent QPS (before → after) Recall@100 %
FP32 1193.2 → 1213.8 133.87 → 167.72 99.999
FP16 2942.5 → 1069.7 212.90 → 219.89 99.957
INT8 3020.1 → 1069.8 317.62 → 317.34 99.119
INT4 3063.0 → 1069.2 352.74 → 509.53 87.353

Quantized Flat peak RSS falls by 64–65%. Separately, read-only live heap falls from 508.3 to 215.3 MiB for FP16 and 398.4 to 105.5 MiB for INT4 relative to the previous PR implementation. The Flat INT8 full run showed lower throughput than the previous PR implementation (381 → 317 concurrent QPS); same-collection query-only runs in final/previous/previous/final order measured 411/389/396/428 QPS. The CSV retains the full-run value; no universal throughput improvement is claimed.

Both docs/benchmark-hnsw.csv and docs/benchmark-flat.csv contain measured results with the tested source revisions. Existing zvec measurements are retained.

Reproduction

Build the desired revision with CGO_ENABLED=0 go build -o /tmp/vector-db-bench ./cmd/vector-db-bench. Use a fresh collection for each precision; change quantization to int4, int8, or fp16, and omit it for FP32. For Flat, use --index-type flat and omit HNSW parameters.

/usr/bin/time -v env GOMAXPROCS=8 GOMEMLIMIT=24GiB \
  taskset -c 0-7 /tmp/vector-db-bench xvec \
  --path /tmp/hnsw-bench-collection \
  --case-type Performance768D100K \
  --dataset-dir /tmp/fp32-bench/dataset --skip-download \
  --index-type hnsw --m 50 --ef-construction 500 --ef-search 300 \
  --quantize-type fp16 --k 100 --batch-size 100 \
  --max-docs-per-segment 10000000 --optimize-concurrency 8 \
  --num-concurrency 8 --concurrency-duration 30s --serial-cooldown 3s \
  --payload-profile ids_only --enable-mmap=true --is-using-refiner=false \
  --output /tmp/hnsw-bench.json

For same-graph queries, reuse the baseline collection path and add --skip-drop-old --skip-load with the final binary.

@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 81.19403% with 63 lines in your changes missing coverage. Please review.
✅ Project coverage is 81.66%. Comparing base (b1cc63b) to head (116c783).

Files with missing lines Patch % Lines
collection.go 73.91% 24 Missing ⚠️
internal/core/algorithm/rotator.go 60.71% 11 Missing ⚠️
encoded_vector.go 82.60% 8 Missing ⚠️
internal/core/algorithm/hnsw_encoded_vectors.go 69.56% 7 Missing ⚠️
internal/core/algorithm/hnsw_algorithm.go 90.00% 6 Missing ⚠️
internal/core/algorithm/quantized_flat_searcher.go 90.76% 6 Missing ⚠️
document.go 92.30% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main      #91      +/-   ##
==========================================
+ Coverage   81.37%   81.66%   +0.29%     
==========================================
  Files         148      150       +2     
  Lines       30253    30534     +281     
==========================================
+ Hits        24618    24937     +319     
+ Misses       5633     5595      -38     
  Partials        2        2              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Deploying xvec with  Cloudflare Pages  Cloudflare Pages

Latest commit: 116c783
Status: ✅  Deploy successful!
Preview URL: https://7e10a460.xvec.pages.dev
Branch Preview URL: https://perf-flat-memory.xvec.pages.dev

View logs

@zhenghaoz zhenghaoz changed the title perf: reduce quantized Flat memory usage perf: reduce Flat and HNSW vector memory Sep 26, 2026
@zhenghaoz
zhenghaoz merged commit 6a8b120 into main Sep 26, 2026
9 checks passed
@zhenghaoz
zhenghaoz deleted the perf/flat-memory branch September 26, 2026 05:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant