Skip to content

perf: reduce Vamana memory through shared vector storage - #97

Open
zhenghaoz wants to merge 1 commit into
mainfrom
perf/vamana-memory
Open

zhenghaoz wants to merge 1 commit into
mainfrom
perf/vamana-memory

Conversation

@zhenghaoz

Copy link
Copy Markdown
Contributor

Vamana build, save, and reopen retained multiple copies of the same vectors and eagerly built fallback indexes. This change shares immutable vector storage and scalar codes, streams persistence, and reuses traversal scratch to lower process peak memory while preserving the file format and query/refinement behavior.

  • Borrow immutable collection FP32 rows during construction; detach on mutation and transfer newly built graphs into quantization without cloning.
  • Reuse encoded document originals when reopening scalar-quantized Vamana, validate them against the artifact, and omit unused edge-distance caches. Use temporary mmap for artifact decoding.
  • Defer exact fallback construction, share quantized Flat views, and reuse construction heaps, query visit marks, and validation scratch.
  • Stream Save through a 64 KiB buffer with incremental CRC32C and atomic replacement. Read metadata instead of cloning documents just to check whether Optimize is needed.
  • Update the four xvec CSV rows and preserve raw reports, commands, checksums, prior measurements, and the analysis.

The persisted bytes remain compatible. Borrowed inputs must remain immutable; mutation uses copy-on-write. Save holds the read lock through persistence, so concurrent Add publication waits until Save finishes. No forced GC or GC tuning was added.

The complete benchmark uses Cohere Performance768D100K (100,000 × 768), K=100, 1,000 serial queries, and 8 concurrent workers for 30 seconds. Environment: e2-standard-8 / AMD EPYC 7B12, Go 1.27.1, CGO_ENABLED=0, GOMAXPROCS=8, GOMEMLIMIT=24GiB, affinity 0–7. Vamana degree/build list/query list = 64/100/200, alpha=1.2, occlusion=750; INT4/INT8 rotation enabled, refinement disabled, mmap enabled. Each run uses a fresh collection. RSS is the entire process high-water mark measured by wait4, including loading, optimization, reopen, and queries; downloads and compilation are excluded.

“Original” below is the historical xvec CSV at 6a8b120; “This PR” is the complete rerun on 5d9b8f5 plus the archived production patch 586bbd53868e. zvec is the unchanged historical v0.7.0+rotate measurement, not a new run.

Precision Measurement Peak RSS (MiB) Optimize (s) Serial QPS Concurrent QPS Recall@100 (%)
INT4 xvec original 3092.26 49.68 330.15 2003.05 87.012
INT4 xvec this PR 1575.29 81.64 212.11 1410.06 87.013
INT4 zvec historical 566.63 31.25 388.83 2661.17 81.039
INT8 xvec original 3145.95 55.29 276.88 1913.68 98.730
INT8 xvec this PR 1508.05 75.84 201.40 1322.81 98.726
INT8 zvec historical 644.31 28.25 551.29 3142.45 98.355
FP16 xvec original 3739.13 94.38 152.71 1279.50 99.363
FP16 xvec this PR 1486.04 77.36 137.84 882.08 99.349
FP16 zvec historical 783.31 77.13 382.94 2270.11 99.198
FP32 xvec original 2797.18 50.42 317.50 1607.42 99.375
FP32 xvec this PR 1956.61 75.02 204.63 978.77 99.374
FP32 zvec historical 770.00 57.12 357.31 1860.60 99.392

Compared with the original CSV, peak RSS falls by 49.1% / 52.1% / 60.3% / 30.1% for INT4 / INT8 / FP16 / FP32. Recall differs by at most 0.014 percentage points. Memory remains 1.9–2.8× the historical zvec values.

The measurements also show lower serial and concurrent QPS than the original xvec CSV in all four modes. Optimize is slower for INT4/INT8/FP32 and faster for FP16. These are single-run historical comparisons, and the original xvec revision predates other changes including SIMD kernels; they do not isolate this patch's effect. zvec has different recall, particularly for INT4, so these are not equal-recall comparisons. The CSV retains the full-run results, including regressions; separate diagnostics do not replace them.

Updated CSV, original CSV, analysis, and latest raw reports and provenance.

Validation completed before submission:

  • go test ./... and the additional Vamana snapshot/update isolation test passed.
  • Regression coverage includes borrowed/owned artifact byte equality, more than 1,000 graph nodes, L2/IP/cosine, all scalar precisions, mmap on/off, filtering, refinement, original-vector isolation, resave, and mutation.
  • All four full 100K benchmark runs exited successfully; CSV shape, unchanged zvec rows, and archived source checksums were verified.
  • pnpm test (21 tests), pnpm check, and pnpm build passed. The build retains the existing large-chunk warning.
  • The race detector was not run because this environment has no C compiler.

@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying xvec with  Cloudflare Pages  Cloudflare Pages

Latest commit: bc87d5a
Status: ✅  Deploy successful!
Preview URL: https://abae4c92.xvec.pages.dev
Branch Preview URL: https://perf-vamana-memory.xvec.pages.dev

View logs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant