Skip to content

Vector: rerank candidate count differs per path — docs oversample×ef_search vs top_k×oversample vs hardcoded ×3 #397

Description

@EnRaiha

Version / build tested against

origin/main @ e235fe5 (2026-09-29)

Deployment mode

Origin — single node (local)

Engine(s) involved

Vector

Summary

Rerank candidate count is computed three different ways: docs state oversample * ef_search; the executor uses top_k * oversample; quantized_search hardcodes (top_k * 3).max(20). A caller-supplied oversample is silently dropped on the quantized path, so tuning is unreliable.

Steps to reproduce

Run the same search through the standard and quantized paths with oversample = k and compare candidate counts; inspect rerank_k computations. (static verification at the pin; runtime repro pending)

Expected behavior

One formula, honouring the configured oversample (and ef_search where documented).

Actual behavior

Three formulas; configured oversample ignored on the quantized path.

What actually happened? (severity facts)

  • Acknowledged/committed data was lost, corrupted, or silently wrong
  • The server crashed, hung, or failed to start
  • A security or isolation boundary was crossed
  • Core functionality is broken with no acceptable workaround
  • A workaround exists (accept unpredictable recall/latency tuning)

Proposed severity

SEV-3 — Medium: tuning contract broken; results remain otherwise correct.

Reproducibility

Always — every attempt

Last known-good version / commit (if a regression)

(unknown / not a regression)

Environment & logs

Linux x86_64; Verified by static code reading at the pin above; runtime reproduction pending.
Code references:

  • (see prior-art line below)

Before submitting


Additional evidence (origin/main @ e235fe55c)

  • What: docs/vectors.md:78 and planner/query_options.rs:65 document the rerank set as oversample × ef_search. The executor computes fetch_k = top_k × oversample (vector_search_exec.rs:212-216, × 2 × oversample.max(20) under RLS). The sealed-segment rerank ignores oversample and computes rerank_k = top_k × 3 .max(20) (collection/search.rs:54).
  • Where: nodedb/src/data/executor/handlers/vector_search_exec.rs:212-216,289,295; nodedb-vector/src/collection/search.rs:54-55; docs/vectors.md:78; nodedb-vector/src/planner/query_options.rs:65.
  • Evidence: quotes above.
  • Impact: a caller-supplied oversample is silently dropped on the quantized path; documented tuning behavior does not match execution, so recall/memory tuning is unreliable. The two formulas also diverge as ef_search and top_k differ.
  • Fix (claim direction): route all pool sizing through one function that honors the resolved oversample and ef_search; update quantized_search to accept the resolved breadth.
  • Prior-art: PR [#270](fix(sql,vector): reject negative LIMIT, bound ANALYZE name, unify oversample #270) ("unify oversample") was closed unmerged; the inconsistency stands at e235fe5. No open issue covers it.
    Why: the tuning contract is broken: docs say oversample * ef_search; the executor uses top_k * oversample; quantized_search hardcodes (top_k * 3).max(20). A caller-supplied oversample is silently dropped on the quantized path.
    Steps to test: run a search with oversample = k on the standard vs quantized paths and compare candidate counts; rg -n "rerank_k|oversample" nodedb-vector/src/collection/search.rs nodedb/src/data/executor/handlers/vector_search_exec.rs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions