Skip to content

Answer natural graph questions with ranked candidate selection - #342

Merged
forhappy merged 3 commits into
mainfrom
codex/natural-query-answers
Sep 30, 2026
Merged

forhappy merged 3 commits into
mainfrom
codex/natural-query-answers

Conversation

@forhappy

@forhappy forhappy commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Natural-language graph questions now continue with a disclosed, deterministically ranked symbol and execute the matching typed query. For example, what does FieldSectionMutationService depend on? includes calls made by its owned methods and imports; ambiguous names return an auto-picked answer plus alternatives. Fuzzy discovery retains useful traversal results with approximate labels and coverage caveats.

  • Route dependency, usage, impact, and connection phrases through native typed execution while preserving explicit discovery filters, bounds, and strict structured-command matching.
  • Default natural discovery to compact 800-token pages, with relationships before declarations and full evidence available on request.
  • Fix two impact traversal cutoffs: store posting chunk envelopes and premature termination after partial owner-level evidence.
  • Add regression coverage and a 25-question qualification suite across five pinned repositories.

Motivation

Implements Workstream 1 of the supplied Compass improvement plan: loose questions previously stopped at ambiguity or returned expensive non-answers despite useful graph evidence. Selection remains local, bounded, deterministic, and explicitly labelled. Workstreams 2–6 remain outside this PR, including attribute-chain extraction and broader impact relationship coverage.

Verification

cargo fmt --all -- --check                                      PASS
cargo clippy --workspace --lib --bins --locked -- -D warnings    PASS
cargo test --workspace --lib --bins --locked                     1,114 passed; 2 ignored

cargo test -p compass-query --test natural_intent --test natural_answers \
  --test natural_query_golden --test code_impact --test query_contract \
  -p compass-output --test agent_query \
  -p compass-cli --test code_query_cli --test compass_product --locked
                                                               105 passed
cargo test -p compass-query --test code_search --test code_traversal \
  --test export_binding_resolution --locked                     27 passed
sh scripts/check_product_boundary.sh                            PASS
python3 -m unittest discover -s benchmarks/agent_query/tests -q  240 passed
git diff --check                                               PASS

Cargo checks used a dedicated target directory on the mounted workspace volume. Language-extraction and CompassQL qualification gates were not rerun locally; those runtime surfaces are unchanged.

Before integration with the Compass 0.4.0 main branch, ran the agent query harness with suite_natural.toml, identical questions and 800-token first pages, and no follow-ups. Compass 0.3.29 (candidate debug binary SHA-256 prefix cb7c4fa20d4013a4) versus Graphify 0.9.67:

Measure Compass Graphify
Source-reviewed first-page oracle passes 24/25 (96%) 17/25 (68%)
Mean output across all questions, estimated tokens 544 757
Median output on the 17 questions both answered, estimated tokens 746 754
Median query wall time, ms 3,843 532

One broad Cobra question still omits a required anchor from the first page. Compass remains slower. Token estimates are UTF-8 output bytes divided by four; anchor recall is a focused proxy, not an independent precision oracle or a population-wide accuracy claim. This run does not qualify cold/warm latency, RSS, or production performance. A separate five-question loose-versus-exact-neighborhood check averaged 782 versus 788 estimated tokens at the same 800-token budget.

Reproduce with the documented harness, supplying the pinned source checkouts and binaries:

python3 benchmarks/agent_query/runner.py run \
  --suite benchmarks/agent_query/suite_natural.toml \
  --workspace <qualification-workspace> \
  --compass-binary <built-compass> --graphify-binary <graphify> \
  --source cobra=<cobra> --source flask=<flask> --source gson=<gson> \
  --source zod=<zod> --source axum=<axum-crate-root>

Merged main at 8a17f7c8 (Compass 0.4.0) and resolved the textual and semantic conflicts. Retained main's fresh-artifact benchmark policy, owner-qualified lookup rules, calls-only trails, and strict answer validation. Added coverage for dependency CLI operands and valid auto-picked candidate answers. The benchmark figures above describe the pre-merge binary; the full real-repository benchmark was not rerun for this conflict-resolution update.

Browser CI qualification follow-up

The startup timing test ran alongside another functional browser test on CI and exceeded the unchanged three-second limit on all three attempts (3,148 / 3,069 / 3,190 ms). Run browser performance qualification in a dedicated single-worker project after functional Chromium tests finish. Keep the original limits, fixtures, and assertions. The focused performance script uses --no-deps so it still runs only the timing checks.

Validation:

  • npm ci and npm run typecheck:js: passed.
  • npm run test:js: 473 unit tests passed; its initial six-worker browser run hit an unrelated transient history-status assertion.
  • CI=true npm run test -w @compass/viewer-tests -- --workers=2: all 120 browser tests passed, with both timing tests scheduled last in their own worker.
  • npm run test:performance -w @compass/viewer-tests -- --repeat-each=3: all six runs passed.
  • node scripts/check_viewer_assets.mjs and git diff --check: passed.

No viewer runtime or generated assets changed. VS Code integration/packaging and platform checks remain covered by CI.

Compatibility and documentation

Natural selection uses query-planner/2; explicit callers/callees/impact/node/search matching remains strict. Raw graph and query schema majors remain unchanged. Compact discovery text cursors advance to version 3, explicitly rejecting older cursors; restart the query to continue. The default text page budget changes from 8,000 to 800, with --text-budget and --cursor retaining caller control. No graph rebuild is needed. Auto-picked typed answers retain candidates/ambiguous Agent View status; owner-qualified selection preserves suffix validation and rejects missing owners or incomplete leaf postings.

Updated COMPATIBILITY.md, MIGRATION.md, CHANGELOG.md, PERFORMANCE.md, the query-engine documentation, and benchmark documentation. No runtime Graphify, Python, embedding, credential, or network dependency is introduced.

Checklist

  • The change is focused and excludes unrelated formatting or generated files
  • Tests cover changed behavior, or this pull request changes documentation only
  • User-facing commands, flags, limits, and examples are documented
  • Compatibility or migration effects are described
  • No credentials, private source code, or sensitive report details are included
  • I agree to license my contribution under MIT OR Apache-2.0
  • I followed the Compass code of conduct

Retain owner-qualified lookup, calls-only trails, and fresh benchmark artifacts from main. Integrate dependency intent with the CLI operand contract and keep auto-picked Agent View answers explicitly ambiguous candidates.
Run the unchanged startup timing gates in a single-worker Chromium project after functional tests complete. Keep the focused performance command independent of project dependencies and document the qualification setup.
@forhappy
forhappy merged commit bfe9f70 into main Sep 30, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant