Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,16 @@

## Unreleased

- Route natural-language usage, dependency, impact, and connection questions
into typed graph queries. Continue ambiguous questions with a disclosed,
deterministically ranked candidate and retain alternative identities.
- Include owned method calls in natural-language class dependency answers.
Present retained approximate results as useful candidates while preserving
provenance, coverage caveats, and strict structured-command matching.
- Default natural discovery to 800-token pages, show relationships first,
and version its text cursor to 3. Fix sub-chunk store relationship reads
that previously cut impact traversal short.

## 0.4.0 - 2026-09-28

- Resolve source-proven Go field selectors to exact named struct fields and
Expand Down
33 changes: 31 additions & 2 deletions COMPATIBILITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -496,8 +496,12 @@ layout declaration. See the supported boundaries in the
Compass adds the additive strict projection `compass.query.agent-view/1` for
typed CLI and MCP consumers. It is derived from, and digest-bound to, the raw
`compass.query/1` or `compass.query.discovery/1` response. The raw CLI `json`
shape, MCP `structuredContent.result`, graph schemas, and discovery
`compass.query.discovery-text-page/2` cursor meaning are unchanged.
shape, MCP `structuredContent.result`, and graph schemas are unchanged.
Discovery text pagination now uses
`compass.query.discovery-text-page/3`: compact pages show relationships before
node inventories. Older discovery cursors fail with an explicit version error;
restart the query to obtain a new cursor. Natural discovery text now defaults
to an 800-token page; `--text-budget N` and `--cursor` retain caller control.

The typed commands accept `--format agent-json`; default text is an
answer-first presentation. MCP keeps `compass.mcp.tool-result/1` and adds the
Expand All @@ -512,6 +516,31 @@ The additive `relationship_inconsistency` diagnostic extends the strict
TypeScript consumers and the checked-in manifest must accept the new value
before interpreting a relationship result that carries it.

### Natural-language answer selection

Natural-language queries use `query-planner/2` to recognize dependency, usage,
impact, and connection phrases. `compass query` preserves its
`compass.query.discovery/1` envelope while delegating recognized, unfiltered
questions to typed execution. Explicit direction, scope, context, and DFS
controls continue through filtered discovery. `compass ask` uses the same
operand selection policy.

Natural language may select a ranked candidate when the name is ambiguous.
The selected identity and alternative names remain explicit; inferred operand
matching does not change the provenance of structural edges. Explicit
`callers`, `callees`, `impact`, `node`, and `search` commands retain strict
identity behavior. Consumers must distinguish a useful candidate answer from
an exact match using the retained seeds and diagnostics. Discovery Agent View
now presents retained fuzzy/ambiguous neighborhoods as candidates, instead of
requiring resolution before showing their results. No graph or query schema
major changes. Existing cursors reject a changed semantic result by digest.

Typed Agent View also uses `resultState: "candidates"` with
`matchState: "ambiguous"` for disclosed auto-picked answers; their witnessed
relationships remain available. Strict unresolved ambiguity still uses
`needs_resolution`. Owner-qualified operands retain suffix validation and
reject missing owners or incomplete leaf-name postings before ranking.

### Typed text pages and store self-check

Typed commands (`ask`, `search`, `callers`, `callees`, `impact`, `explore`, and
Expand Down
15 changes: 15 additions & 0 deletions MIGRATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -609,3 +609,18 @@ compass install --platform codex --project
```

Keep the old Graphify installation and `graphify-out/` directory until the new `compass-out/` graph has passed your project checks. The two tools don't share runtime output paths.

## Natural-language candidate selection

`compass query` and `compass ask` now continue with a ranked symbol when a
natural-language operand has multiple matches. Review the auto-picked or
approximate label and the alternatives before attributing the result to the
original question. Use an exact node ID to override the choice. Explicit
structured commands retain their strict matching behavior. No graph rebuild
is needed for this query change; cursors created for a different semantic
answer fail explicitly and must be restarted.

Discovery text cursors now use version 3 because compact pages show relationships
before declarations. Restart a query if a saved version 2 cursor is rejected.
Natural discovery pages default to 800 tokens. Increase `--text-budget` or
follow the printed `--cursor` to retrieve more of the same bounded answer.
35 changes: 35 additions & 0 deletions PERFORMANCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,33 @@ Compare a proposed change with a previously approved Compass result captured on
the same runner and corpus. A median regression above 10% requires explicit
review and evidence explaining the tradeoff.

## Natural-language answer cost qualification

The developer-side [agent query harness](benchmarks/agent_query/README.md)
qualifies answer recall and output cost separately from runtime performance.
Run `suite_natural.toml` on its five pinned repositories with the same question
and an 800-token first page for each tool; no follow-up pages are counted.
The harness records binary identities, source revisions, bounded stdout,
source-reviewed oracle results, and query wall time. Token estimates use
UTF-8 output bytes divided by four, rather than a model tokenizer.

On 2026-09-30, a macOS debug build of Compass 0.3.29 with natural-language
candidate selection passed 24 of 25 questions, compared with 17 of 25 for
Graphify 0.9.67. Mean output cost across all questions was 544 versus 757
estimated tokens. On the 17 questions both tools answered, median output cost
was 746 versus 754 estimated tokens. One broad Cobra question still omitted
a required anchor from its first page.

These measurements precede integration with the Compass 0.4.0 main branch;
they remain historical evidence for that candidate binary, not measurements
of the merged build.

This focused anchor-recall check is not an independent precision oracle or a
population-wide accuracy claim. Compass's median query wall time was 3,843 ms,
compared with 532 ms for Graphify. This run does not qualify a latency
improvement, cold/warm cache behavior, peak RSS, or a production performance
baseline; those remain governed by the baseline policy above.

## Community detection performance

The 2026-09-12 Leiden hot-path qualification used Compass `0.3.24` candidate
Expand Down Expand Up @@ -1256,6 +1283,14 @@ controls, the static layout, the 200-row community DOM bound, and the visible
edge disclosure to appear within three seconds. These are runner-specific
diagnostic observations, not a cross-platform latency guarantee.

Browser wall-clock qualification runs in the single-worker
`chromium-performance` Playwright project after the functional Chromium tests
finish, so concurrent test pages cannot consume its startup budget. The
one-second small-graph and three-second large-graph limits and readiness
assertions remain unchanged. Run only this qualification with
`npm run test:performance -w @compass/viewer-tests`; the normal `npm run test:js`
includes both projects in order.

### Django parallel fact-state qualification

The 2026-08-05 large-repository fact-state hardening was measured from Compass
Expand Down
3 changes: 2 additions & 1 deletion benchmarks/agent_query/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,11 +9,12 @@ question/evidence matrix, including ask, communities, clusters, and god nodes.
The audit report distinguishes completed checks from surfaces still awaiting
source or design-quality judgments.

Four suites share the harness:
Five suites share the harness:

| Suite | Questions | Shape |
| --- | ---: | --- |
| `suite.toml` | 47 | The first five-repository suite, including Compass's compact and paged projections |
| `suite_natural.toml` | 25 | Same natural-language question on both tools, source-reviewed v2 oracles, 800-token pages, no follow-ups |
| `suite_v2.toml` | 50 | A blackbox-fair extension: same questions for both tools, default output forms, no tool-specific projections |
| `suite_fd.toml` | 12 | Separate pinned `sharkdp/fd` sample, recorded from source before either tool's first extraction/query run |
| `suite_ask.toml` | 10 | Same natural-language caller/callee questions and 2,000-token budget for Compass `ask` and Graphify `query` across five languages |
Expand Down
Loading
Loading