Skip to content

The CLI, MCP server and skill say only what is so - #610

Merged
danbarua merged 9 commits into
mainfrom
surfaces-say-what-is-so
Oct 1, 2026
Merged

danbarua merged 9 commits into
mainfrom
surfaces-say-what-is-so

Conversation

@danbarua

@danbarua danbarua commented Oct 1, 2026

Copy link
Copy Markdown
Owner

The CLI help, the MCP server's text and the MCP skill said things that were not so. Each item below says what changed and the output that shows it. Measurements are from bin/labkit built from this branch, against throwaway records, on 2026-10-01.

1. labkit --db X mcp serves the record in X

The CLI locates the record once (--db, LABKIT_DB_URL, LABKIT_HOME, then the working directory) and every command reads that answer, mcp included. The MCP server no longer reads LABKIT_TENANT or locates a record itself. --db help and docs/environment.md now give the real default order, which includes the git repository root. The LABKIT_TENANT and LABKIT_SOURCE rows in docs/environment.md are gone: the web seeder that read them was removed earlier, and no code reads either variable.

tests/mcp/mcp-stdio.test.ts starts cli.ts --db <dir> mcp from a different, empty working directory and asserts the record lands in <dir> and nothing lands in the working directory. With mcp ignoring --db, the same test fails:

✗ `labkit --db <dir> mcp` serves the record in that directory, not the working directory's
 7 pass
 1 fail

Over stdio against the binary, started from a different directory: labkit: creating a new record at …/mcpdb2/.labkit/pglite, and no .labkit in the working directory.

2. register_session is gone

Removed: the tool, SESSION_TOOLS, the write gate, the "Before anything" group, the registeredSession report schema, and the registry's reconstructedFrom slot, which only that tool set. The ACP tool set keeps its per-session registry. MCP writes are attributed as an unregistered write was. A live note over stdio returns:

"attribution_label": "mock-session",
"attribution_id": "mock-session-0",
"attribution_how": "claimed",

Tests that asserted the refusal or the registration are removed.

3. MCP text

The initialize instructions, the docs tool and the docs resource are built from the tools a server registers. From the binary:

instructions: LabKit is a research record: questions, the lines of enquiry pursuing them, what was measured, what was concluded, and the conditions results are held to. `why`, `search` and `work_list` read it. `note`, `record_observations`, `record_analysis`, `conclude` and `evaluate_criterion` change it. `docs` describes each tool.

With --read-only: "why, search and work_list read it. This server does not change the record. docs describes each tool." The docs page lists no write tool. A test asserts both.

  • why (MCP and CLI help) describes every kind of handle it explains, and says a proposition resolves only when exactly one claim asserts it. Measured: two claims asserting "y holds" give labkit: "y holds" is claimed 2 times; name one: CLM_7, CLM_8. The CLI help points at labkit get <handle>.
  • note says what it does, on both surfaces.
  • conclude states no rule about replacing. The old text said replacing is accepted only on an analysis that supersedes the one that concluded the finding; conclude COMP_5 --replacing CLM_4 on an analysis that revises nothing succeeded. The sentence about superseding without replacing on a revising analysis is also gone: no command records a revision.
  • The docs resource description no longer claims to list arguments and returns.

resources/list from the binary: "description": "What each tool on this server does, in prose, generated from the tool declarations."

4. CLI help

  • The root help's "Start with" list names labkit enquiries in place of labkit known, which does not exist.
  • is prints plain text, not **…**.
  • accept says it takes a line of enquiry and records a decision on its question.
  • amend --citing reads "the claim the amendment rests on"; nothing requires it once the condition has been evaluated.
  • analyse --from is optional. labkit analyse LOE_1 --method "on paper" printed COMP_13 EU_13 ART_13.

Every labkit <command> named in the root help and in all 35 command and subcommand help pages was checked against the binary's command list; each one exists.

5. Colour

--no-ansi, a non-empty NO_COLOR and FORCE_COLOR=0 turn colour off on stdout and stderr. Error lines were written with console.error, which Bun colours red regardless; they now go through one colour decision. picocolors read FORCE_COLOR=0 as on. Stderr bytes of labkit why CLM_999:

FORCE_COLOR=1                 033 [ 3 1 m l a b k i t : …
FORCE_COLOR=1 --no-ansi       l a b k i t : …
FORCE_COLOR=1 NO_COLOR=1      l a b k i t : …
FORCE_COLOR=0                 l a b k i t : …

labkit --no-ansi get NOPE_1 and labkit --no-ansi list nope print 0 bytes on stdout and an uncoloured line on stderr. tests/cli/colour.test.ts covers these cases, plus empty stdout on a refusal. Run against the previous code, the three "off" cases fail.

6. close enquiry --answered-by is repeatable

answeredBy is a list. The closing decision gets an ANSWERS edge for each claim and BASED_ON edges for the findings under each. answered is a list on the close result, on an enquiry's status and on each closed pursuit in the survey, and empty when the enquiry was abandoned. An enquiry's status counts as resting on confirmatory work only when every answering claim does, which is the survey's rule for established. Old single-id events are not supported, as the owner asked.

$ labkit close enquiry LOE_7 --answered-by CLM_10 --answered-by CLM_11
DEC_12
$ labkit --json why LOE_7 | jq -c '.report.enquiry.answered'
[{"claim":"CLM_10","asserts":"primer helps"},{"claim":"CLM_11","asserts":"primer lasts"}]

7. Accepted enquiries are listed as accepted

The enquiry list reads an ACCEPTS decision on the enquiry's question.

$ labkit enquiries
Enquiries — 1
accepted  LOE_1  does the coating slow corrosion?

accept help says labkit enquiries lists the enquiry as accepted and labkit --json why carries the reason and the reopening condition ({"acceptedBecause":"no more samples","reopensIf":"a new batch"}). Plain labkit why does not print them.

8. Prose

  • README.md is deleted.
  • The mcp-in-context skill, scripts/mcp-call.sh and references/response-anatomy.md named labkit gate, now and gate_status tools, and a register_session refusal. They now use why and work_list, with payloads captured from this binary.
  • The MCP inspector takes --db and --read-only for its own options when they follow the binary, so labkit started without them. For example, … --cli bin/labkit mcp --read-only --tool-name note wrote a note. mcp-call.sh puts the server's flags in a wrapper script, takes the record from LABKIT_RECORD as --db, and starts the server with --read-only unless MCP_CALL_WRITE=1. The read-only default is my choice, not something you asked for: before this change the script was read-only only because writes needed a registration. A note through it returns tool_not_found by default; with MCP_CALL_WRITE=1 it writes NOTE_4.
  • build-scenario-docs.ts writes labkit why, not labkit why-supported. Regenerated: S-3b, S-3c, S-11, S-11e.

A repo-wide grep (excluding node_modules, .git, bin) for register_session, gate_status, why-supported, labkit known, SESSION_TOOLS and LABKIT_TENANT finds nothing. A grep for known as a command name finds comments in scripts/checks/smoke-cli.sh:114, scripts/db/probe-bonsai-replay.sh:16, packages/core-domain/write/work.ts:75, packages/app-cli/views/format.ts:31, tests/cli/views.test.ts:507 and tests/domain/survey-after-standing-changes.test.ts:2. None of them runs the command, and they are left alone. known is also a field of the now report.

Checks

  • bun run test on main at 3de1f9a: 1639 pass, 2 skip, 1 fail (1642). On this branch: 1644 pass, 2 skip, 1 fail (1647). The failure is the same on both: packages/app-acp/http.test.ts "the CLI serves --http until SIGTERM…" receives an escape sequence after the URL on stderr when FORCE_COLOR is set in the environment. It is outside this change.
  • bun run check:quick: all 25 passed. check:binary, check:cli and check:scenario-docs pass.

Found, not changed

  • labkit mcp accepts --author, --reconstructed-from and --date and ignores them. MCP now has no way to set reconstructedFrom. Both belong to attribution.
  • After undo of a close, enquiryStatus still reports the enquiry closed, because it does not filter a retracted closing decision. The survey does filter it. The enquiry list does not filter it either.
  • dump and restore still read --db themselves rather than the located record (dump.ts was not touched).
  • The record daemon's log lines still go through console.error.
  • infra/ci/triggers.tf still lists the root README.md among paths that skip a build.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY

danbarua and others added 9 commits October 1, 2026 01:10
The CLI locates the record once, from --db, LABKIT_DB_URL, LABKIT_HOME and
the working directory, and every command reads it from there, mcp included.
The MCP server no longer reads LABKIT_TENANT or locates a record of its own;
--tenant and --db are the root's flags.

docs/environment.md: the --db/LABKIT_HOME default names the git-root step it
skipped; the LABKIT_TENANT and LABKIT_SOURCE rows went with the web seeder
that read them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
…sters

register_session, its write gate, the "Before anything" group, the
registeredSession report and the registry's reconstructedFrom slot are
removed. MCP writes carry the stand-in session (mock-session), as an
unregistered write did. The ACP tool set keeps its per-session registry.

The initialize instructions, the docs tool and the docs resource are built
from the tools a server registers, so a read-only server names no write
tool. The instructions no longer name a `now` tool. `why` says what it
explains for every kind of handle and that a proposition resolves only when
exactly one claim asserts it, on both surfaces; the CLI help points at
`get`. `note` says what it does. `conclude` no longer states a rule about
`replacing` that nothing enforces, and says when a conclusion supersedes
without it. The docs resource no longer claims to list arguments and
returns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
The root help's "Start with" list names `labkit enquiries` in place of
`labkit known`, which does not exist. `is` prints plain text. `accept` says
it takes a line of enquiry and what it records. `amend --citing` no longer
says it is required once the condition has been evaluated; nothing requires
it. `analyse --from` is optional, as the domain already allowed.

The enquiry list carries `accepted`, read from an ACCEPTS decision on the
enquiry's question, and `labkit enquiries` shows an open enquiry in that
state as `accepted` rather than `running`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
One decision, colourWanted, colours stdout and the CLI's error lines. The
error lines were written with console.error, which Bun colours red whatever
--no-ansi or NO_COLOR say; they now go through process.stderr.write. The
stdout decision came from picocolors, which reads FORCE_COLOR=0 as on.
core-db's "connection lost" line is written plain for the same reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
The CLI kept only the last --answered-by. `answeredBy` is now a list: the
closing decision gets an ANSWERS edge for each claim and a BASED_ON edge
for the findings under each, and its reason names every answer.

`answered` is a list on the close's result, on an enquiry's status and on
each closed pursuit in the survey, empty when the enquiry was abandoned.
An enquiry's status rests on confirmatory work only when every answering
claim does, the rule the survey uses for established.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
The mcp-in-context skill, its mcp-call.sh and response-anatomy.md named a
`labkit gate` command, `now` and `gate_status` tools and a register_session
refusal, none of which exist. They now use `why` and `work_list`, with
payloads captured from the current binary. The inspector keeps --db and
--read-only for itself when they follow the binary, so mcp-call.sh passes
them through a wrapper script and serves --read-only unless
MCP_CALL_WRITE=1; writes over the inspector are no longer refused, and the
skill says so.

The scenario docs name `labkit why`, not `labkit why-supported`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
`labkit why` on an accepted enquiry prints that its question is accepted,
not the reason or the reopening condition; `--json` carries both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
…othing on stdout

`conclude` no longer describes superseding without `replacing` on a revising
analysis: no verb records a revision. `why` names the evidence unit among the
kinds it walks. The colour test also asserts that `get` and `why` on a missing
handle print nothing on stdout.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
@github-actions github-actions Bot added area:record The record: core-domain, core-db, the CLI, the MCP server, the domain model and workspaces area:agent The agent and its hosts: core-agent, app-acp, app-vscode, acp-fake, acp-scenarios, acp-transcripts labels Oct 1, 2026
@danbarua
danbarua enabled auto-merge (squash) October 1, 2026 00:17
@danbarua
danbarua merged commit e4250a4 into main Oct 1, 2026
10 checks passed
@danbarua
danbarua deleted the surfaces-say-what-is-so branch October 1, 2026 00:18
danbarua added a commit that referenced this pull request Oct 1, 2026
…ted enquiries say why (#611)

The follow-up to #608, #609 and #610. Measured with the compiled binary
(`bun run cli:build`) against throwaway records, always with `--db`. The
"before" binary was built from `main` at e4250a4.

## 1. Reads stop handling shapes only the retired verbs wrote

For each shape, I grepped every staged write in
`packages/core-domain/write/*.ts` (`unitOfWork.node`/`edge`).
`UnitOfWork.set` and `setEdge` have no callers, so no live write sets
any property, `retracted` included. None of the shapes below is produced
by a live write:

| shape | written only by | what went |
|---|---|---|
| `review` kind in `why` | `recordReview` | **kept**, see below.
Removed: the `ReviewRef` alias and the minted `reviewRef` codec |
| sharpening arm of `originAlready` (Decision MOTIVATES Question +
SHARPENS) | `sharpen` | the arm; `originAlready` now reads only the note
|
| `closed` gate state (Decision CLOSES Gate) | `closeGate` | the state,
`GateStatus.closure`, the `closed` lookups in `gateStatus`/`gateList`,
the `explainGate` case, `closed` in `GATE_STATES` and in the `why`
descriptions |
| `undecided` standing (GRADES) | `isUndecided` | the verdict, the
standing, the GRADES lookup in `standingsConferred`, the view branch,
the vocabulary word |
| retracted nodes | `undo` | every `retracted IS NULL` / `retracted !==
true` filter in `explain`, `happened`, `inventory`, `standing`, `core`;
`everGated` and the unlabelled `ever` match in `workList`; `retractedIn`
and the `retracting` line in `happened` |
| `revisedBy` (Decision MOTIVATES/SUPERSEDES Computation) and the review
behind it | `replaceAnalysis` | `revisedBy`, `impliedSupersession`,
`conclusionsOf`, the BASED_ON-review edge in `concluding`; the Review
match and `because` on `analysisRevision`; the Review lookups in
`whySupported` (EVALUATES, BASED_ON Review, INVALIDATED_BY) |

**Not done: the `review` kind.** `search` walks every searchable label
in `SEARCHABLE_TEXT` (core-db/domain.ts) and refuses a label with no
kind. `Review` is one of them. With the kind removed, every search
failed:

```
$ labkit --db …/rec1 --no-ansi search seed
labkit: Review is searchable but names no research-concept kind
exit 1
```

Removing the kind needs one of two changes that this PR does not make:
take `Review` out of the schema's searchable annotations, or make
`search` skip a label with no kind (its comment says that case should
fail loudly). The kind is back, and so is ", review" in the `why`
descriptions.

Tests: the undecided view test and the undecided verdict case are gone;
the `workStateFrom` test lost its retracted-gate case;
`tests/events/retraction.test.ts` is gone, since it checked that reads
hide retracted nodes. `closed` leaves the CLI vocabulary because
`tests/cli/vocabulary.test.ts` failed once no schema could say it (`+
"closed"`). The enquiry list still prints `closed`, uncoloured.

The same reads over two records, before and after: the throwaway record
(`happened` 52 lines, `now` 16 lines, `why CLM_6` on a replaced claim)
and the rebuilt Bonsai record (`now` 82 lines, `claims` 51, `enquiries`
29, `work` 11, `gates` 20). `diff` printed nothing for any of them. The
one visible change is `gates --state closed`:

```
before: Gates — 0 / nothing (exit 0)
after:  labkit: Invalid option: expected one of "never-evaluated"|"incomplete"|"blocked"|"satisfied" (exit 1)
```

`bun run bonsai:record` with this binary:

```
107/107 criteria recorded
107/107 evaluations recorded, 57 citing a measurement
gate GATE_243, task TASK_135, enquiry LOE_134, 107 criteria/evaluations from experiments/stage2b_denoising/gates.toml
OK: the live event stream and graph state are exactly what the scripts produce, commit hashes and clocks aside.
OK: record built (…/bin/labkit --db …/tmp/bonsai), replay clean.
```

### What nothing writes any more

Graph schema and migrations are unchanged. Declared: 14 node labels and
31 edge labels; written: 13 and 25.

- **Label:** `Review`.
- **Edge labels:** `REVERIFIES`, `GRADES`, `KEEPS`, `SHARPENS`,
`EVALUATES`, `INVALIDATED_BY`.
- **Endpoint pairs on edges still written elsewhere:** MOTIVATES
Decision→Question, MOTIVATES Decision→Computation, SUPERSEDES
Decision→Computation, CLOSES Decision→Gate, BASED_ON Decision→Review,
GATES Gate→Computation, PRODUCES Task→Computation, PRODUCES
Task→Artefact, and CONCERNS/MENTIONS Note→Review.
- **Properties:** `retracted` (nothing calls `UnitOfWork.set`).
`ClaimProps.kind` still allows `"undecided"`. `provisioning.ts` still
installs the `<label>_hide_retracted` RLS policy, and with
`retraction.test.ts` gone no test exercises it.

Reads of these that were outside this brief and are still in place:
`reverifiedBy`/REVERIFIES in `whySupported` and the "Re-checked by"
view; the Computation lineage, KEEPS and the changed/restated/unpaired
lists in `analysisRevision`, which now always return the empty shape;
and the withdrawn-verdict helper in `checks.ts`.

## 2. `labkit gates` lines its columns up with colour on

`rows()` padded by string length, escape codes included. It now pads by
visible width. The test renders `renderGateList` with colour on, strips
the escape codes, and compares the result with the plain rendering. It
failed before the fix:

```
- "blocked     GATE_1  no publication
+ "blocked              GATE_1  no publication
```

Compiled binary, `FORCE_COLOR=1 labkit gates | strip-ansi`:

```
before:                                     after:
never-evaluated      GATE_4  no publication   never-evaluated  GATE_4  no publication
holding up  TASK_3  write it up               holding up       TASK_3  write it up
```

## 3. The record daemon's log lines take the CLI's colour decision

`colourWanted` moves to `packages/core-db/colour.ts`, because core-db
cannot import app-cli. `stderrLine` sits beside it and writes one line
to stderr, in red when `colourWanted` says so. The daemon's `log()`, the
wire host's two lines, the daemon entry's usage line, the client's
"replacing it" line and the CLI's refusals all go through it. core-db
has no access to `--no-ansi`, so the daemon decides from the environment
alone. Bun colours `console.error` whenever `FORCE_COLOR` is set, even
beside `NO_COLOR=1`:

```
$ echo 'console.error("x")' > ce.ts
$ NO_COLOR=1 bun ce.ts 2>&1 | od -c     # FORCE_COLOR=3 inherited from the shell
033 [ 0 m 033 [ 3 1 m x 033 [ 0 m \n
```

Daemon log, fresh `--db` each, under `FORCE_COLOR=3 NO_COLOR=1`:

```
before: ^[[0m^[[31m2026-10-01T00:30:39.226Z [labkit daemon 63989] serving …^[[0m
after:  2026-10-01T00:30:42.025Z [labkit daemon 64191] serving …
```

`tests/persistence/daemon-log-colour.test.ts` runs the daemon entry
under `FORCE_COLOR=1` (the control, which must be coloured),
`FORCE_COLOR=3 NO_COLOR=1` and `FORCE_COLOR=0`.

## 4. The `--http` launcher test clears the colour variables for its
child

The child now gets `process.env` without `FORCE_COLOR`, `NO_COLOR`, `CI`
and `TERM`, which are the variables `colourWanted` reads.

```
before: FORCE_COLOR=3 bun test packages/app-acp/http.test.ts
  Received: "http://127.0.0.1:55791/acp\u001B[0m"
  13 pass, 1 fail
after:  FORCE_COLOR=3 bun test packages/app-acp/http.test.ts
  14 pass, 0 fail
```

## 5. `labkit why` on an accepted enquiry prints the reason and the
reopening condition

```
before:
LOE_1 is open — its question is accepted as unresolved because
  - (Q_1)  does the schedule move convergence? — currently accepted
after:
LOE_1 is open — its question is accepted as unresolved because
  - (Q_1)  does the schedule move convergence? — currently accepted
accepted because: the effect is below seed noise
reopens if: a run with ten seeds
```

The `accept` help text now points at `labkit why` and no longer at
`labkit --json why`. `tests/cli/enquiries.test.ts` asserts both lines
through the CLI.

## 6. The deleted root README

`infra/ci/triggers.tf` no longer lists `README.md` in `ignored_files`.
`infra/ci/README.md` had the same list in prose, and it is updated too.
No other file referred to the root README; every other `README.md` hit
is a package README or test data.

## Not in this PR: removing the MCP read-only mode

The coordinator also passed on a request to remove `labkit mcp
--read-only`, the `readOnly` paths in `app-mcp/server.ts` and `docs.ts`
with their tests, and the `MCP_CALL_WRITE` switch in the skill's
`mcp-call.sh`. I made the change, and typecheck, the MCP tests (22 pass)
and `mcp-call.sh tools/list` (9 tools, 4 with `readOnlyHint`) all
passed. The permission system then refused the commit as test removal.
The change is not in this branch. It is saved as a patch outside the
repo for Dan to decide on.

## Checks

- `bun test`, before: 1644 pass, 2 skip, **1 fail** (`http.test.ts`
under this shell's `FORCE_COLOR=3`), 1647 tests across 221 files.
- `bun test`, after: **1647 pass, 2 skip, 0 fail**, 1649 tests across
221 files. The difference: +1 gates-colour test, +3 daemon-colour tests
in a new file, −1 undecided view test, −1 `retraction.test.ts` (one
test, one file).
- `bun run check:quick`: all 25 passed.
- Rebased onto `origin/main` (e4250a4) before pushing.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:agent The agent and its hosts: core-agent, app-acp, app-vscode, acp-fake, acp-scenarios, acp-transcripts area:record The record: core-domain, core-db, the CLI, the MCP server, the domain model and workspaces

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant