The CLI, MCP server and skill say only what is so - #610
Merged
Merged
Conversation
The CLI locates the record once, from --db, LABKIT_DB_URL, LABKIT_HOME and the working directory, and every command reads it from there, mcp included. The MCP server no longer reads LABKIT_TENANT or locates a record of its own; --tenant and --db are the root's flags. docs/environment.md: the --db/LABKIT_HOME default names the git-root step it skipped; the LABKIT_TENANT and LABKIT_SOURCE rows went with the web seeder that read them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
…sters register_session, its write gate, the "Before anything" group, the registeredSession report and the registry's reconstructedFrom slot are removed. MCP writes carry the stand-in session (mock-session), as an unregistered write did. The ACP tool set keeps its per-session registry. The initialize instructions, the docs tool and the docs resource are built from the tools a server registers, so a read-only server names no write tool. The instructions no longer name a `now` tool. `why` says what it explains for every kind of handle and that a proposition resolves only when exactly one claim asserts it, on both surfaces; the CLI help points at `get`. `note` says what it does. `conclude` no longer states a rule about `replacing` that nothing enforces, and says when a conclusion supersedes without it. The docs resource no longer claims to list arguments and returns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
The root help's "Start with" list names `labkit enquiries` in place of `labkit known`, which does not exist. `is` prints plain text. `accept` says it takes a line of enquiry and what it records. `amend --citing` no longer says it is required once the condition has been evaluated; nothing requires it. `analyse --from` is optional, as the domain already allowed. The enquiry list carries `accepted`, read from an ACCEPTS decision on the enquiry's question, and `labkit enquiries` shows an open enquiry in that state as `accepted` rather than `running`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
One decision, colourWanted, colours stdout and the CLI's error lines. The error lines were written with console.error, which Bun colours red whatever --no-ansi or NO_COLOR say; they now go through process.stderr.write. The stdout decision came from picocolors, which reads FORCE_COLOR=0 as on. core-db's "connection lost" line is written plain for the same reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
The CLI kept only the last --answered-by. `answeredBy` is now a list: the closing decision gets an ANSWERS edge for each claim and a BASED_ON edge for the findings under each, and its reason names every answer. `answered` is a list on the close's result, on an enquiry's status and on each closed pursuit in the survey, empty when the enquiry was abandoned. An enquiry's status rests on confirmatory work only when every answering claim does, the rule the survey uses for established. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
The mcp-in-context skill, its mcp-call.sh and response-anatomy.md named a `labkit gate` command, `now` and `gate_status` tools and a register_session refusal, none of which exist. They now use `why` and `work_list`, with payloads captured from the current binary. The inspector keeps --db and --read-only for itself when they follow the binary, so mcp-call.sh passes them through a wrapper script and serves --read-only unless MCP_CALL_WRITE=1; writes over the inspector are no longer refused, and the skill says so. The scenario docs name `labkit why`, not `labkit why-supported`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
`labkit why` on an accepted enquiry prints that its question is accepted, not the reason or the reopening condition; `--json` carries both. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
…othing on stdout `conclude` no longer describes superseding without `replacing` on a revising analysis: no verb records a revision. `why` names the evidence unit among the kinds it walks. The colour test also asserts that `get` and `why` on a missing handle print nothing on stdout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY
danbarua
enabled auto-merge (squash)
October 1, 2026 00:17
danbarua
added a commit
that referenced
this pull request
Oct 1, 2026
…ted enquiries say why (#611) The follow-up to #608, #609 and #610. Measured with the compiled binary (`bun run cli:build`) against throwaway records, always with `--db`. The "before" binary was built from `main` at e4250a4. ## 1. Reads stop handling shapes only the retired verbs wrote For each shape, I grepped every staged write in `packages/core-domain/write/*.ts` (`unitOfWork.node`/`edge`). `UnitOfWork.set` and `setEdge` have no callers, so no live write sets any property, `retracted` included. None of the shapes below is produced by a live write: | shape | written only by | what went | |---|---|---| | `review` kind in `why` | `recordReview` | **kept**, see below. Removed: the `ReviewRef` alias and the minted `reviewRef` codec | | sharpening arm of `originAlready` (Decision MOTIVATES Question + SHARPENS) | `sharpen` | the arm; `originAlready` now reads only the note | | `closed` gate state (Decision CLOSES Gate) | `closeGate` | the state, `GateStatus.closure`, the `closed` lookups in `gateStatus`/`gateList`, the `explainGate` case, `closed` in `GATE_STATES` and in the `why` descriptions | | `undecided` standing (GRADES) | `isUndecided` | the verdict, the standing, the GRADES lookup in `standingsConferred`, the view branch, the vocabulary word | | retracted nodes | `undo` | every `retracted IS NULL` / `retracted !== true` filter in `explain`, `happened`, `inventory`, `standing`, `core`; `everGated` and the unlabelled `ever` match in `workList`; `retractedIn` and the `retracting` line in `happened` | | `revisedBy` (Decision MOTIVATES/SUPERSEDES Computation) and the review behind it | `replaceAnalysis` | `revisedBy`, `impliedSupersession`, `conclusionsOf`, the BASED_ON-review edge in `concluding`; the Review match and `because` on `analysisRevision`; the Review lookups in `whySupported` (EVALUATES, BASED_ON Review, INVALIDATED_BY) | **Not done: the `review` kind.** `search` walks every searchable label in `SEARCHABLE_TEXT` (core-db/domain.ts) and refuses a label with no kind. `Review` is one of them. With the kind removed, every search failed: ``` $ labkit --db …/rec1 --no-ansi search seed labkit: Review is searchable but names no research-concept kind exit 1 ``` Removing the kind needs one of two changes that this PR does not make: take `Review` out of the schema's searchable annotations, or make `search` skip a label with no kind (its comment says that case should fail loudly). The kind is back, and so is ", review" in the `why` descriptions. Tests: the undecided view test and the undecided verdict case are gone; the `workStateFrom` test lost its retracted-gate case; `tests/events/retraction.test.ts` is gone, since it checked that reads hide retracted nodes. `closed` leaves the CLI vocabulary because `tests/cli/vocabulary.test.ts` failed once no schema could say it (`+ "closed"`). The enquiry list still prints `closed`, uncoloured. The same reads over two records, before and after: the throwaway record (`happened` 52 lines, `now` 16 lines, `why CLM_6` on a replaced claim) and the rebuilt Bonsai record (`now` 82 lines, `claims` 51, `enquiries` 29, `work` 11, `gates` 20). `diff` printed nothing for any of them. The one visible change is `gates --state closed`: ``` before: Gates — 0 / nothing (exit 0) after: labkit: Invalid option: expected one of "never-evaluated"|"incomplete"|"blocked"|"satisfied" (exit 1) ``` `bun run bonsai:record` with this binary: ``` 107/107 criteria recorded 107/107 evaluations recorded, 57 citing a measurement gate GATE_243, task TASK_135, enquiry LOE_134, 107 criteria/evaluations from experiments/stage2b_denoising/gates.toml OK: the live event stream and graph state are exactly what the scripts produce, commit hashes and clocks aside. OK: record built (…/bin/labkit --db …/tmp/bonsai), replay clean. ``` ### What nothing writes any more Graph schema and migrations are unchanged. Declared: 14 node labels and 31 edge labels; written: 13 and 25. - **Label:** `Review`. - **Edge labels:** `REVERIFIES`, `GRADES`, `KEEPS`, `SHARPENS`, `EVALUATES`, `INVALIDATED_BY`. - **Endpoint pairs on edges still written elsewhere:** MOTIVATES Decision→Question, MOTIVATES Decision→Computation, SUPERSEDES Decision→Computation, CLOSES Decision→Gate, BASED_ON Decision→Review, GATES Gate→Computation, PRODUCES Task→Computation, PRODUCES Task→Artefact, and CONCERNS/MENTIONS Note→Review. - **Properties:** `retracted` (nothing calls `UnitOfWork.set`). `ClaimProps.kind` still allows `"undecided"`. `provisioning.ts` still installs the `<label>_hide_retracted` RLS policy, and with `retraction.test.ts` gone no test exercises it. Reads of these that were outside this brief and are still in place: `reverifiedBy`/REVERIFIES in `whySupported` and the "Re-checked by" view; the Computation lineage, KEEPS and the changed/restated/unpaired lists in `analysisRevision`, which now always return the empty shape; and the withdrawn-verdict helper in `checks.ts`. ## 2. `labkit gates` lines its columns up with colour on `rows()` padded by string length, escape codes included. It now pads by visible width. The test renders `renderGateList` with colour on, strips the escape codes, and compares the result with the plain rendering. It failed before the fix: ``` - "blocked GATE_1 no publication + "blocked GATE_1 no publication ``` Compiled binary, `FORCE_COLOR=1 labkit gates | strip-ansi`: ``` before: after: never-evaluated GATE_4 no publication never-evaluated GATE_4 no publication holding up TASK_3 write it up holding up TASK_3 write it up ``` ## 3. The record daemon's log lines take the CLI's colour decision `colourWanted` moves to `packages/core-db/colour.ts`, because core-db cannot import app-cli. `stderrLine` sits beside it and writes one line to stderr, in red when `colourWanted` says so. The daemon's `log()`, the wire host's two lines, the daemon entry's usage line, the client's "replacing it" line and the CLI's refusals all go through it. core-db has no access to `--no-ansi`, so the daemon decides from the environment alone. Bun colours `console.error` whenever `FORCE_COLOR` is set, even beside `NO_COLOR=1`: ``` $ echo 'console.error("x")' > ce.ts $ NO_COLOR=1 bun ce.ts 2>&1 | od -c # FORCE_COLOR=3 inherited from the shell 033 [ 0 m 033 [ 3 1 m x 033 [ 0 m \n ``` Daemon log, fresh `--db` each, under `FORCE_COLOR=3 NO_COLOR=1`: ``` before: ^[[0m^[[31m2026-10-01T00:30:39.226Z [labkit daemon 63989] serving …^[[0m after: 2026-10-01T00:30:42.025Z [labkit daemon 64191] serving … ``` `tests/persistence/daemon-log-colour.test.ts` runs the daemon entry under `FORCE_COLOR=1` (the control, which must be coloured), `FORCE_COLOR=3 NO_COLOR=1` and `FORCE_COLOR=0`. ## 4. The `--http` launcher test clears the colour variables for its child The child now gets `process.env` without `FORCE_COLOR`, `NO_COLOR`, `CI` and `TERM`, which are the variables `colourWanted` reads. ``` before: FORCE_COLOR=3 bun test packages/app-acp/http.test.ts Received: "http://127.0.0.1:55791/acp\u001B[0m" 13 pass, 1 fail after: FORCE_COLOR=3 bun test packages/app-acp/http.test.ts 14 pass, 0 fail ``` ## 5. `labkit why` on an accepted enquiry prints the reason and the reopening condition ``` before: LOE_1 is open — its question is accepted as unresolved because - (Q_1) does the schedule move convergence? — currently accepted after: LOE_1 is open — its question is accepted as unresolved because - (Q_1) does the schedule move convergence? — currently accepted accepted because: the effect is below seed noise reopens if: a run with ten seeds ``` The `accept` help text now points at `labkit why` and no longer at `labkit --json why`. `tests/cli/enquiries.test.ts` asserts both lines through the CLI. ## 6. The deleted root README `infra/ci/triggers.tf` no longer lists `README.md` in `ignored_files`. `infra/ci/README.md` had the same list in prose, and it is updated too. No other file referred to the root README; every other `README.md` hit is a package README or test data. ## Not in this PR: removing the MCP read-only mode The coordinator also passed on a request to remove `labkit mcp --read-only`, the `readOnly` paths in `app-mcp/server.ts` and `docs.ts` with their tests, and the `MCP_CALL_WRITE` switch in the skill's `mcp-call.sh`. I made the change, and typecheck, the MCP tests (22 pass) and `mcp-call.sh tools/list` (9 tools, 4 with `readOnlyHint`) all passed. The permission system then refused the commit as test removal. The change is not in this branch. It is saved as a patch outside the repo for Dan to decide on. ## Checks - `bun test`, before: 1644 pass, 2 skip, **1 fail** (`http.test.ts` under this shell's `FORCE_COLOR=3`), 1647 tests across 221 files. - `bun test`, after: **1647 pass, 2 skip, 0 fail**, 1649 tests across 221 files. The difference: +1 gates-colour test, +3 daemon-colour tests in a new file, −1 undecided view test, −1 `retraction.test.ts` (one test, one file). - `bun run check:quick`: all 25 passed. - Rebased onto `origin/main` (e4250a4) before pushing. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Open
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The CLI help, the MCP server's text and the MCP skill said things that were not so. Each item below says what changed and the output that shows it. Measurements are from
bin/labkitbuilt from this branch, against throwaway records, on 2026-10-01.1.
labkit --db X mcpserves the record in XThe CLI locates the record once (
--db,LABKIT_DB_URL,LABKIT_HOME, then the working directory) and every command reads that answer,mcpincluded. The MCP server no longer readsLABKIT_TENANTor locates a record itself.--dbhelp anddocs/environment.mdnow give the real default order, which includes the git repository root. TheLABKIT_TENANTandLABKIT_SOURCErows indocs/environment.mdare gone: the web seeder that read them was removed earlier, and no code reads either variable.tests/mcp/mcp-stdio.test.tsstartscli.ts --db <dir> mcpfrom a different, empty working directory and asserts the record lands in<dir>and nothing lands in the working directory. Withmcpignoring--db, the same test fails:Over stdio against the binary, started from a different directory:
labkit: creating a new record at …/mcpdb2/.labkit/pglite, and no.labkitin the working directory.2.
register_sessionis goneRemoved: the tool,
SESSION_TOOLS, the write gate, the "Before anything" group, theregisteredSessionreport schema, and the registry'sreconstructedFromslot, which only that tool set. The ACP tool set keeps its per-session registry. MCP writes are attributed as an unregistered write was. A livenoteover stdio returns:Tests that asserted the refusal or the registration are removed.
3. MCP text
The
initializeinstructions, thedocstool and the docs resource are built from the tools a server registers. From the binary:With
--read-only: "why,searchandwork_listread it. This server does not change the record.docsdescribes each tool." The docs page lists no write tool. A test asserts both.why(MCP and CLI help) describes every kind of handle it explains, and says a proposition resolves only when exactly one claim asserts it. Measured: two claims asserting "y holds" givelabkit: "y holds" is claimed 2 times; name one: CLM_7, CLM_8. The CLI help points atlabkit get <handle>.notesays what it does, on both surfaces.concludestates no rule aboutreplacing. The old text saidreplacingis accepted only on an analysis that supersedes the one that concluded the finding;conclude COMP_5 --replacing CLM_4on an analysis that revises nothing succeeded. The sentence about superseding withoutreplacingon a revising analysis is also gone: no command records a revision.resources/listfrom the binary:"description": "What each tool on this server does, in prose, generated from the tool declarations."4. CLI help
labkit enquiriesin place oflabkit known, which does not exist.isprints plain text, not**…**.acceptsays it takes a line of enquiry and records a decision on its question.amend --citingreads "the claim the amendment rests on"; nothing requires it once the condition has been evaluated.analyse --fromis optional.labkit analyse LOE_1 --method "on paper"printedCOMP_13 EU_13 ART_13.Every
labkit <command>named in the root help and in all 35 command and subcommand help pages was checked against the binary's command list; each one exists.5. Colour
--no-ansi, a non-emptyNO_COLORandFORCE_COLOR=0turn colour off on stdout and stderr. Error lines were written withconsole.error, which Bun colours red regardless; they now go through one colour decision. picocolors readFORCE_COLOR=0as on. Stderr bytes oflabkit why CLM_999:labkit --no-ansi get NOPE_1andlabkit --no-ansi list nopeprint 0 bytes on stdout and an uncoloured line on stderr.tests/cli/colour.test.tscovers these cases, plus empty stdout on a refusal. Run against the previous code, the three "off" cases fail.6.
close enquiry --answered-byis repeatableansweredByis a list. The closing decision gets an ANSWERS edge for each claim and BASED_ON edges for the findings under each.answeredis a list on the close result, on an enquiry's status and on each closed pursuit in the survey, and empty when the enquiry was abandoned. An enquiry's status counts as resting on confirmatory work only when every answering claim does, which is the survey's rule for established. Old single-id events are not supported, as the owner asked.7. Accepted enquiries are listed as accepted
The enquiry list reads an ACCEPTS decision on the enquiry's question.
accepthelp sayslabkit enquirieslists the enquiry as accepted andlabkit --json whycarries the reason and the reopening condition ({"acceptedBecause":"no more samples","reopensIf":"a new batch"}). Plainlabkit whydoes not print them.8. Prose
README.mdis deleted.mcp-in-contextskill,scripts/mcp-call.shandreferences/response-anatomy.mdnamedlabkit gate,nowandgate_statustools, and aregister_sessionrefusal. They now usewhyandwork_list, with payloads captured from this binary.--dband--read-onlyfor its own options when they follow the binary, so labkit started without them. For example,… --cli bin/labkit mcp --read-only --tool-name notewrote a note.mcp-call.shputs the server's flags in a wrapper script, takes the record fromLABKIT_RECORDas--db, and starts the server with--read-onlyunlessMCP_CALL_WRITE=1. The read-only default is my choice, not something you asked for: before this change the script was read-only only because writes needed a registration. Anotethrough it returnstool_not_foundby default; withMCP_CALL_WRITE=1it writesNOTE_4.build-scenario-docs.tswriteslabkit why, notlabkit why-supported. Regenerated: S-3b, S-3c, S-11, S-11e.A repo-wide grep (excluding
node_modules,.git,bin) forregister_session,gate_status,why-supported,labkit known,SESSION_TOOLSandLABKIT_TENANTfinds nothing. A grep forknownas a command name finds comments inscripts/checks/smoke-cli.sh:114,scripts/db/probe-bonsai-replay.sh:16,packages/core-domain/write/work.ts:75,packages/app-cli/views/format.ts:31,tests/cli/views.test.ts:507andtests/domain/survey-after-standing-changes.test.ts:2. None of them runs the command, and they are left alone.knownis also a field of thenowreport.Checks
bun run teston main at 3de1f9a: 1639 pass, 2 skip, 1 fail (1642). On this branch: 1644 pass, 2 skip, 1 fail (1647). The failure is the same on both:packages/app-acp/http.test.ts"the CLI serves --http until SIGTERM…" receives an escape sequence after the URL on stderr whenFORCE_COLORis set in the environment. It is outside this change.bun run check:quick: all 25 passed.check:binary,check:cliandcheck:scenario-docspass.Found, not changed
labkit mcpaccepts--author,--reconstructed-fromand--dateand ignores them. MCP now has no way to setreconstructedFrom. Both belong to attribution.undoof a close,enquiryStatusstill reports the enquiry closed, because it does not filter a retracted closing decision. The survey does filter it. The enquiry list does not filter it either.dumpandrestorestill read--dbthemselves rather than the located record (dump.tswas not touched).console.error.infra/ci/triggers.tfstill lists the rootREADME.mdamong paths that skip a build.🤖 Generated with Claude Code
https://claude.ai/code/session_01FWZHuy7xJbVqxyLipnvwdY