Skip to content

A reopened session shows what was shown live: tool cards from the journal's records, thinking recorded - #577

Closed
danbarua wants to merge 2 commits into
mainfrom
replay-from-journal
Closed

danbarua wants to merge 2 commits into
mainfrom
replay-from-journal

Conversation

@danbarua

Copy link
Copy Markdown
Owner

Problem

A session reopened with session/load does not show what the client was shown live (#543). Two causes:

  1. Replay read the wrong record. Every tool call's raw outcome is in the journal (tool records), but replay() in app-acp/rpc/updates.ts rebuilt cards from the conversation summary (MessageView). It hardcoded kind other, used the bare tool name as the title with no locations, and sent the model-facing result text as rawOutput.
  2. The journal drops thinking. The thinking text streamed to the client is never recorded on the step that produced it. The only thinking stored is the provider's resend payload: empty text for Claude Opus 5.5, nothing at all for openai-chat providers (the Qwen turn in the tool-permutations session). What was shown live is gone.

Rule this PR works to: nothing shown live is lost, and nothing is recomputed or invented on reload.

Change (so far)

  • 3fe2464 Tool cards on reload are built from each call's tool record, through the same functions the live path uses (core-agent/host/tool-display.ts for the announcement, locations and settled outcome; the ACP card code in rpc/updates.ts). A refused, failed or located call now gives the same card live and after reload.
  • view-model-parity.test.ts compares the whole view model live and after reload, less the labkit.dev/reconstructed marker, for a plain answer, a refused call, and a located call plus a failed call.

Not done yet (this PR stays draft until they are)

  • Record the thinking text (and the partial answer of a step that did not finish) the client was shown, on model_settled, for every provider; replay it. Parity test with thinking from an Anthropic and an openai-chat provider.
  • Record each call's kind and locations when shown, instead of recomputing them from today's tool definitions on reload (breaks when a tool changes).
  • Stop inventing an id and failed status for a call with no tool record.
  • Replay permission requests and their answers, and plans (held Replay plan tool results on session load #576 overlaps).
  • The id of the user's prompt block: the client makes it up live (user:0); the three parity tests still fail on this one field only.

Verification

  • bun run check: 1660 pass, 3 fail, the three parity tests above, each on the prompt block id only. Typecheck, format and lint pass.
  • Before the change the same tests failed on ids, kind, locations, title and rawOutput.

Linked: #543, #544, #576

🤖 Generated with Claude Code

https://claude.ai/code/session_01GVjyjkaADPAqLvtnf3xUNs

… live card code

session/load rebuilt each tool card from the conversation summary: title the bare
tool name, kind always "other", no locations, and rawOutput the model-facing
result text instead of what the client was shown. Replay now reads each call's
raw `tool` record and sends it through the same functions the live path uses
(core-agent host/tool-display.ts: announcement, locations, settled outcome; and
the ACP card code in rpc/updates.ts), so a refused, failed or located call shows
the same card live and after a reload.

view-model-parity.test.ts now compares the whole view model live and after
session/load, less the reconstructed marker. It still fails on one field: the
id of the user's own prompt block, which the client makes up live.

Still breaks the rule that a reload shows only what was recorded: kind and
locations are recomputed from the current tool definitions, and a call with no
tool record is shown with a made-up id and status "failed". Next commits store
those facts when they are shown, along with the thinking text the client saw.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVjyjkaADPAqLvtnf3xUNs
@danbarua

danbarua commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner Author

Early notes on the first commit, aimed at the unchecked items, since they are cheaper to settle before the next commits:

  • The direction holds. Live and reload now share toolAnnouncement, settledTool and showTool, so a card has one definition. The refused/failed/located parity test compares whole cards less _meta, which is the right bar.
  • Recording what was shown should stay out of what the model reads. The partial answer of a step that did not finish, and thinking text, are things the client saw. The model's conversation is a fold over the journal (effectiveToolResult, projectConversation). If these land as new record fields, check that the fold ignores them, so a reopened session does not start showing the model a cut-off answer it never received as history. A test that the provider request after reload is unchanged by the new fields would pin it.
  • A journal format change. Recording kind and locations when shown (item 2) and thinking (item 1) adds fields to tool and model_settled records. Say in the PR whether older journals without them still load, and what they replay as (today's recomputed values, or nothing).
  • Until item 2 lands, reload runs the live tool's code. recordedCall calls tool.parseInput(call.args) and tool.locations(...) from today's definitions for every located call. For the workspace tools that is schema parsing only. The only written no-I/O promise for reload covers diff baselines (protocol-reference.md:137-139), so running tool code on reload is a new dependency either way: a tool whose parseInput or locations does I/O, or has changed since the session was recorded, now affects what a reopened session shows.
  • Plans (Replay plan tool results on session load #576): the notes on A reloaded session shows tool calls differently from the live session #543 still apply. Live, the plan is published before the call completes (protocol-reference.md:216), and two docs say plans are not replayed (protocol-reference.md:217-218, docs/agent/vscode-acp.md:196-197).
  • The terms-of-reference sample transcript is missing its last tool update and its commands and info updates #544: as noted on A reloaded session shows tool calls differently from the live session #543, the missing session_info_update and available_commands_update in that transcript are sent after replay and before the load response, so this change may not explain them.

… was shown

Each step's record carries what the client was shown of its stream (thinking text,
and the answer text of a step that did not finish), and replay sends it. Not
finished and not checked: kept so the approach is on record. The agent core is
being rewritten, and this branch will not be merged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVjyjkaADPAqLvtnf3xUNs
@github-actions github-actions Bot added the area:agent The agent and its hosts: core-agent, app-acp, app-vscode, acp-fake, acp-scenarios, acp-transcripts label Oct 1, 2026
@danbarua

danbarua commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Closing without merging: the agent core this changes is being rewritten. The branch keeps both commits, including the unfinished work on recording thinking, for reference.

What this PR found still holds. These are requirements for the new core: a fact is recorded when it happens, never recomputed when the session is drawn again.

  • A reopened session shows exactly what was shown live, drawn from what the journal recorded, never rebuilt from the conversation summary (MessageView) or from the model-facing result text.
  • The thinking text the client was shown is recorded, for every provider. Today only the provider's resend payload is kept: empty text for Anthropic's models, nothing at all for openai-chat providers. The Qwen session d950ffb6-71f9-4b48-b082-9098a15279f8 (thinking: high) reopens with no thinking block for this reason.
  • The answer text streamed by a step that did not finish (failed or cancelled) is recorded too.
  • A tool call's kind, title and locations are recorded when they are shown, not recomputed on reload from today's tool definitions, which breaks when a tool changes.
  • A call with no record is not given an invented id or a failed status.
  • Permission requests and their answers, and plans, are replayed (overlaps Replay plan tool results on session load #576).
  • The user's prompt keeps one id, live and on reload, instead of one the client makes up (user:0).
  • The test is parity: the whole view model, live and after reopening, compared field by field (app-acp/view-model-parity.test.ts here).

The ACP side of replay is set out in docs/agent/acp-language.md: session/load replays the entire conversation with the same update kinds a live turn produces, and answers only after the last of them.

Linked: #543, #544, #576

@danbarua danbarua closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:agent The agent and its hosts: core-agent, app-acp, app-vscode, acp-fake, acp-scenarios, acp-transcripts

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant