Usage is very high #18
Description
Activity
Plan
Cut per-run token usage: trim the harness prompt, stop the CLI loading ambient operator context, cap per-run spend, and report cache reads separately from fresh input
Context
Issue #18 reports single sessions totalling ~1M tokens "up and down", causing Claude rate limits, and asks that we prompt the model with only the minimal data it needs.
Auditing the code turns up two separate things, and both need addressing:
- The 1M number is mostly cache reads.
claude.Result.TokensIn()(internal/claude/runner.go:110) sumsinput_tokens + cache_creation_input_tokens + cache_read_input_tokensinto one figure, which is stored asruns.tokens_inand rendered by the web console as plain "Tokens in" (internal/web/assets/app.js:649,685) and by Discord (internal/discord/notifier.go:269). On an agentic run the whole conversation prefix is re-read from cache on every turn, so a 40-turn run with a 25K prefix reports ~1M input tokens while only tens of thousands were ever processed fresh. Cache reads are billed at a fraction of fresh input and are not what drives a rate limit. The number is real but it is not the number the operator thinks it is, and today there is no way to see the split. - There is genuine, fixable bloat, in two places: what the harness puts into the prompt, and what the Claude Code CLI loads into every turn's system prompt because we never tell it not to. Neither is 1M tokens on its own, but both are paid on every turn of every run.
The intended outcome: the same work, with a materially smaller context re-sent per turn, a hard per-run spend ceiling the daemon can enforce, and a dashboard that distinguishes cache traffic from fresh input so "usage is very high" can be judged from data rather than a conflated total.
Approach
Part A — Report cache reads separately from fresh input
Nothing here changes what is sent to Claude; it changes what the operator sees, and it is the part that actually answers "is 1M tokens a problem?".
internal/claude/runner.go: keepTokensIn()as the total (so existingruns.tokens_inrows keep their meaning) and add three accessors alongside it, mirroring the existingTokensOut()style:FreshTokensIn()(Usage.InputTokens),CacheWriteTokens()(CacheCreationInputTokens),CacheReadTokens()(CacheReadInputTokens). TheUsagestruct already parses all four fields.internal/store/store.go: append a third entry to themigrationsslice — the file already establishes this pattern at line 285 (ALTER TABLE runs ADD COLUMN kind ...) and the migration loop bumpsPRAGMA user_versionper entry, so existing databases upgrade in place:Add the two fields toALTER TABLE runs ADD COLUMN tokens_cache_read INTEGER NOT NULL DEFAULT 0; ALTER TABLE runs ADD COLUMN tokens_cache_write INTEGER NOT NULL DEFAULT 0;
store.Run, to therunColumnsconst, and toscanRun.RecordUsagealready takes seven positional arguments; adding two more makes it unreadable. Replace the tail arguments with a smallstore.RunUsage{ModelID, SessionID, CostUSD, TokensIn, TokensOut, CacheRead, CacheWrite, Turns}value and update the four call sites (internal/orchestrator/loop.go:660,:740,internal/orchestrator/prcomments.go:385,:435). Keep the existing "empty model/session does not overwrite"CASE WHENsemantics — the pre-record call atloop.go:660depends on it.- Surface the split:
internal/web/assets/app.js— runs table cell (line 649) and detail grid (lines 685-686): render1.0M in (18K new · 24K written · 950K cached) / 42K out, reusing the existingfmtTokenshelper (line 132). No CSS or server-side change is needed:/runsalready marshals the wholestore.Run.internal/discord/notifier.go:269— same breakdown in the Tokens embed field.internal/orchestrator/loop.go:794— extend theclaude_doneevent detail (and the equivalent inprcomments.go:479) withfresh_in=… cached_in=… out=…so the per-run audit trail carries it.
Part B — Trim what the harness puts in the prompt
All in
internal/orchestrator/prompt.go. The current worst case forissueContextis ~36K chars (12K body + 12 × 2K comments) andimplementTaskPromptappends the approved plan untruncated on top of that;prCommentTaskPromptis unbounded in the number of comments it renders.- Stop feeding the harness's own comments back to the model.
issueContext(line 190) renders every comment, including the plan comment, the PR announcement, and every failure comment the daemon itself wrote. Filter with the existingisAgentCommenthelper (phase.go:37) before themaxCommentsIncluwindow is applied. This is the single largest and safest win: on a re-plan or implement run the plan is currently sent twice (truncated inside the discussion, in full under "Approved plan"), and failure comments are pure noise. The re-plan path's wording ("see the newest comment in the discussion above") still holds — reviewer feedback is a human comment and survives the filter. Optionally also drop bareimplementapprovals viaisApproval(phase.go:51); they carry no information the "Approved plan" header does not already state. - Tighten the constants (lines 11-15), which are the only knobs the truncation uses:
maxBodyChars 12000 → 6000,maxCommentChars 2000 → 1200,maxCommentsInclu 12 → 6. Every one already routes through the existingtruncatehelper, which appendstruncationSuffix— andextractPlan(phase.go:134) keys off that suffix to refuse a truncated plan, so the plan cap below must stay well clear of any real plan length. - Cap the plan. Add
maxPlanChars = 20000and applytruncateto the plan inimplementTaskPrompt(line 265) and topreviousPlaninplanTaskPrompt(line 244). 20K chars is roughly 5K tokens and comfortably above a normal plan, soTestImplementTaskPromptCarriesTheApprovedPlanVerbatimkeeps passing; the cap only bites on a runaway plan. - Bound
prCommentTaskPrompt(line 140), which today has no limit on comment count: addmaxPRCommentsInclu = 10(keep the newest, state how many were dropped, mirroring theissueContext"showing the last N of M" line), give diff hunks their own smallermaxDiffHunkChars = 800, and cap review summaries at 3.pendingMentions(prcomments.go:106) can legitimately return dozens of inline comments after a big review pass.
Note that the system prompts themselves (
systemPrompt,planSystemPrompt,prCommentSystemPrompt) are ~1.5K chars each and are not worth touching — they are the cheapest part of the prompt and each line in them is load-bearing for a safety rule.Part C — Stop the CLI loading the operator's ambient context
This is the lever with the largest per-turn effect and it is currently entirely unmanaged. The daemon shells out to
claudewith the operator's own environment (internal/claude/runner.go:198-222), so every run inherits their user-level MCP servers, plugins, skills, and settings. MCP tool schemas and the skills listing sit in the system prompt on every turn — this is exactly the "prompting the AI with data it does not need" the issue describes, and it is invisible in the harness's own code.In
internal/claude/runner.go, extendOptionsand theargsconstruction (after the existing--permission-modeblock, beforeopts.ExtraArgsso an operator can still override):Flag Why Default --strict-mcp-configWith no --mcp-config, this loads zero MCP servers. The harness's runs need none — they get the built-in file/bash tools.on --disable-slash-commandsDrops the skills listing from the system prompt; the harness drives the model from stdin and never types a slash command. on --exclude-dynamic-system-prompt-sectionsMoves per-machine sections (cwd, env, git status) out of the system prompt, improving prompt-cache reuse across runs. Applies here because we use --append-system-prompt, not--system-prompt.on --autocompact <tokens>Caps how large the conversation grows before it is compacted. This is what bounds the per-turn cache-read cost that produced the 1M figure. Accepts 100k–1M. 200000 --setting-sources <list>Controls whether user/project/local settings load at all. leave unset — see Risks Expose these as first-class
config.ClaudeConfigfields (StrictMCPConfig,DisableSlashCommands,ExcludeDynamicSystemPrompt,AutocompactTokens,SettingSources) rather than asking operators to know the flags, with the defaults above set inconfig.Default()(internal/config/config.go:244).config.MigrateisDefault()-driven and generic, so existingconfig.jsonfiles pick the new fields up through-migrate-configwith no change tointernal/config/migrate.go.Validate
AutocompactTokensinConfig.Validate(): zero means "don't pass the flag", otherwise it must be within 100000–1000000, matching what the CLI accepts.Part D — A hard per-run spend ceiling
The README states as a principle that there is "no cost budget", and the gate only reacts after Claude reports a limit. The CLI offers
--max-budget-usd <amount>(print mode only), which stops a run at a dollar figure instead of letting a pathological run consume an hour of usage before therun.timeoutfires.- Add
Options.MaxBudgetUSD float64tointernal/claude/runner.goand pass--max-budget-usdwhen it is > 0. - Add
run.max_budget_usdtoconfig.RunConfigand set it fromexecute(loop.go:700) andexecutePRComments(prcomments.go:403). - Default:
0(off), so behaviour is unchanged unless the operator opts in.config.example.jsonshould ship a non-zero suggestion in the documented reference so the knob is discoverable. - A budget-terminated run surfaces through the existing failure path:
diagnose(runner.go:293) already renderssubtype,terminal_reasonand the CLI's own message into the error, so the run is recorded as failed with an explanatory reason and backs off normally.
There is no
--max-turnsin the installed CLI (2.1.x) — do not reach for it.Files touched
internal/claude/runner.go— new usage accessors; newOptionsfields; new args.internal/orchestrator/prompt.go— comment filtering, tightened constants, plan and PR-comment caps.internal/orchestrator/loop.go,internal/orchestrator/prcomments.go—RecordUsagecall sites, new claudeOptions, richerclaude_doneevents.internal/store/store.go— third migration entry, twoRunfields,runColumns,scanRun,RecordUsagesignature.internal/config/config.go— newClaudeConfigandRunConfigfields, defaults, validation.internal/web/assets/app.js,internal/discord/notifier.go— token breakdown display.config.example.json(also the embedded default viaembedded.go);models.jsonuntouched.README.md— "How Claude Code is invoked" (line 319, theclaudeinvocation block is verbatim and must be updated), the "Prompts" section's description of what goes into the prompt, the configuration reference, and the "Usage-limit-aware, not schedule-aware" principle at line 86, which no longer holds literally oncerun.max_budget_usdexists.- Tests:
internal/orchestrator/prompt_test.go(add cases for agent-comment filtering and the PR-comment cap),internal/store/store_test.go(a migration test in the shape ofTestMigrationAddsKindAndPRCommentTasks),internal/claude/runner_test.go(assert the new flags appear, in the style of the existing arg assertions).
Verification
go build ./... && go test ./...— the store migration test and the prompt tests are the ones that should visibly change.- Prompt size, measured rather than assumed: add a temporary test (or a
-runone-off) that buildsimplementTaskPromptfrom a fixture issue with a long body, 20 comments (half of them agent comments) and a long plan, and printslen()before and after. Expect roughly a 2-3× reduction. ./coding-agent-loop -dry-run -once— exercises discovery, phase decision and model selection without spending anything; confirms the config additions load and validate.- A real single run against a scratch issue:
./coding-agent-loop -no-mutate -once(runs Claude for real, pushes nothing). Then inspect the transcript:The first line confirmsjq -c 'select(.type=="system") | {tools: (.tools|length), mcp: (.mcp_servers//[]|length)}' ~/.agent-loop/logs/<run>.jsonl | head -1 jq 'select(.type=="result") | .usage' ~/.agent-loop/logs/<run>.jsonl
--strict-mcp-configtook effect (no MCP servers); the second gives the four-way usage split to compare against a pre-change run on the same issue. - Open
/ui, check the run row and detail page show thein (new · written · cached) / outbreakdown, and confirm a Discord notification (if enabled) carries the same.
Risks and decisions for the reviewer
--setting-sourcesis deliberately left unset. Restricting it would cut the most ambient context, but user-level settings can carry auth helpers, permission rules and env that abypassPermissionsdaemon run may depend on. Turning it off could break runs on the operator's machine in a way that is hard to diagnose. It is exposed as config, defaulted to unchanged behaviour. Decide whether to default it toprojectinstead — that is the aggressive choice and would need a manual run to confirm auth still works.--disable-slash-commandsalso disables skills. If a target repository ships.claude/skillsthat the agent is expected to use, this would silently remove them. The harness has never invoked a skill explicitly, so this is judged safe, but it is a behaviour change worth a conscious yes.--autocompact 200000changes behaviour on long runs: the model's earlier context gets summarized rather than carried verbatim. This is the trade the issue is asking for (fewer tokens per turn) but it can cost quality on a long implement run. 200K is a conservative starting point — the CLI's own default is higher.- Redefining
tokens_inwas considered and rejected. Makingtokens_inmean "fresh input only" would give a more honest headline number, but silently changes the meaning of every existing row. The plan keeps the total and adds the breakdown instead. Say so if you would rather have the cleaner semantics and accept the discontinuity. - Retry behaviour is deliberately untouched.
retryDelay(loop.go:384) caps back-off atretry_backoff_max(24h default) and never gives up, so a permanently-broken issue costs a full plan-or-implement run every day forever. Capping that would contradict the documented "the trigger label is the only thing that decides whether it is worked" principle, so it is flagged rather than changed. If reducing usage matters more than that principle, the cheap version is: after N consecutive failures, drop the trigger label and say so in the failure comment. - The plan/implement split inherently pays for exploration twice — the implement run cannot see the plan run's session (
--no-session-persistence, and a human gate sits between them). Nothing in this plan changes that; a shared-session design would be a much larger change and would weaken the "the plan a human approved is the one on the issue" guarantee.
Reply with exactly
implementto approve this plan and start the change. Reply with anything else and the plan will be revised to address it.coding-agent-loop run
ac782780-2909-4afc-a4ef-a8d3f1dcc03b, modelclaude-opus-5, cost $2.8418- The 1M number is mostly cache reads.
- addedagent-plannedManaged by coding-agent-loopManaged by coding-agent-loopand removed
on Aug 29, 2026 Every solution doesn't require you extend the config.json. This tool needs to be slim and great at what its intended for out of the box. We don't need a configuration tweak for every little fix. Revise your plan such that all of this is baked in my default - no configuration overrides.
- added and removedagent-plannedManaged by coding-agent-loopManaged by coding-agent-loop
on Aug 30, 2026 Plan
Cut per-run token usage with baked-in defaults: trim the harness prompt, stop the CLI loading ambient operator context, and split cache reads out of the reported "tokens in"
What changed since the previous plan
The reviewer's feedback was: "Every solution doesn't require you extend the config.json. This tool needs to be slim and great at what its intended for out of the box... all of this is baked in by default — no configuration overrides."
Accordingly:
- No new
config.jsonfields. The previous plan added fiveclaude.*fields and onerun.*field. All are gone.internal/config/config.go,config.example.json,embedded.goandinternal/config/migrate.goare untouched, and no-migrate-configstep is required to get the benefit. - No new
claude.Optionsfields either. The CLI flags are appended unconditionally insideRunner.Run, next to the--print/--no-session-persistenceflags that are already hardcoded there. That is the existing precedent in the file for "this is how the harness invokes Claude, full stop." - The per-run dollar budget (
--max-budget-usd) is dropped entirely. It only made sense as a number an operator picks; a hardcoded dollar figure would either be so high it never fires or would kill legitimate long runs with a confusing error, and on an OAuth subscription the reported cost is imputed rather than billed.run.timeout(45m) already bounds a run in wall-clock, and--autocompactbelow bounds what each turn costs. Dropping it also means the README's "no time-of-day window and no cost budget" principle (line 86) stays true and needs no rewrite. - Everything else — the prompt trimming and the token-reporting split — was already constants-and-code with no configuration surface, and survives as-is.
The one operator escape hatch that remains is the one that already exists:
claude.extra_argsis still appended last, so--mcp-config …can be added back by anyone who genuinely needs an MCP server in agent runs.Context
Issue #18 reports single sessions totalling ~1M tokens "up and down", hitting Claude rate limits, and asks that the model be prompted with only the minimal data it needs.
The audit turns up two separate things:
- The 1M number is mostly cache reads.
claude.Result.TokensIn()(internal/claude/runner.go:110) sumsinput_tokens + cache_creation_input_tokens + cache_read_input_tokensinto one figure, stored asruns.tokens_inand rendered as plain "Tokens in" by the web console (internal/web/assets/app.js:649,685) and Discord (internal/discord/notifier.go:269). On an agentic run the whole conversation prefix is re-read from cache every turn, so a 40-turn run with a 25K prefix reports ~1M input tokens while only tens of thousands were ever processed fresh. Cache reads are billed at a fraction of fresh input and are not what drives a rate limit. The number is real, but it is not the number the operator thinks it is, and today there is no way to see the split. - There is genuine, fixable bloat, in two places: what the harness puts in the prompt, and what the Claude Code CLI loads into every turn's system prompt because we never tell it not to. Neither is 1M tokens on its own, but both are paid on every turn of every run.
Intended outcome: the same work, with a materially smaller context re-sent per turn, and a dashboard that distinguishes cache traffic from fresh input so "usage is very high" can be judged from data. All of it on by default, with nothing to configure.
Part A — Report cache reads separately from fresh input
Changes what the operator sees, not what is sent to Claude. It is the part that actually answers "is 1M tokens a problem?".
internal/claude/runner.go— keepTokensIn()as the total (so existingruns.tokens_inrows keep their meaning) and add three accessors alongside it in the style of the existingTokensOut():FreshTokensIn()(Usage.InputTokens),CacheWriteTokens()(CacheCreationInputTokens),CacheReadTokens()(CacheReadInputTokens). TheUsagestruct already parses all four fields.internal/store/store.go— append a third entry to themigrationsslice. The file already establishes this pattern (entry 2,ALTER TABLE runs ADD COLUMN kind …, line 285) and the migration loop bumpsPRAGMA user_versionper entry, so existing databases upgrade in place rather than being deleted:ALTER TABLE runs ADD COLUMN tokens_cache_read INTEGER NOT NULL DEFAULT 0; ALTER TABLE runs ADD COLUMN tokens_cache_write INTEGER NOT NULL DEFAULT 0;
Add the two fields to
store.Run(line 83), to therunColumnsconst (line 579), and toscanRun(line 582) in the same positions.RecordUsagesignature — it already takes seven positional args (store.go:453); two more makes it unreadable. Replace the tail with astore.RunUsage{ModelID, SessionID, CostUSD, TokensIn, TokensOut, CacheRead, CacheWrite, Turns}value and update the four call sites:internal/orchestrator/loop.go:660and:740,internal/orchestrator/prcomments.go:385and:435. Keep the existing "empty model/session does not overwrite"CASE WHENsemantics — the pre-record call atloop.go:660depends on it.Surface the split:
internal/web/assets/app.js— runs table cell (line 649) and detail grid (lines 685-686): render1.0M in (18k new · 24k written · 950k cached) / 42k out, reusing the existingfmtTokenshelper (line 132). No server-side or CSS change:/runsalready marshals the wholestore.Run.internal/discord/notifier.go:269— same breakdown in the Tokens embed field.internal/orchestrator/loop.go:794andprcomments.go:479— extend theclaude_doneevent detail withfresh_in=… cached_in=… out=…so the per-run audit trail carries it.
Part B — Trim what the harness puts in the prompt
All in
internal/orchestrator/prompt.go. Worst case today forissueContextis ~36K chars (12K body + 12 × 2K comments), andimplementTaskPromptappends the approved plan untruncated on top of that;prCommentTaskPrompthas no cap at all on comment count.-
Stop feeding the harness's own comments back to the model.
issueContext(line 190) renders every comment, including the plan comment, the PR announcement, and every failure comment the daemon itself wrote. Filter with the existingisAgentCommenthelper (phase.go:37) — and also drop bare approvals viaisApproval(phase.go:51), which carry no information the "Approved plan" header doesn't already state — before themaxCommentsIncluwindow is applied, and report "showing the last N of M" against the filtered count.This is the single largest and safest win. On a re-plan or implement run the plan is currently sent twice (truncated inside the discussion, in full under "Approved plan"), and failure comments are pure noise. The re-plan wording ("see the newest comment in the discussion above") still holds: reviewer feedback is a human comment and survives the filter.
Edge case to handle: when filtering leaves zero comments (e.g. the first implement run, whose only comments are the plan and the word
implement), omit the### Discussionheader entirely rather than emitting an empty section. -
Tighten the constants (lines 11-15), which are the only knobs the truncation uses:
maxBodyChars 12000 → 6000,maxCommentChars 2000 → 1200,maxCommentsInclu 12 → 6. Everything already routes through the existingtruncatehelper. -
Cap the plan. Add
maxPlanChars = 20000and applytruncateto the plan inimplementTaskPrompt(line 265) and topreviousPlaninplanTaskPrompt(line 244). Note the coupling:truncateappendstruncationSuffix, andextractPlan(phase.go:134) deliberately refuses to recover a plan ending in that suffix — 20K chars (~5K tokens) is comfortably above any real plan, so the cap only bites on a runaway one andTestImplementTaskPromptCarriesTheApprovedPlanVerbatimkeeps passing. -
Bound
prCommentTaskPrompt(line 140), which today renders every commentpendingMentions(prcomments.go:106) returns — that can be dozens after a big review pass. AddmaxPRCommentsInclu = 10(keep the newest, state how many were dropped, mirroring theissueContextline), give diff hunks their own smallermaxDiffHunkChars = 800instead of reusingmaxCommentChars, and cap review summaries at 3.TestPRCommentTaskPromptIncludesEveryCommentAndDiffHunkuses two comments and one review, so it is unaffected.
Not touched:
systemPrompt,planSystemPrompt,prCommentSystemPromptare ~1.5K chars each, are the cheapest part of the prompt, and each line is load-bearing for a safety rule.
Part C — Stop the CLI loading the operator's ambient context
The lever with the largest per-turn effect, and currently entirely unmanaged. The daemon shells out to
claudewith the operator's own environment (internal/claude/runner.go:198-222), so every run inherits their user-level MCP servers, plugins and skills. MCP tool schemas and the skills listing sit in the system prompt on every turn — exactly the "prompting the AI with data it does not need" the issue describes, and invisible from inside the harness's own code.In
internal/claude/runner.go, add these to the hardcodedargsliteral (line ~178), alongside--print/--output-format/--verbose/--no-session-persistence, i.e. beforeopts.ExtraArgsis appended:Flag Why Verified on 2.1.240 --strict-mcp-configWith no --mcp-config, loads zero MCP servers. Harness runs need none — they use the built-in file/bash tools. Cuts every inherited MCP server's tool schemas out of the system prompt.✅ accepted --disable-slash-commands"Disable all skills" — drops the skills listing from the system prompt. The harness drives the model from stdin and never types a slash command. ✅ accepted --exclude-dynamic-system-prompt-sectionsMoves per-machine sections (cwd, env, git status) out of the system prompt into the first user message. This is a prompt-cache-reuse win, not a size reduction — be honest about that in the commit message. Applies here because we use --append-system-prompt, not--system-prompt.✅ accepted --autocompact 200000Caps how large the conversation grows before compaction. This is what bounds the per-turn cache-read cost that produced the 1M figure. ✅ accepted (CLI validates the range: bogusis rejected with "must be 'auto', or between 100k and 1M")Deliberately not passed:
--setting-sources(see Risks),--safe-mode,--bare,--tools.Since these are literals rather than
Optionsfields, no validation code is needed — the CLI validates--autocompactitself at parse time, before any API call.
Approaches considered and rejected for Part C
--safe-modelooks like the perfect "slim by default, one flag" answer: it disables CLAUDE.md, skills, plugins, hooks, MCP servers, custom commands, agents and output styles in one go, while leaving auth, model selection, built-in tools and permissions working. It is rejected because it also disables the target repository'sCLAUDE.md, which is genuine project context the harness's own system prompt asks the agent to honour ("Follow the conventions already present in the repository"). The four targeted flags get most of the saving without giving that up.--bareis rejected outright: it forces Anthropic auth toANTHROPIC_API_KEY/apiKeyHelperand never reads OAuth or the keychain. The daemon authenticates via~/.claude/.credentials.json, so this would break every run.--tools <subset>would shave the built-in tool definitions, but the agent legitimately needs the full file/bash/search set and losing web access silently would be a real capability regression. Not worth the tokens.
Files touched
internal/claude/runner.go— three new usage accessors; four new hardcoded args.internal/orchestrator/prompt.go— agent-comment filtering, tightened constants, plan cap, PR-comment caps.internal/orchestrator/loop.go,internal/orchestrator/prcomments.go—RecordUsagecall sites, richerclaude_doneevents.internal/store/store.go— third migration entry, twoRunfields,runColumns,scanRun,RunUsagestruct +RecordUsage.internal/web/assets/app.js,internal/discord/notifier.go— token breakdown display.README.md— theclaude …invocation block at line 327 is verbatim and must gain the four flags plus a sentence on why; the Prompts section (line ~312) describesissueContexttruncation and should mention that harness-authored comments are filtered out. No change to the Configuration reference, the principles list, ormodels.json.- Not touched:
internal/config/config.go,internal/config/migrate.go,config.example.json,embedded.go. - Tests:
internal/orchestrator/prompt_test.go(agent-comment filtering; the empty-discussion edge case; the PR-comment cap),internal/store/store_test.go(migration test shaped like the existingTestMigrationAddsKindAndPRCommentTasks; the two existingRecordUsagecall sites at lines 154 and 444 need updating for the new signature),internal/claude/runner_test.go(assert the four flags appear in the invocation).
Verification
go build ./... && go test ./...— the store migration test and the prompt tests are the ones that should visibly change.- Measure the prompt, don't assume it. Add a test in
prompt_test.gothat buildsimplementTaskPromptfrom a fixture issue with a long body, 20 comments (half of them carryingmarkerPrefix) and a long plan, and assertslen(p)is under a fixed ceiling. Expect roughly a 2-3× reduction versus today; keeping it as a real assertion stops the constants drifting back up. ./coding-agent-loop -dry-run -once— exercises discovery, phase decision and model selection without spending anything.- A real single run against a scratch issue:
./coding-agent-loop -no-mutate -once(runs Claude for real, pushes nothing), then inspect the transcript:The first line confirmsjq -c 'select(.type=="system") | {tools: (.tools|length), mcp: (.mcp_servers//[]|length)}' ~/.agent-loop/logs/<run>.jsonl | head -1 jq 'select(.type=="result") | .usage' ~/.agent-loop/logs/<run>.jsonl
--strict-mcp-configtook effect (zero MCP servers); the second gives the four-way usage split to compare against a pre-change run on the same issue. - Open
/ui: confirm the run row and detail page show thein (new · written · cached) / outbreakdown, and that an existing run recorded before the migration renders0 new · 0 written · 0 cachedrather than breaking.
Risks and decisions for the reviewer
--disable-slash-commandsalso disables skills. If a target repository ships.claude/skillsthe agent is expected to use, this silently removes them. The harness has never invoked a skill explicitly, so this is judged safe — but with no config knob it is now unconditional, so it is worth a conscious yes.- These flags are unconditional and boolean flags cannot be un-set from
extra_args.--autocompactis a value flag and a second occurrence inextra_argswill very likely win (commander is last-wins), but--strict-mcp-configand--disable-slash-commandscannot be turned back off. An operator who needs MCP in agent runs adds--mcp-config <file>toextra_args, which--strict-mcp-configthen honours. This is the accepted cost of "baked in, no overrides." --autocompact 200000changes behaviour on long runs: earlier context gets summarized rather than carried verbatim. That is the trade the issue asks for, but it can cost quality on a long implement run. 200K is conservative relative to the CLI's own default; it is a one-token edit inrunner.goif it proves too tight.--setting-sourcesis deliberately not passed. Restricting it would cut the most ambient context of all, but user-level settings can carry auth helpers, permission rules and env that abypassPermissionsdaemon run may depend on. With no config escape hatch, turning it off would be an unconditional risk of breaking every run on the operator's machine in a hard-to-diagnose way. Left at CLI default.- Filtering agent comments also hides prior failure comments from the model. A retry no longer sees why the last attempt failed. I judge that an improvement (a clean retry beats one anchored on a stale failure), but it is a real behaviour change.
- Redefining
tokens_inwas considered and rejected. Making it mean "fresh input only" would give a more honest headline number but silently changes the meaning of every existing row. The plan keeps the total and adds the breakdown. Say so if you would rather have the cleaner semantics and accept the discontinuity. - Retry behaviour is deliberately untouched.
retryDelay(loop.go:384) caps back-off atretry_backoff_max(24h) and never gives up, so a permanently-broken issue costs a full run every day forever. Capping that would contradict the documented "the trigger label is the only thing that decides whether it is worked" principle, so it is flagged rather than changed. If usage matters more than that principle, the baked-in version is: after N consecutive failures, drop the trigger label and say so in the failure comment. - The plan/implement split inherently pays for exploration twice — the implement run cannot see the plan run's session (
--no-session-persistence, and a human gate sits between them). Nothing here changes that; a shared-session design would be much larger and would weaken the "the plan a human approved is the one on the issue" guarantee. - Minor, pre-existing:
truncate(prompt.go:288) slices by bytes and can split a UTF-8 rune. Harmless today and out of scope, but the constants being lowered makes it fire more often; afor !utf8.ValidString(...)back-off is a two-line fix if you want it folded in.
Reply with exactly
implementto approve this plan and start the change. Reply with anything else and the plan will be revised to address it.coding-agent-loop run
5a7f4f64-e22a-45a8-9379-84bfd3c1a05b, modelclaude-opus-5, cost $1.5158- No new
- addedagent-plannedManaged by coding-agent-loopManaged by coding-agent-loopand removed
on Aug 30, 2026 implement
Opened a draft pull request for this issue: #19
Tests failed (
make test) — see the PR for output.Comment
implementagain if you want another attempt at this issue.coding-agent-loop run
b328b7bb-cc7a-451c-b698-00fbcad54205- added and removedagent-plannedManaged by coding-agent-loopManaged by coding-agent-loop
on Aug 30, 2026 - added a commit that references this issue
on Aug 30, 2026
Ensure that we are only prompting the AI with the minimal data it needs. Upon auditing i'm finding a single session taking upwards of 1 millions tokens collectively between tokens up and tokens down. This is way too much and its causing claude rate limits.