+
+
+
+
+
+Stardate 2026.267: The recall budget jam and its fix — BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
+
+
STARDATE 2026.267
+
The recall budget jam and its fix
+
+
Date
America/Denver
+
Status
Finished
+
Format
Did · Learned · Open · Log
+
+
Recall fell back to its old behaviour on 1,539 of 1,685 live attempts because its budget ledger was full of phantom charges, while real spend was about $0.0014. We traced the cause and fixed the accounting.
+
+
+
01 What we did
+
Recall is the step where JEV, our small judge model, picks which memories go into an agent's context before it answers. Every JEV lane (one job the judge does) has a daily budget. When the budget says no, the lane falls back to its old behaviour. Nothing breaks, but the agent answers without the judge's help.
+
Each lane keeps a ledger: before each call it books an estimated cost, called a reservation.
+
On 2026-09-24 we found recall falling back far too often. Why was a lane that spends almost nothing always out of money? We looked at 1,685 live recall attempts, the lane's ledger, and the timing of each step inside one recall.
+
+
+
02 What we learned
+
The budget was full of phantom charges. 1,539 of the 1,685 attempts, about 91%, fell back because the budget said no. The ledger showed $0.4989 charged against a $0.50 limit. Real spend was about $0.0014.
+
The chain:
+
A recall step has 2.0 seconds in total.
About 0.4 s went on two Inbox lookups of the lane's grant (its signed permission), and about 0.4 s on fetching its key from the vault.
The lane kept 1.0 s in reserve for a retry, and the call itself got only about 0.3 s.
Large requests take 0.4 to 0.5 s, so they timed out with an unknown outcome.
An unknown outcome was booked at a worst case that grew with request size and question count, then doubled. With about 86 questions in a large recall, each booking was 16 to 750 times the real bill.
These bookings, and those for refused calls (HTTP 403), never aged out of what was meant to be a daily window. They piled up until the ledger was full.
+
What surprised us: the lane was not short of money at all. Its own bookkeeping blocked it. The same pessimistic accounting sits in the shared client used by other JEV lanes, including the incoming filter and nightly dreaming. Any lane with real traffic and occasional timeouts will jam the same way.
+
The fix is a corrected client for the lane. A booking is capped by the request's size in bytes, old bookings age out of a true rolling window, and a refused request (a 4xx error) costs nothing. The lesson we keep: check a lane's ledger before blaming the provider.
+
A later readout, from our system map on 2026-09-25, shows recall applying on 590 of 591 prompts over six hours. That is a later reading, not proof that this fix alone caused it.
+
+
+
03 Open questions
+
Other lanes shared the same accounting when we found this. Should they move to the corrected client before real traffic jams them?
The lookups eat into a 2.0 s step. How much of that time can we win back?
The later readout covers six hours. Does it hold over days and under heavy load?
+
+
+
04 Lab book log
+
Found 1,539 of 1,685 live recall attempts falling back on budget.
timing
Grant lookups took about 0.4 s, the key fetch about 0.4 s; the call got about 0.3 s.
root cause
Unknown outcomes were booked at 16 to 750 times the real bill and never aged out.
fix
Corrected client: capped bookings, rolling-window aging, no charge for refused calls.
Later readout from the system map: recall applies on 590 of 591 prompts over six hours.
+
+
+
+
+
+
+
diff --git a/site/assets/stardate/2026-268-clm-side-test.html b/site/assets/stardate/2026-268-clm-side-test.html
new file mode 100644
index 0000000..7517c4c
--- /dev/null
+++ b/site/assets/stardate/2026-268-clm-side-test.html
@@ -0,0 +1,66 @@
+
+
+
+
+
+
+
+
+Stardate 2026.268: The CLM side test, and why we stay on JEV — BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
+
+
STARDATE 2026.268
+
The CLM side test, and why we stay on JEV
+
+
Date
America/Denver
+
Status
Finished
+
Format
Did · Learned · Open · Log
+
+
We tested whether CLM, an 8B model run on one of our own studios, could pick recall context instead of JEV. It could not; it found the needed rows far less often and was fast enough only with a precomputed cache.
+
+
+
01 What we did
+
Recall picks which memory rows go into an agent's context before it answers, and today JEV, our small judge model, does that job. Could CLM, an 8B model run locally on an Apple Silicon studio, do it instead, at least as well and fast enough?
+
We used a frozen holdout of 600 episodes. An episode is one question with a pool of candidate rows; a holdout is a test set kept apart from all tuning. In 300 episodes the needed evidence was present. CLM's input format and a score cutoff were chosen on a separate practice set, and the rule for switching was written before any scoring. CLM ran zero-shot, as it comes, with no training on our data. Its authors say this kind of use needs fine-tuned heads, so a zero-shot result is not the model's ceiling.
+
+
+
02 What we learned
+
Quality: no. Routing picks 8 rows from the pool. How often did those 8 hold every row the answer needs, over the 300 evidence-present episodes?
+
Router
Full cover
Versus the JEV arm
JEV arm's routing (cheap word matching)
230 (76.7%)
baseline
Installed current workflow
192 (64.0%)
−12.7 points
CLM, row text as is
81 (27.0%)
−49.7 points
CLM, row prefixed with its file name
146 (48.7%)
−28.0 points
+
Both CLM gaps have uncertainty ranges entirely below zero, the pre-registered signal to stay on JEV.
+
End to end, CLM's own arm declined every question: on the practice set, no score cutoff beat declining everything. Its 30 most confident answers were 26 critical errors and 4 fully supported.
+
What surprised us: one CLM variant beat a baseline by 6.3 points of fully supported answers, yet a control using 8 random rows matched it in all three seeds. The gain came from the framework's word-matching repair pass, which fired in 440 of 600 episodes behind CLM against 197 behind the word-matching router.
+
Speed: only with a cache. On a shared GPU that was 92 to 99% busy, encoding the query and every row took 3.0 to 4.0 s at the median, against a recall budget of about 1.6 s. With each row's embedding (numbers that stand for its meaning) cached ahead of time, it took 123 to 125 ms. Cold start was 16 s to ready, and peak memory was 20.7 GB.
+
Not a port defect. Our local build matched the maintainers' reference outputs, with the largest gap under 0.001 in cosine terms.
+
+
+
03 Open questions
+
The framework has no recording of JEV's own reranking. A shadow run comparing JEV's top picks with a candidate's on live pools would fill that gap.
The questions name topic terms from the answering row, which favours word matching. Would real, more semantic queries narrow a 28-point gap?
A fine-tuned head of about 19 million parameters might help, in a separate project that never trains on this holdout.
+
+
+
04 Lab book log
+
The source report records no run dates; this entry was filed on this day.
pre registration
Switch rule written before any calibration or holdout scoring.
parity
Local build matched the maintainers' outputs within 0.001 cosine.
dev
Format and cutoff chosen; no cutoff beat declining everything.
holdout
600 episodes scored; CLM routing 49.7 and 28.0 points behind.
latency
60 real pools timed on a shared GPU; 3.0 to 4.0 s median uncached.
control
Random rows matched the best CLM variant.
decision
Stay on JEV.
+
+
+
+
+
+
+
diff --git a/site/assets/stardate/2026-268-deep-dream-day-one.html b/site/assets/stardate/2026-268-deep-dream-day-one.html
new file mode 100644
index 0000000..9c6c91c
--- /dev/null
+++ b/site/assets/stardate/2026-268-deep-dream-day-one.html
@@ -0,0 +1,66 @@
+
+
+
+
+
+
+
+
+Stardate 2026.268: Deep dream, day one — BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
+
+
STARDATE 2026.268
+
Deep dream, day one
+
+
Date
America/Denver
+
Status
Running
+
Format
Did · Learned · Open · Log
+
+
Our first full clean-up pass over the main studio's memory retired 3,133 of 117,506 live memories by 09:36, all reversibly. The near-duplicate wave paused at the day's JEV budget, and we resumed it the same afternoon.
+
+
+
01 What we did
+
Deep dream is the on-demand, full-pass version of our nightly memory clean-up. It looks for duplicates and useless notes, such as tool echoes ("Wall time 1.2 seconds") and old snapshots of a machine's state. Our question: how much can one pass retire, with every change undoable?
+
Nothing is deleted. A note is retired with a retirement marker, a flag that hides it from recall without deleting it. One command restores a whole wave.
+
Three independent code reviews came first, and every required fix was applied. A rehearsal on a copy of the memory applied and undid every stage, byte for byte. The live plan, made at 06:22, covered 117,506 live memories in four waves. In every wave except exact duplicates, JEV, our small judge model, must clear each retirement.
+
+
+
02 What we learned
+
By 09:36, 3,133 memories were retired, 2.7% of the plan.
+
Wave
Candidates
Retired
Why the rest stayed
Machine-state snapshots
13
9
3 kept by JEV; 1 not yet 7 days old
Tool echoes
498
371
90 kept by JEV; 3 held back for privacy; 34 pointed at by another marker
Exact duplicates
535 extra copies
11
4 protected decisions; the rest refused by the planner's compatibility rules
Near-duplicates
9,642 pairs
2,742
6,010 pairs judged so far; 142 contradictions logged, not merged; 103 held by the chain rule
+
For near-duplicates, JEV must say "same fact" or "supersedes" with at least 90% confidence; then the shorter or older copy is retired. The chain rule keeps any note that another marker points to.
+
The checks held. Recall, sampled 40 times after each wave, never returned a retired copy, and found the kept fact about 90% of the time. Every changed memory changed only in its marker and labels. A live undo test on the snapshot wave restored 9, and all 13 notes matched the before-snapshot byte for byte. Spot checks of merges down to a similarity of 0.92 found only true rewordings.
+
It was cheap. JEV calls since 07:05 cost $0.176, for 6,173 calls. Deep dream was set to stop at $0.23 of a rolling 24-hour budget of $0.25 that it shares with the nightly. The near-duplicate wave stopped at the day's JEV budget. The remaining 3,632 pairs were first set for the next morning; instead we raised the rolling budget to $1.00, with an undo, and resumed them at 16:09.
+
What surprised us: exact duplicates, which sound easy, retired only 11 of 535 extra copies. And on most near-duplicate pairs JEV was at least 90% sure the notes were redundant, but split its answer between "same fact" and "newer replaces older". Neither answer reached 90%, so both copies stayed.
+
+
+
03 Open questions
+
Which copy should we keep when JEV splits its answer? One proposal: the newer copy when it holds everything the older says, otherwise the longer.
385 exact-duplicate groups hold the same text saved as different kinds of note. Nothing merges them yet.
The lane re-checks its permission with about 6 Inbox hub calls per pair, which caused 9 safe stops when the hub was slow. Can it check less often, safely?
+
+
+
04 Lab book log
+
review
Three independent code reviews, then a rehearsal on a copy; every stage undone byte for byte.
Live plan made over 117,506 live memories.
Cost window opens; $0.176 over 6,173 JEV calls from here.
3,133 retired, 2.7% of the plan; near-duplicates paused at the day's JEV budget.
Rolling budget raised to $1.00, with an undo; the remaining near-duplicate pairs resumed.
In 96 headless Grok runs, a memory-search tool got the fewest knowledge questions wrong and answered about 3× cheaper and 4× faster. Injecting recall into the prompt made Grok worse.
+
+
+
01 What we did
+
How should Grok reach our shared memory? We compared four setups, called arms. (A hook is a small program that runs at set points in an agent's loop and can add text to its prompt.)
+
A: no memory hooks.
R: recall injected into the prompt, as Claude gets it on the main studio.
M: a read-only memory-search tool.
C: recall plus advice hooks from JEV, our small judge model.
+
The 12 tasks were six fleet-knowledge questions whose answers live in memory, and six small coding tasks with tests. Each ran twice per arm: 96 headless runs (scripted, nobody at the keyboard), same settings throughout. Each run had a fresh private home folder, and grading happened elsewhere, so no agent saw the answer key.
+
+
+
02 What we learned
+
All 48 coding runs passed in every arm; memory neither helped nor hurt. Knowledge questions told a different story (12 runs per arm, 95% ranges in parentheses):
+
Arm
Wrong
Mean cost
Median time
M, search tool
25% (9 to 53%)
$0.119
24 s
A, no hooks
50% (25 to 75%)
$0.386
106 s
C, recall and advice
50% (25 to 75%)
$0.390
276 s
R, recall injected
75% (47 to 91%)
$0.373
113 s
+
The search tool won on cost and speed: about 3× cheaper and about 4× faster, clear against every arm (p ≤ 0.04). On accuracy, only its gap to injected recall is clear (p = 0.04); against no hooks it is suggestive, not proven (p = 0.40).
Injected recall made Grok worse. 8 of its 12 knowledge runs hit the 20-turn limit without answering. The recall block sent Grok hunting through the host to check what it said: 216 tool calls on host paths, against 7 for the search tool.
Advice hooks gave no clear gain. Turns took about 14 s instead of about 6 s, causing 4 of the 5 timeouts, perhaps partly because each test run used a fresh state folder.
Stale memories mislead. Both of M's misses on which machines take no new work came from old machine-state notes treated as current. The search tool returns facts without their dates.
+
What surprised us most: recall meant to help sent Grok checking instead of answering.
+
Caveats. With 12 runs per arm, treat the accuracy order as a strong hint; cost and speed are solid. "No memory" is not knowledge-free: in 10 of 12 no-hook knowledge runs, Grok read the host's own notes.
+
Two lessons for anyone running agents: an agent without a sandbox can read anything its user account can, and a killed run can leave its tool processes running.
+
+
+
03 Open questions
+
Would the search tool, recommended but not yet built, help lane-host agents in production?
Would returning each memory with its date and type stop the stale-note misses?
Can JEV advice calls get fast enough for Grok lanes? Until then we would keep them off.
+
+
+
04 Lab book log
+
setup
Harness built and validated for $3.96: 12 tasks, 4 arms, 2 repeats.
runs
96 headless Grok runs on one lane host; recorded cost $16.30.
cleanup
Stopped ten leftover scans that ran up to 55 minutes after killed runs.
grading
Graded elsewhere; one task regraded by hand, same rule for every arm.
Report written: $20.26 recorded, about $22 to 25 counting five timed-out runs.
+
+
+
+
+
+
+
diff --git a/site/assets/stardate/2026-268-incoming-filter-on.html b/site/assets/stardate/2026-268-incoming-filter-on.html
new file mode 100644
index 0000000..12758b1
--- /dev/null
+++ b/site/assets/stardate/2026-268-incoming-filter-on.html
@@ -0,0 +1,67 @@
+
+
+
+
+
+
+
+
+Stardate 2026.268: The incoming filter, from shadow to on — BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
+
+
STARDATE 2026.268
+
The incoming filter, from shadow to on
+
+
Date
America/Denver
+
Status
Finished
+
Format
Did · Learned · Open · Log
+
+
After 24 hours in shadow mode we read all 116 would-be drops by hand and found no durable facts. We then switched the filter on, and its first real drop was written to an owner-only recovery copy before being dropped.
+
+
+
01 What we did
+
Agents write facts to our shared memory as they work. Most new rows, 80 to 98%, arrive through one capture path, and some of it is noise: echoes of tool output and scraps of telemetry. We wanted JEV, our small judge model, to stop that noise before it is stored, without losing a real fact.
+
So the filter started in shadow mode: it judges every fact but drops nothing. After an independent review we set the rules for leaving shadow:
+
at least 24 hours of shadow after a credential fix, with credential errors near zero;
at least 100 judged candidates;
every would-be drop read by hand, with zero durable facts among them;
drops counted under the live filter's caps;
latency measured over at least 20 captures.
+
Once on, the filter may drop at most half of any one capture, and at most 8 facts an hour across the whole service. It may only drop two kinds of noise: telemetry fragments and echoes of commands that ran.
+
+
+
02 What we learned
+
The shadow day. In 24 hours the filter saw 418 captures and 2,498 candidate facts, and made 1,894 JEV calls. Credential errors were 0.4%.
+
The would-be drops. With both caps applied, it would have dropped 83 facts; the caps spared 33 more. We read all 116 by hand. Every one was an echo of test, CI (automatic build-and-test runs), build or script output, plus one fragment of system time. None was a durable fact.
+
Latency. 95% of captures finished within 45.0 s, against 41.3 s before: 3.7 s more.
+
What surprised us was how uniform the list was: nothing but output echoes and one scrap of system time.
+
Going on. At 13:00 on 2026-09-25 we switched the filter on, and the service restarted in about 4 seconds. At 14:07 came the first real drop, a script error echo. It was written first to an owner-only recovery copy. We verified that the fact was absent from what was stored, and that the receipt links to the copy by hash.
+
Safety nets. If a recovery copy cannot be written and read back, nothing is dropped. Copies are kept for 30 days and purged only by hand, after review. A dropped fact can be restored by storing it again, word for word, in the same scope. The fast rollback is a switch in the service settings and a restart; the full one restores the exact previous files.
+
+
+
03 Open questions
+
Facts that match the capture hook's own drop patterns are kept unjudged, by design, so some echoes still get stored until the deep dream catches them. Should the filter judge those too?
The caps spared 33 would-be drops in shadow. Are they tighter than they need to be?
Will zero durable facts among the drops hold as the work changes?
+
+
+
04 Lab book log
+
Credential fix in place; the 24-hour shadow clock counts from here.
Capture path starts in shadow mode: JEV judges, nothing is dropped.
Revised build installed after a third review; recovery copies forced to disk.
All 116 would-be drops read by hand; zero durable facts.
Switched to on; the service restarted in about 4 seconds.
First real drop, a script error echo, saved first to a recovery copy and verified.
+
+
+
+
+
+
+
diff --git a/site/assets/stardate/2026-268-judge-kit-reviews.html b/site/assets/stardate/2026-268-judge-kit-reviews.html
new file mode 100644
index 0000000..67a3b5f
--- /dev/null
+++ b/site/assets/stardate/2026-268-judge-kit-reviews.html
@@ -0,0 +1,65 @@
+
+
+
+
+
+
+
+
+Stardate 2026.268: The judge kit and its two independent reviews — BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
+
+
STARDATE 2026.268
+
The judge kit and its two independent reviews
+
+
Date
America/Denver
+
Status
Finished
+
Format
Did · Learned · Open · Log
+
+
We added seven quality items around our agents. Two independent code reviews found real defects, and every high and medium finding is now fixed.
+
+
+
01 What we did
+
The judge kit is our set of quality checks around the agents. We added seven items, each with a canary (a small automatic check that runs all the time) in our regression harness.
+
Item
What it does
First result
Outcome grading
Hourly, asks whether each recalled memory was used
0.8% of 4,451 clearly used
Claim checks
Hourly, asks whether a "done" claim has evidence
about 5% of claim turns flagged
Pre-send guard
Checks recipients and holds before a client message; only orchestration-tier models may send
tests 20 of 20
Brief checks
Blocks a task brief with no acceptance check, artifact or stop condition
45% of real briefs lacked acceptance checks
Hook triage
Measures the text our hooks add to prompts
about 18.7 KB per prompt
Drift alarms
Every 15 minutes, compares each lane with its last 24 hours
in a backtest, caught the 09-24 incidents within the hour
Routing scoreboard
Hourly, observed success per model, account and machine
one model, 90% on one studio and 65% on another
+
Six items had passed their own tests, canaries and live samples, but no second reader had seen the code. Our question: would independent reviewers find what our own tests missed?
+
+
+
02 What we learned
+
The first review found real defects. A reviewer on another studio read all the code, ran every suite and wrote 12 probe scripts.
+
Five high-severity and five medium defects: three in the pre-send guard, two in the drift alarms, and partial grading in the hourly checks.
The kind of thing it caught: a lane failing every call never raised an alarm, and turns still in progress were graded on a partial reply and never graded again.
+
All 5 high and 5 medium findings were fixed, and 15 of the 16 low. We also made subagents (helpers started by another agent) draft-only for client messages.
+
The second review, of the fixes, found three more: one high and two medium, in the pre-send guard and the brief check. All are fixed, along with 11 low findings.
+
What surprised us: code that had passed its own tests, canaries and live samples still hid high-severity gaps.
+
+
+
03 Open questions
+
Only 0.8% of injected memories are clearly used. Is recall picking poorly, or is the grader strict?
A tested patch for the Inbox check-in saved 4.7 MB a day (38%) in a replay, and it is now live on the main studio. Will it hold up across the fleet?
+
+
+
04 Lab book log
+
build
Seven items built, each with a canary.
review one
First review of six items: 5 high, 5 medium, 16 low.
fixes
All high and medium fixed, 15 of 16 low; probes rerun clean.
review two
Second review, of the fixes: one high, two medium, 11 low.
refix
Second-review findings fixed; three low items accepted as they are; 100 tests pass; four superseded grants revoked.
Report written: all seven items live, regression 12 of 12.
We found 527 duplicate markers that had quietly switched off because the copy they pointed to was later retired. The fix is installed, the repair has run, and the first protected nightly run is next.
+
+
+
01 What we did
+
When our memory holds two copies of the same fact, the nightly dream keeps one and retires the other with a retirement marker, a flag that hides a note from recall without deleting it. Each marker points to its kept copy, the one that stays visible.
+
On 2026-09-25 we asked a narrow question: do those markers keep working over time? A marker hides its duplicate only while its kept copy is live and unmarked. So we ran a census, a read-only count of every marker and the state of its kept copy, on the main studio.
+
+
+
02 What we learned
+
Markers were quietly switching off. At 13:44 the census found 7,222 marked notes. 527 exact-duplicate markers were inactive, because their kept copy had itself been retired later. Nothing was lost, but each of those duplicates was visible to recall again.
+
503 were exact-on-exact chains, built up since 09-06 at about 26 a night: a later exact-duplicate pass retired an earlier marker's kept copy.
24 came from 4 kept copies that the nightly's near-duplicate step retired.
None came from deep dream. Its chain rule, which keeps any note another marker points to, held.
+
The cause. The nightly's exact-duplicate step picked the winner of each group of identical copies by source quality, then by id. A new identical copy with a smaller id won, and the old kept copy was retired. The retention planner, which proposes what to retire, also skips any note that carries a marker, even an inactive one. So no later run ever fixed a chain.
+
What surprised us: nothing reported it. A marker written by any tool can be undone, silently, by a later run that retires its kept copy.
+
The fix. Three changes:
+
the planner takes a protected list, so a kept copy is never proposed for retirement and always wins its group;
the nightly scans every marker's kept copy before it plans, which took 0.8 s live, and guards its near-duplicate proposals too;
restore can target chosen markers only.
+
A repair restores only the chained markers in the nightly's journal, and the next protected nightly marks them again.
+
Where it stands. An independent review returned accept, with 0 blocking findings. The fix is installed with a one-command undo. The repair restored the 527 markers with 0 other changes to the stored notes, and the census then showed 0 inactive markers.
+
+
+
03 Open questions
+
The first protected nightly run, set for 09-26 03:30, will re-mark up to 500 of the restored duplicates. Will it do so cleanly?
Any new tool that writes retention markers must protect the kept copies of all existing markers. How do we make sure every future writer does?
Dream health should be measured by the count of inactive markers, not just markers applied. What count should raise an alarm?
+
+
+
04 Lab book log
+
Start of the count: 503 exact-on-exact chains built up from here, about 26 a night.
Census on the main studio: 7,222 marked notes, 527 inactive duplicate markers.
cause
Traced to the nightly's exact step letting a newer identical copy win its group.
fix
Planner protects every kept copy; the nightly scans them before each plan.
review
Independent review: accept, with 0 blocking findings.
install
Fix installed with a one-command undo.
repair
527 markers restored with 0 other changes; the census then showed 0 inactive.
State recorded; the first protected nightly is set for 09-26 03:30.
+
+
+
+
+
+
+
diff --git a/site/assets/stardate/2026-268-memory-reach.html b/site/assets/stardate/2026-268-memory-reach.html
new file mode 100644
index 0000000..d10f9d3
--- /dev/null
+++ b/site/assets/stardate/2026-268-memory-reach.html
@@ -0,0 +1,66 @@
+
+
+
+
+
+
+
+
+Stardate 2026.268: Memory rarely reaches the agents that need it — BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
+
+
STARDATE 2026.268
+
Memory rarely reaches the agents that need it
+
+
Date
America/Denver
+
Status
Finished
+
Format
Did · Learned · Open · Log
+
+
On lane hosts, recall delivered something on only 5 to 10% of prompts, and almost nothing for Grok. The facts were in memory, but queries built from the raw prompt missed them.
+
+
+
01 What we did
+
Our agents share one memory. Recall is the step that hands an agent the useful memories before it answers. Before judging whether memory helps, we asked a plainer question: does the right memory reach the agent at all?
+
At about 01:00 on 2026-09-25 we measured three things:
+
Lane hosts, the studios that run our agents' delegated jobs. From the recall hook's log since 09-18 (a hook is a small program that runs at set points in an agent's loop), we counted how many prompts got any memory.
The main studio, where Claude gets recall on almost every prompt, reranked by JEV, our small judge model.
Retrieval. For six fleet-knowledge questions, did recall carry the answering fact, and could a focused search find it?
+
+
+
02 What we learned
+
On lane hosts, recall rarely delivered anything. Across all runtimes it delivered something on about 5 to 10% of prompts.
+
Lane host
Claude prompts with recall
Codex prompts with recall
one studio
63 of 1,205
63 of 641
another
32 of 401
not listed
a third
9 of 90
20 of 85
a fourth
0 of 23
not listed
a fifth
0 of 15
not listed
+
Grok got almost nothing: its recall step before tool calls logged 730 skips on one lane host. The cause is a word-matching and project gate. It keeps a memory only if the prompt shares its words, and drops memories filed under a different project, which is just the name of the agent's working folder.
+
On the main studio, recall arrives but is rarely used. It injects 3 to 5 memories on almost every prompt. Our outcome grader found only about 0.8% clearly used.
+
The facts are there; the queries miss them. Recall on the main studio carried the answering fact for 1 of the 6 knowledge questions, or 2 of 6 with file content added to the query. A focused memory search found the same facts at the top. The fact that one studio takes no new agent work, for example, scored 0.83 to 0.86.
+
What surprised us is why. Prompts often name the key thing only inside a file, and recall at prompt time never sees it. So when memory looks unhelpful, the problem is mostly delivery and query building, not missing knowledge. Our rule now: before judging memory's value in any test, check that recall delivered the relevant fact.
+
+
+
03 Open questions
+
Would a search tool that agents call themselves fix delivery on lane hosts? Our Grok A/B test points that way.
Would running recall again after the agent reads its files close the gap?
Six questions is a small test. How often does recall carry the answer on a larger set?
+
+
+
04 Lab book log
+
Start of the lane-host recall log used for this count.
Counted delivery per lane host and runtime: about 5 to 10% of prompts got anything.
grok check
Grok's recall before tool calls logged 730 skips on one lane host.
main recall
The main studio's recall injects 3 to 5 memories on almost every prompt; about 0.8% clearly used.
retrieval test
Recall carried the answer for 1 of 6 questions, or 2 with file content added.
search test
A focused search found the same facts at the top, scoring 0.83 to 0.86.
+
+
+
+
+
+
+
diff --git a/site/assets/stardate/index.html b/site/assets/stardate/index.html
new file mode 100644
index 0000000..d452b5b
--- /dev/null
+++ b/site/assets/stardate/index.html
@@ -0,0 +1,84 @@
+
+
+
+
+
+
+
+
+Stardate log — the lab book of the BORG Collective
+
+
+
+
+
+
+Skip to content
+
+BORGCOLLECTIVE
+
+
+
+
+
LAB BOOK / THE COLLECTIVE
+
STARDATE LOG.
+
Every experiment the collective runs is written up here, the same way every time: what we did, what we learned, what we still don’t know, and the lab book log.
+
+
Reading a stardate
2026.268 is the year, then the day of the year in America/Denver. Day 268 of 2026 is 25 September 2026.
+
Status
Running means the work is still going. Finished means it reached a result. Superseded means a later entry replaced it.
+
Entries
8 so far, newest first. Numbers come from the collective’s own reports; names, machines and accounts are left out.
We found 527 duplicate markers that had quietly switched off because the copy they pointed to was later retired. The fix is installed, the repair has run, and the first protected nightly run is next.
After 24 hours in shadow mode we read all 116 would-be drops by hand and found no durable facts. We then switched the filter on, and its first real drop was written to an owner-only recovery copy before being dropped.
Our first full clean-up pass over the main studio's memory retired 3,133 of 117,506 live memories by 09:36, all reversibly. The near-duplicate wave paused at the day's JEV budget, and we resumed it the same afternoon.
On lane hosts, recall delivered something on only 5 to 10% of prompts, and almost nothing for Grok. The facts were in memory, but queries built from the raw prompt missed them.
We tested whether CLM, an 8B model run on one of our own studios, could pick recall context instead of JEV. It could not; it found the needed rows far less often and was fast enough only with a precomputed cache.
In 96 headless Grok runs, a memory-search tool got the fewest knowledge questions wrong and answered about 3× cheaper and 4× faster. Injecting recall into the prompt made Grok worse.
Recall fell back to its old behaviour on 1,539 of 1,685 live attempts because its budget ledger was full of phantom charges, while real spend was about $0.0014. We traced the cause and fixed the accounting.