research(experiments): the 08-13 restatement restated on v0.13.3 (#737, #738) - #739
Open
eaitbrahim wants to merge 3 commits into
Open
research(experiments): the 08-13 restatement restated on v0.13.3 (#737, #738)#739eaitbrahim wants to merge 3 commits into
eaitbrahim wants to merge 3 commits into
Conversation
Its verdict holds and its numbers do not, and the reason is not the one this change set out to find. **0 of 240 configurations clear n>=100 AND profit_factor > 1.0 at the taker rate, priced per product.** Zero at the maker rate. Five shipped signal families x 24 assets x 3 fees x 2 slippage regimes, plus the transfer check and the frequency axis: 960 trials, no argmax in Arm A, zero free parameters. **The engine never moved the 08-13 numbers; 565 new candles did.** Run over the truncated 08-13 corpus, today's engine reproduces three of the six cells that document printed BIT-IDENTICALLY -- turtle XRP-USD 157/1.223, PAXG-USDT 238/1.145, FET-USD 269/1.206 -- plus rsi_meanrev's 24-asset anchor to four decimals (median gross 1.1251 at median n=42). #442's gap-through-stop exit fill, its ratchet policy, and #523's re-derived cap are all inert here: turtle carries a static stop and neither rule declares trail_atr_mult/be_roll_rr. Recorded because the honest expectation was the opposite, and because a conservative-only correction that turns out to be a no-op is worth knowing before the next one is deferred on the same reasoning. **The deletion test, run in reverse on real data.** 08-13 sec.2 argued ZEC's edge was tail-carried by DELETING three trades. Adding three moved pullback_continuation on ZEC-USD from gross 1.025 to 1.695, and six moved turtle_breakout from 1.411 to 1.700 -- across the maker rate at 1.139, which 08-13 states flatly that none of its seven gross-positive cells survived. A profit factor a fortnight of ordinary data moves by +0.67 on 174 trades is measuring three trades. And per-product pricing deletes the episode anyway: ZEC is the fourth-thinnest name in the universe at 107.3bp, 21.5x the floor, and priced at its own liquidity its gross falls to 0.889 -- a loss before any fee. Corrections of record against 08-13, none of which rewrite it: - sec.6's "every signal rule the codebase ships has been measured" is FALSE. Three were measured; five ship. This document makes the sentence true again. - sec.3's "all seven die at the maker rate" was an artifact of the flat floor. Priced per product, SIX OF THE SEVEN are dead at zero fee. So is the WLD/TON observation that follows it (WLD 1.061 -> 0.626 at 120.9bp; TON 0.774 -> 0.317 at the 183.8bp cap). - sec.3.2's 0.034 transfer gap re-measures at 0.1172 flat / 0.1362 per-product. "Stably unprofitable" holds; "not overfit" is weaker than the document claims. - The 08-20 note's break-even-inside-the-allowance row is superseded. It is priced at psi=5bp, the floor 0 of 24 assets reach. Re-derived at the measured median psi=52.3bp with the note's own arithmetic: p_be 20.52% against a reconstructed 14.9% win rate, 5.62 points UNDERWATER inside the allowance. Rail 14 is still worth 14.3 points of break-even, more than any rule change measured in this directory -- but it is the boundary between decisively and clearly negative, not between negative and break-even. Arm A independently reproduces the 09-01 per-product restatement from scratch four days later on a longer corpus: median PF 0.3120 flat -> 0.2192 per-product, delta -0.0928 against its -0.090, every per-family figure within 0.005. The one cell above 1.0 at the fee actually paid is Arm B's turtle on ZEC-USD, PF 1.0275 at n=96 -- four trades below the admission floor, on the asset already known to be regime-bound and tail-carried. Reported because it exists. Nothing promoted, no rule row added, no config, allowlist or shipped parameter touched. 08-13 keeps its numbers and gains a forward pointer; records here are appended to, never revised (#247). Ledger verifies at 94 rows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXz13qp1UM3pBa6BqsRqvC
The file's own preamble makes the index a claim about the directory -- "Index is newest first, by the date each document carries in its filename" -- and it was missing a quarter of its records, with no 2026-09 section at all. An index missing records is worse than none, because it reads as complete. Adds a 2026-09 section (the 09-05 restatement, the equities DCA benchmark, the equities cost-fidelity study, the per-product slippage restatement, and the triple-barrier and CUSUM first measurements), and indexes 2026-08-27's pooled review preview, which was never listed either. Three of these are the records that supersede figures other documents in this directory still carry, so their absence from the index was the expensive kind. Verified mechanically rather than by eye: every link in the index resolves to a file that exists, and every .md and .py in the directory now has an entry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXz13qp1UM3pBa6BqsRqvC
…rs.py requires (#737) The guard is right and this record was in scope: it reports profit factors and states slippage assumptions, so a reader comparing figures across this directory needs to know which of its columns are priced at the floor. The note is not the boilerplate its neighbours carry, because this is the one record that is not uniformly flat-priced -- it reports BOTH regimes. Saying "the figures below are priced at the flat 5bp floor" would be false here. It says instead what is true: the flat columns are a bridge to the numbers 08-13 printed rather than results, every verdict on the page is stated at per-product pricing and is unaffected by the correction, and the one place a flat figure reads like a finding -- ZEC crossing the maker rate in sec.2.1 -- is exactly what sec.2.2 exists to remove. Found by CI, which is the failure worth recording: the suite tests this repository's DOCUMENTATION conventions, not only its code, so adding a file to docs/experiments/ is itself a change the suite has opinions about. Running only the three test files whose code I had touched was the wrong model of the blast radius. Full local suite: 5750 passed, 1 deselected, 1 failed -- both the deselected and the failed test are credential-dependent and fail identically on a clean main checkout of this machine (test_scope.py::test_attest_writes_none_when_no_current_credential_resolves and test_executor.py::test_confirming_replaces_a_stale_fingerprint_rather_than_carrying_it_forward both assert no credential resolves; this machine has real ones, CI does not). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXz13qp1UM3pBa6BqsRqvC
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #737. Closes #738.
The verdict
0 of 240 configurations clear
n ≥ 100 ∧ profit_factor > 1.0at the taker rate this account actually pays, priced per product. Zero at the maker rate. 960 trials — five shipped signal families × 24 assets × 3 fees × 2 slippage regimes, plus the transfer check and the frequency axis. No argmax in Arm A, zero free parameters.The 08-13 null holds across a population 2.7× larger, under a strictly more expensive cost model, on a corpus 23 days longer.
The finding, which is not that number
The engine never moved the 08-13 numbers. 565 new candles did.
The deployment DB had grown since 08-13, so a naive re-run would have absorbed that and reported it as an engine effect. A truncated-corpus control separates them:
turtleXRP-USDturtlePAXG-USDTturtleFET-USDturtleCRV-USDturtleZEC-USDpullbackZEC-USDPlus a fourth identity from Arm A:
rsi_meanrevat its shippedoversold=20reproduces §4's anchor to four decimals — median gross 1.1251 at median n=42.#442's gap-through-stop exit fill, its ratchet policy, and #523's re-derived cap are all inert for these rules. Recorded because the honest expectation was the opposite, and because a conservative-only correction that turns out to be a no-op is worth knowing before the next one is deferred on the same reasoning.
The deletion test, run in reverse on real data
§2 of 08-13 argued ZEC's edge was tail-carried, evidencing it by deleting three trades. Adding three moved
pullback_continuationon ZEC-USD from gross 1.025 → 1.695. A profit factor a fortnight of ordinary data moves by +0.67 on 174 trades is measuring three trades, not an edge.And per-product pricing deletes the episode anyway. ZEC is the fourth-thinnest name in the universe — 107.3bp, 21.5× the floor, on $1.08M/day. Priced at its own liquidity its gross falls to 0.889: a loss before any fee.
Corrections of record against 08-13
None of these rewrite it — it keeps its numbers and gains a forward pointer (#247).
Independent reproduction
Arm A re-derives the 09-01 per-product restatement from scratch, four days later on a longer corpus: median PF 0.3120 flat → 0.2192 per-product, delta −0.0928 against its −0.090, every per-family figure within 0.005.
The one cell above 1.0
Arm B's
turtle_breakouton ZEC-USD, per-product taker: PF 1.0275 at n=96 — four trades below the admission floor, on the asset already known to be regime-bound and tail-carried. Reported because it exists, not because it means anything.Also here (#738)
docs/experiments/README.mdwas missing six records and had no 2026-09 section, while its own preamble calls it a complete index. Verified mechanically: every link resolves, every.md/.pyhas an entry.Verification
keel trials verify: chain intact at 94 rows; ledger diff is1 insertion, 0 deletionstests/test_fee_reality_block.py16 passed ·tests/research/test_ledger.py12 passedruff checkclean on both driversNot changed
Nothing promoted, no rule row added, no config, allowlist or shipped parameter touched.
🤖 Generated with Claude Code
https://claude.ai/code/session_01EXz13qp1UM3pBa6BqsRqvC