Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@
- [ ] **E3 on ffmpeg_read** (2026-09-27) — png is null (`docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md`: 12W/8L, p=0.50, Δ inside the rep noise). Vendor + build `ffmpeg_read` (`tools/vendor_ffmpeg.sh`, `tools/build_targets.sh`). No target set in `tools/lib/eval_set.py` contains ffmpeg: add a new named set (e.g. `ffmpeg_signal`, register it in `TARGET_SETS`; do not edit an existing set) holding the built `ffmpeg_read.so` (or `ffmpeg_read_nosan.so`) path with `-m 65536`, then run `--arms elo-mcts,elo-alphabeta --set ffmpeg_signal --reps 2 --lock-single-thread`. If null again, retire `--alphabeta` in favour of `--mcts`.
- [ ] **`--vendor-tracecmp` cannot configure vendored libpng** (2026-09-27) — `tools/build_targets.sh` passes `-fsanitize-coverage=trace-pc-guard` in `CFLAGS` to libpng's `./configure`; configure's link test has no provider for `__sanitizer_cov_trace_pc_guard*`, so it dies with "C compiler cannot create executables" (and the `-L../zlib` zlib probe fails the same way against the instrumented `libz.a`). Worked around by hand: configure plain, then `make libpng16.la CFLAGS=<instrumented>`. Fix in the script the same way. Also: `tools/vendor_libpng.sh` fetches zlib as a GitHub tarball, which the sandbox proxy 403s; `git clone --branch v1.3.1` works.
- [ ] **A/B the Zipf saturation-gate veto** (2026-09-27) — `core/zipf.py` is wired (summary, report, stats file, gate veto) and calibrated only for speed: 20k-exec fuzzgoat runs never reach Chao2 ≥ 0.99, so the veto never fired. png_read (vendored instrumented libpng, clang, `_noasan.so`) cannot test it either: Chao2 plateaus at 87–89% from 40k to 158k execs on the generated seed corpus and at 84–87% from an empty one (80k), and the gate only engages at execs 1–8, where m < 2 makes Chao2 report 1.0 by construction. With the classifier guards the png spectrum reads `not_power_law`, so the veto is inert there regardless. Needs a target whose Chao2 actually crosses 0.99, and long campaigns where the gate engages, comparing edges with/without the veto (no CLI toggle yet — add one for `bench_paired.py` arms). Also open: `coverage_growth_model()` assumes exponential saturation, which contradicts a Zipf tail; compare its `projected_total` against the Heaps projection on real runs before choosing one.
- [ ] **A/B the minimax arms** (2026-09-27) — Phases 2-5 are wired, not measured. `bench_paired.py run --arms elo,elo-op-minimax` (Phase 4; also check EPS, Hard Rule 41), `--arms baseline,wall-order --set cmplog` (Phase 3; the claim is ≥3 sequential walls, e.g. PNG sig → IHDR → IDAT CRC), `--arms minimize-2k,minimax-select` (Phase 5; also measure coverage drop when one kept seed is removed). Phase 2: `analyse --risk-matrix` over a heterogeneous set (png, jpeg, sqlite, ffmpeg). Keep or delete each per `docs/handover/handover_minimax_implementation_2026-09-01.md` §2.
- [ ] **Minimax arms: png is null, decide `--hail-mary` membership** (2026-09-27) — png, 20 seeds (`docs/learnings/2026-09-27-minimax-phases-png.md`): op-minimax 10W/10L, minimax-select 10W/10L, wall-order 6W/14L (p=0.115, Δ -4.5). All three are on under `--hail-mary`. Next: replicate wall-order with `--reps 2 --lock-single-thread` on `--set cmplog` (log condstmt_solve pick counts first — exposure is unverified); Phase 2 risk matrix needs jpeg/sqlite/ffmpeg built. Drop from hail-mary or delete per result.
- [ ] **`test_transitions_populate_from_public_interface` is flaky** (2026-09-27) — 15/300 `RandPool` seeds leave `transition_counts` empty: always-success rewards let one op win every pick, so no transition between distinct ops is recorded. The test seeds `random`, but `MonteCarloScheduler` draws from the unseeded global pool. Inject a seeded pool.
- [ ] **Gravity splice fit never refit in any A/B** (2026-09-26) — `--splice-donor gravity` is null on fuzzgoat (2W/3L) and `container_signal` (20W/19L/1T, p=1), but every measured run used prior exponents: splice rounds yield ≤23 hits per 2048, below `MIN_POSITIVES=32`. `bench_paired.py` does not capture the `Gravity splice:` line, so the paired cells cannot show it. Before a further A/B: record hits/refits per cell, or find a target where splicing pays; otherwise retire the arm.
- [ ] **A/B the record-phase position arm, and re-examine the 64-byte TE cap** (2026-09-19) — `phase_lock` extrapolates the causal byte map to `offset + k*stride` across the buffer, but is wired and tested, not measured. Two things to check before trusting it. First, the evidence it extrapolates from is still whatever `update_te_causal_map` collected under `max_pos = min(64, ...)`, so a stride above ~32 gives the fold no repeated records to aggregate — the lock is then an extrapolation from a single arc, which the Rayleigh test cannot distinguish from a real one because it only asks about uniformity. Compare lock rates and yield bucketed by inferred stride before and after; if all the value sits at small strides, say so in the docstring rather than leaving the general claim. Second, the cap itself: raising it costs one `transfer_entropy` call per position per update, which is why it is 64. Measure that cost against the yield of the positions past it, because if the cap can move the fold stops being a workaround.
Expand Down
5 changes: 5 additions & 0 deletions docs/handover/handover_minimax_implementation_2026-09-01.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,11 @@ Wired in PR #20 (2026-09-27) behind `--risk-matrix` (P2), `--wall-order` (P3),
three fuzz flags are on under `--hail-mary`. Bench arms: `elo-op-minimax`,
`wall-order`, `minimax-select`. The "Gap" column below is the pre-wiring state.

**png result (2026-09-27, 20 seeds x 10k execs):** P4 `elo-op-minimax` 10W/10L,
Δ 0; P5 `minimax-select` 10W/10L, Δ -0.5; P3 `wall-order` 6W/14L, McNemar
p=0.115, Δ -4.5 (negative lean, not significant). P2 unmeasured (needs a
heterogeneous target set). Details: `docs/learnings/2026-09-27-minimax-phases-png.md`.
Comment on lines +37 to +40

| Phase | Symbol | Gap |
|---|---|---|
| 2 risk matrix | `core/analyzers/analyzer_elo.py::EloTracker.select_minimax_scheduler`, `record_match(target=)` | Live fuzzer builds `BayesianEloTracker` (`core/analyzer_registry.py`); `EloTracker(use_minimax=True)` never constructed. No `--use-minimax` CLI flag. Offline half works: `bench_paired.py --risk-matrix`. |
Expand Down
36 changes: 36 additions & 0 deletions docs/learnings/2026-09-27-minimax-phases-png.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# Minimax Phases 3-5 on png — null (wall-order leans negative)

2026-09-27. `tools/lib/bench_paired.py run --set container_signal --targets
png_read_noasan.so`, seeds 0-19, 10k execs, 1 rep, each arm paired against
its `ARM_BASELINES` entry. 3 parallel shards on 4 cores (not
`--lock-single-thread`); eps is contended and only roughly comparable.

| arm (vs baseline) | W | L | McNemar | Wilcoxon | med Δ edges |
|---|---|---|---|---|---|
| `elo-op-minimax` vs `elo` (P4) | 10 | 10 | 1.0 | 0.68 | 0.0 |
| `wall-order` vs `baseline` (P3) | 6 | 14 | 0.115 | 0.17 | -4.5 |
| `minimax-select` vs `minimize-2k` (P5) | 10 | 10 | 1.0 | 0.98 | -0.5 |

| arm | mean edges | sd | median | eps | corpus |
|---|---|---|---|---|---|
| baseline | 121.2 | 25.3 | 128.5 | 43.8 | 58.7 |
| wall-order | 116.2 | 28.3 | 127.5 | 42.6 | 55.4 |
| elo | 122.5 | 32.0 | 132.5 | 37.4 | 62.8 |
| elo-op-minimax | 127.2 | 16.4 | 129.5 | 36.6 | 63.0 |
| minimize-2k | 123.1 | 22.6 | 130.5 | 37.7 | 40.7 |
| minimax-select | 130.5 | 56.3 | 126.5 | 45.8 | 39.1 |

Notes:
- minimax-select seed 17 scored 348 edges; a rerun of that cell gave 126
(minimize-2k rerun: 139). Not reproducible: one lucky draw. Mean without
it is 119.1; the paired tests are rank-based and unaffected.
Comment on lines +24 to +26
- minimax-select did not inflate the corpus (39.1 vs 40.7 kept seeds).
- wall-order: cmplog is live on png (733 pairs), so `condstmt_solve` is in
the pool, but per-operator pick counts are not logged; how often the
wall order actually decided a pick is unmeasured.
- op-minimax: no eps regression visible (36.6 vs 37.4, within contention).
- Phase 2 (risk matrix) needs a heterogeneous target set; only png builds
here (jpeg/ffmpeg libs absent), so it is unmeasured.

Verdict: no evidence any of the three helps on png. wall-order's 6W/14L is
the only lean, and it is negative. All three are on under `--hail-mary`.
Loading