From d59e8c17f289528cbac3e3b9896fd033036dfca2 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 16:38:24 +0000 Subject: [PATCH] =?UTF-8?q?docs(minimax):=20Phases=203-5=20on=20png=20?= =?UTF-8?q?=E2=80=94=20null,=20wall-order=20leans=20negative?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 20 seeds x 10k execs vs each arm's baseline: op-minimax 10W/10L, minimax-select 10W/10L, wall-order 6W/14L (McNemar p=0.115, delta -4.5). Phase 2 unmeasured (single buildable target). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt --- docs/TODO.md | 2 +- ...dover_minimax_implementation_2026-09-01.md | 5 +++ .../2026-09-27-minimax-phases-png.md | 36 +++++++++++++++++++ 3 files changed, 42 insertions(+), 1 deletion(-) create mode 100644 docs/learnings/2026-09-27-minimax-phases-png.md diff --git a/docs/TODO.md b/docs/TODO.md index 0a0ce111..5800369d 100644 --- a/docs/TODO.md +++ b/docs/TODO.md @@ -28,7 +28,7 @@ - [ ] **E3 on ffmpeg_read** (2026-09-27) — png is null (`docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md`: 12W/8L, p=0.50, Δ inside the rep noise). Vendor + build `ffmpeg_read` (`tools/vendor_ffmpeg.sh`, `tools/build_targets.sh`). No target set in `tools/lib/eval_set.py` contains ffmpeg: add a new named set (e.g. `ffmpeg_signal`, register it in `TARGET_SETS`; do not edit an existing set) holding the built `ffmpeg_read.so` (or `ffmpeg_read_nosan.so`) path with `-m 65536`, then run `--arms elo-mcts,elo-alphabeta --set ffmpeg_signal --reps 2 --lock-single-thread`. If null again, retire `--alphabeta` in favour of `--mcts`. - [ ] **`--vendor-tracecmp` cannot configure vendored libpng** (2026-09-27) — `tools/build_targets.sh` passes `-fsanitize-coverage=trace-pc-guard` in `CFLAGS` to libpng's `./configure`; configure's link test has no provider for `__sanitizer_cov_trace_pc_guard*`, so it dies with "C compiler cannot create executables" (and the `-L../zlib` zlib probe fails the same way against the instrumented `libz.a`). Worked around by hand: configure plain, then `make libpng16.la CFLAGS=`. Fix in the script the same way. Also: `tools/vendor_libpng.sh` fetches zlib as a GitHub tarball, which the sandbox proxy 403s; `git clone --branch v1.3.1` works. - [ ] **A/B the Zipf saturation-gate veto** (2026-09-27) — `core/zipf.py` is wired (summary, report, stats file, gate veto) and calibrated only for speed: 20k-exec fuzzgoat runs never reach Chao2 ≥ 0.99, so the veto never fired. png_read (vendored instrumented libpng, clang, `_noasan.so`) cannot test it either: Chao2 plateaus at 87–89% from 40k to 158k execs on the generated seed corpus and at 84–87% from an empty one (80k), and the gate only engages at execs 1–8, where m < 2 makes Chao2 report 1.0 by construction. With the classifier guards the png spectrum reads `not_power_law`, so the veto is inert there regardless. Needs a target whose Chao2 actually crosses 0.99, and long campaigns where the gate engages, comparing edges with/without the veto (no CLI toggle yet — add one for `bench_paired.py` arms). Also open: `coverage_growth_model()` assumes exponential saturation, which contradicts a Zipf tail; compare its `projected_total` against the Heaps projection on real runs before choosing one. -- [ ] **A/B the minimax arms** (2026-09-27) — Phases 2-5 are wired, not measured. `bench_paired.py run --arms elo,elo-op-minimax` (Phase 4; also check EPS, Hard Rule 41), `--arms baseline,wall-order --set cmplog` (Phase 3; the claim is ≥3 sequential walls, e.g. PNG sig → IHDR → IDAT CRC), `--arms minimize-2k,minimax-select` (Phase 5; also measure coverage drop when one kept seed is removed). Phase 2: `analyse --risk-matrix` over a heterogeneous set (png, jpeg, sqlite, ffmpeg). Keep or delete each per `docs/handover/handover_minimax_implementation_2026-09-01.md` §2. +- [ ] **Minimax arms: png is null, decide `--hail-mary` membership** (2026-09-27) — png, 20 seeds (`docs/learnings/2026-09-27-minimax-phases-png.md`): op-minimax 10W/10L, minimax-select 10W/10L, wall-order 6W/14L (p=0.115, Δ -4.5). All three are on under `--hail-mary`. Next: replicate wall-order with `--reps 2 --lock-single-thread` on `--set cmplog` (log condstmt_solve pick counts first — exposure is unverified); Phase 2 risk matrix needs jpeg/sqlite/ffmpeg built. Drop from hail-mary or delete per result. - [ ] **`test_transitions_populate_from_public_interface` is flaky** (2026-09-27) — 15/300 `RandPool` seeds leave `transition_counts` empty: always-success rewards let one op win every pick, so no transition between distinct ops is recorded. The test seeds `random`, but `MonteCarloScheduler` draws from the unseeded global pool. Inject a seeded pool. - [ ] **Gravity splice fit never refit in any A/B** (2026-09-26) — `--splice-donor gravity` is null on fuzzgoat (2W/3L) and `container_signal` (20W/19L/1T, p=1), but every measured run used prior exponents: splice rounds yield ≤23 hits per 2048, below `MIN_POSITIVES=32`. `bench_paired.py` does not capture the `Gravity splice:` line, so the paired cells cannot show it. Before a further A/B: record hits/refits per cell, or find a target where splicing pays; otherwise retire the arm. - [ ] **A/B the record-phase position arm, and re-examine the 64-byte TE cap** (2026-09-19) — `phase_lock` extrapolates the causal byte map to `offset + k*stride` across the buffer, but is wired and tested, not measured. Two things to check before trusting it. First, the evidence it extrapolates from is still whatever `update_te_causal_map` collected under `max_pos = min(64, ...)`, so a stride above ~32 gives the fold no repeated records to aggregate — the lock is then an extrapolation from a single arc, which the Rayleigh test cannot distinguish from a real one because it only asks about uniformity. Compare lock rates and yield bucketed by inferred stride before and after; if all the value sits at small strides, say so in the docstring rather than leaving the general claim. Second, the cap itself: raising it costs one `transfer_entropy` call per position per update, which is why it is 64. Measure that cost against the yield of the positions past it, because if the cap can move the fold stops being a workaround. diff --git a/docs/handover/handover_minimax_implementation_2026-09-01.md b/docs/handover/handover_minimax_implementation_2026-09-01.md index 4a045822..839da4e5 100644 --- a/docs/handover/handover_minimax_implementation_2026-09-01.md +++ b/docs/handover/handover_minimax_implementation_2026-09-01.md @@ -34,6 +34,11 @@ Wired in PR #20 (2026-09-27) behind `--risk-matrix` (P2), `--wall-order` (P3), three fuzz flags are on under `--hail-mary`. Bench arms: `elo-op-minimax`, `wall-order`, `minimax-select`. The "Gap" column below is the pre-wiring state. +**png result (2026-09-27, 20 seeds x 10k execs):** P4 `elo-op-minimax` 10W/10L, +Δ 0; P5 `minimax-select` 10W/10L, Δ -0.5; P3 `wall-order` 6W/14L, McNemar +p=0.115, Δ -4.5 (negative lean, not significant). P2 unmeasured (needs a +heterogeneous target set). Details: `docs/learnings/2026-09-27-minimax-phases-png.md`. + | Phase | Symbol | Gap | |---|---|---| | 2 risk matrix | `core/analyzers/analyzer_elo.py::EloTracker.select_minimax_scheduler`, `record_match(target=)` | Live fuzzer builds `BayesianEloTracker` (`core/analyzer_registry.py`); `EloTracker(use_minimax=True)` never constructed. No `--use-minimax` CLI flag. Offline half works: `bench_paired.py --risk-matrix`. | diff --git a/docs/learnings/2026-09-27-minimax-phases-png.md b/docs/learnings/2026-09-27-minimax-phases-png.md new file mode 100644 index 00000000..ec07b128 --- /dev/null +++ b/docs/learnings/2026-09-27-minimax-phases-png.md @@ -0,0 +1,36 @@ +# Minimax Phases 3-5 on png — null (wall-order leans negative) + +2026-09-27. `tools/lib/bench_paired.py run --set container_signal --targets +png_read_noasan.so`, seeds 0-19, 10k execs, 1 rep, each arm paired against +its `ARM_BASELINES` entry. 3 parallel shards on 4 cores (not +`--lock-single-thread`); eps is contended and only roughly comparable. + +| arm (vs baseline) | W | L | McNemar | Wilcoxon | med Δ edges | +|---|---|---|---|---|---| +| `elo-op-minimax` vs `elo` (P4) | 10 | 10 | 1.0 | 0.68 | 0.0 | +| `wall-order` vs `baseline` (P3) | 6 | 14 | 0.115 | 0.17 | -4.5 | +| `minimax-select` vs `minimize-2k` (P5) | 10 | 10 | 1.0 | 0.98 | -0.5 | + +| arm | mean edges | sd | median | eps | corpus | +|---|---|---|---|---|---| +| baseline | 121.2 | 25.3 | 128.5 | 43.8 | 58.7 | +| wall-order | 116.2 | 28.3 | 127.5 | 42.6 | 55.4 | +| elo | 122.5 | 32.0 | 132.5 | 37.4 | 62.8 | +| elo-op-minimax | 127.2 | 16.4 | 129.5 | 36.6 | 63.0 | +| minimize-2k | 123.1 | 22.6 | 130.5 | 37.7 | 40.7 | +| minimax-select | 130.5 | 56.3 | 126.5 | 45.8 | 39.1 | + +Notes: +- minimax-select seed 17 scored 348 edges; a rerun of that cell gave 126 + (minimize-2k rerun: 139). Not reproducible: one lucky draw. Mean without + it is 119.1; the paired tests are rank-based and unaffected. +- minimax-select did not inflate the corpus (39.1 vs 40.7 kept seeds). +- wall-order: cmplog is live on png (733 pairs), so `condstmt_solve` is in + the pool, but per-operator pick counts are not logged; how often the + wall order actually decided a pick is unmeasured. +- op-minimax: no eps regression visible (36.6 vs 37.4, within contention). +- Phase 2 (risk matrix) needs a heterogeneous target set; only png builds + here (jpeg/ffmpeg libs absent), so it is unmeasured. + +Verdict: no evidence any of the three helps on png. wall-order's 6W/14L is +the only lean, and it is negative. All three are on under `--hail-mary`.