docs(minimax): E3 alphabeta vs mcts on png — null - #25
Conversation
20 seeds x 10k execs, container_signal png_read_noasan.so: 12W/8L, McNemar p=0.50, median delta +8 edges, equal to the median within-cell rep spread (8). ffmpeg_read still owed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
Reviewer's GuideThis docs-only PR records the png E3 paired comparison as inconclusive/null: alpha-beta shows no measurable advantage over MCTS, with the observed difference at the replication-noise floor. It updates the minimax handover status and adds an ffmpeg_read follow-up that will determine whether the alpha-beta arm should be retired. File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 1 issue
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="docs/TODO.md" line_range="25" />
<code_context>
+- [ ] **E3 on ffmpeg_read** (2026-09-27) — png is null (`docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md`: 12W/8L, p=0.50, Δ at the rep-noise floor). Vendor + build `ffmpeg_read`, run `--arms elo-mcts,elo-alphabeta` with `--reps 2 --lock-single-thread`. If null again, retire `--alphabeta` in favour of `--mcts`.
</code_context>
<issue_to_address>
**issue:** The documented follow-up cannot be run through `bench_paired.py` as described because `ffmpeg_read` is not present in any current benchmark target set; `container_signal` contains only `png_read_noasan.so` and `jpeg_read_noasan.so`, so selecting the ffmpeg target is rejected as unknown or produces no matrix cell.
**Triggers:** When someone follows the TODO after building `ffmpeg_read` but does not first modify the benchmark target set.
**Suggested fix:** Add the built ffmpeg artifact to an explicit benchmark target set, or document the required target-set update and exact artifact basename alongside the TODO.
</issue_to_address>Sourcery assessment
Approval pending. 1 finding to address first.
Blocking findings: docs/TODO.md:25
No eval_set target set contains ffmpeg_read; the follow-up needs a new named set before bench_paired can select it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Resolve the documented handover consistency and benchmark reproducibility issues.
Review effort: Lite
Findings: 2
Open (2)
What changed in this PR
Docs-only follow-up documenting the null PNG E3 comparison of alpha-beta and MCTS.
Changes:
- Records benchmark results, caveats, and statistical analysis.
- Updates minimax handover and findings.
- Adds the FFmpeg follow-up TODO.
| File | Summary |
|---|---|
docs/TODO.md |
Tracks the FFmpeg evaluation. |
docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md |
Records E3 data and methodology; replicate-count details need clarification. |
docs/handover/handover_minimax_implementation_2026-09-01.md |
Updates minimax status; validation wording needs clarification. |
docs/handover/handover_FINDINGS.md |
Marks PNG E3 as null; summary and status counts need updating. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
9 mcts cells are replicated, not 10 (seed 0 has 3 runs). Per-cell range over all 9: median 10, not 8; median delta +8 stays inside it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
…abeta-mcts-ooqvo5 # Conflicts: # docs/TODO.md

This PR records the result of the E3 A/B from the minimax handover. It is a docs-only follow-up to #20.
Setup:
bench_paired.pywith armselo-mctsandelo-alphabeta, target setcontainer_signal(png_read_noasan.so), 20 seeds × 10k execs.Result: no measurable difference. Paired, alphabeta vs mcts went 12 wins / 8 losses, McNemar p=0.50, Wilcoxon p=0.67. The median difference is +8 edges. Repeated runs of the same mcts cell already span a median of 10 edges (range 0–39 over 9 cells), so the difference is inside the noise.
Caveats:
--lock-single-thread: three shards ran in parallel on 4 cores.Changes:
handover_FINDINGS.md.docs/learnings/2026-09-27-alphabeta-vs-mcts-png.mdwith the full numbers.ffmpeg_read. That first needs a new target set ineval_set.py, because no existing set includes ffmpeg. If that comparison is also null, retire--alphabeta.🤖 Generated with Claude Code
https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
Summary by Sourcery
Document the null PNG alpha-beta-versus-MCTS evaluation and track the remaining ffmpeg comparison before deciding whether to retire alpha-beta.
Enhancements:
Documentation: