Skip to content

docs(minimax): E3 alphabeta vs mcts on png — null - #25

Merged
daedalus merged 4 commits into
masterfrom
claude/minimax-alphabeta-mcts-ooqvo5
Sep 27, 2026
Merged

daedalus merged 4 commits into
masterfrom
claude/minimax-alphabeta-mcts-ooqvo5

Conversation

@daedalus

@daedalus daedalus commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

This PR records the result of the E3 A/B from the minimax handover. It is a docs-only follow-up to #20.

Setup: bench_paired.py with arms elo-mcts and elo-alphabeta, target set container_signal (png_read_noasan.so), 20 seeds × 10k execs.

arm mean edges sd median
elo-mcts 131.0 20.8 136.5
elo-alphabeta 130.6 27.0 134.0

Result: no measurable difference. Paired, alphabeta vs mcts went 12 wins / 8 losses, McNemar p=0.50, Wilcoxon p=0.67. The median difference is +8 edges. Repeated runs of the same mcts cell already span a median of 10 edges (range 0–39 over 9 cells), so the difference is inside the noise.

Caveats:

  • The runs were not --lock-single-thread: three shards ran in parallel on 4 cores.
  • The container restarted twice mid-run; the harness resumed from the saved runs.
  • Edge counts are at a fixed exec budget, so contention should not change them. EPS numbers are not comparable across runs.
  • mcts has repeated runs on 9 of the 20 seeds (seed 0 has 3); alphabeta has 1 run on every seed.

Changes:

  • Updated the minimax handover §1 and §2, plus the E3 row in handover_FINDINGS.md.
  • Added docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md with the full numbers.
  • Added a TODO for running E3 on ffmpeg_read. That first needs a new target set in eval_set.py, because no existing set includes ffmpeg. If that comparison is also null, retire --alphabeta.

🤖 Generated with Claude Code

https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt

Summary by Sourcery

Document the null PNG alpha-beta-versus-MCTS evaluation and track the remaining ffmpeg comparison before deciding whether to retire alpha-beta.

Enhancements:

  • Record the PNG E3 comparison as inconclusive/null, with no evidence that alpha-beta outperforms MCTS.
  • Update minimax handover findings to reflect the completed PNG evaluation and clarify that the remaining phases are wired but unmeasured.

Documentation:

  • Add detailed results and caveats for the alpha-beta versus MCTS PNG benchmark.
  • Document a follow-up evaluation on ffmpeg_read, including the need for a dedicated target set and criteria for retiring alpha-beta.

20 seeds x 10k execs, container_signal png_read_noasan.so:
12W/8L, McNemar p=0.50, median delta +8 edges, equal to the
median within-cell rep spread (8). ffmpeg_read still owed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
@sourcery-ai

sourcery-ai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Reviewer's Guide

This docs-only PR records the png E3 paired comparison as inconclusive/null: alpha-beta shows no measurable advantage over MCTS, with the observed difference at the replication-noise floor. It updates the minimax handover status and adds an ffmpeg_read follow-up that will determine whether the alpha-beta arm should be retired.

File-Level Changes

Change Details Files
Recorded the png E3 benchmark as a null result for alpha-beta versus MCTS and documented the experimental limitations.
  • Added paired benchmark metrics, statistical tests, and within-cell noise comparison.
  • Captured parallel execution, container restart, fixed-budget, and replication caveats.
  • Updated the minimax handover and findings tables to mark png complete and ffmpeg pending.
docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md
docs/handover/handover_minimax_implementation_2026-09-01.md
docs/handover/handover_FINDINGS.md
Added the next-step decision rule for validating or retiring the alpha-beta arm.
  • Added an ffmpeg_read E3 TODO using locked single-thread paired runs.
  • Specified retiring --alphabeta in favor of --mcts if the follow-up is also null.
docs/TODO.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@daedalus
daedalus marked this pull request as ready for review September 27, 2026 13:15
Copilot AI lite review requested due to automatic review settings September 27, 2026 13:15

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 1 issue

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="docs/TODO.md" line_range="25" />
<code_context>
+- [ ] **E3 on ffmpeg_read** (2026-09-27) — png is null (`docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md`: 12W/8L, p=0.50, Δ at the rep-noise floor). Vendor + build `ffmpeg_read`, run `--arms elo-mcts,elo-alphabeta` with `--reps 2 --lock-single-thread`. If null again, retire `--alphabeta` in favour of `--mcts`.
</code_context>
<issue_to_address>
**issue:** The documented follow-up cannot be run through `bench_paired.py` as described because `ffmpeg_read` is not present in any current benchmark target set; `container_signal` contains only `png_read_noasan.so` and `jpeg_read_noasan.so`, so selecting the ffmpeg target is rejected as unknown or produces no matrix cell.

**Triggers:** When someone follows the TODO after building `ffmpeg_read` but does not first modify the benchmark target set.

**Suggested fix:** Add the built ffmpeg artifact to an explicit benchmark target set, or document the required target-set update and exact artifact basename alongside the TODO.
</issue_to_address>

Sourcery assessment

Approval pending. 1 finding to address first.

Blocking findings: docs/TODO.md:25


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread docs/TODO.md Outdated
No eval_set target set contains ffmpeg_read; the follow-up needs a new
named set before bench_paired can select it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sourcery assessment

Approved.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Resolve the documented handover consistency and benchmark reproducibility issues.

Review effort: Lite
Findings: 2 Low severity

Open (2)
What changed in this PR

Docs-only follow-up documenting the null PNG E3 comparison of alpha-beta and MCTS.

Changes:

  • Records benchmark results, caveats, and statistical analysis.
  • Updates minimax handover and findings.
  • Adds the FFmpeg follow-up TODO.
File Summary
docs/​TODO.md Tracks the FFmpeg evaluation.
docs/​learnings/​2026-09-27-alphabeta-vs-mcts-png.md Records E3 data and methodology; replicate-count details need clarification.
docs/​handover/​handover_minimax_implementation_2026-09-01.md Updates minimax status; validation wording needs clarification.
docs/​handover/​handover_FINDINGS.md Marks PNG E3 as null; summary and status counts need updating.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread docs/handover/handover_minimax_implementation_2026-09-01.md
Comment thread docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md Outdated
9 mcts cells are replicated, not 10 (seed 0 has 3 runs). Per-cell
range over all 9: median 10, not 8; median delta +8 stays inside it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
…abeta-mcts-ooqvo5

# Conflicts:
#	docs/TODO.md
@daedalus
daedalus merged commit 1a71a8f into master Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants