Skip to content

feat(minimax): wire Phases 2-5 behind flags - #20

Merged
daedalus merged 2 commits into
masterfrom
claude/minimax-alphabeta-mcts-ooqvo5
Sep 27, 2026
Merged

daedalus merged 2 commits into
masterfrom
claude/minimax-alphabeta-mcts-ooqvo5

Conversation

@daedalus

@daedalus daedalus commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Closes the "Phases 2–5 have no callers" item in docs/handover/handover_minimax_implementation_2026-09-01.md. The E3 alphabeta-vs-mcts A/B on png is still running; its result will land in a follow-up commit on this branch.

Wiring

Phase Flag Consumer
2 risk matrix bench_paired.py analyse --risk-matrix prints the minimax-robust arm via EloTracker.select_minimax_scheduler
3 comparison wall fuzz --wall-order _op_condstmt_solve solves the head of solve_comparison_wall over the next 5 unsolved branches
4 operator sequencing fuzz --op-minimax (needs --mc-bandit) bandit strategy calls select_op_minimax
5 robust corpus fuzz --minimax-select, minimize --minimax-robust _select_budget in auto-minimize, and cmin pruning

The new flags are left out of --hail-mary until they are measured. New bench arms: elo-op-minimax (vs elo), wall-order (vs baseline), minimax-select (vs minimize-2k).

Bugs fixed along the way

  • P4: the root only searched operators[:4] in list order. It also scored ops by the posterior mean, so it never explored. It now uses Thompson draws (shared with select_op) and heapq.nlargest.
  • P5: a seed's loss was scored as its size instead of the edges only it covers. The pruning bound compared a seed count to an edge count. Selection now minimizes (max unique loss, count at max), stops when that no longer strictly decreases, and costs O(|edges|) per candidate.
  • P2: the matrix took the worst seed, which is about 1.0 for every arm, so the pick fell back to list order. It also read missing data as zero regret and dropped the baseline arm. It now uses the mean over seeds, treats missing data as 1.0, and keeps the baseline.

Tests

Five new test files. Each feature has a falsification test, an adversarial test and a control. The P5 core tests fail on the old rate_distortion.py (6 of 9 fail).

Pre-existing failures (also red on the base commit)

test_integration ASAN/EPS tests, *constructor_flag* ordering tests, crash_file_protection, no_op_mutations, scheduler_operator_reach[BOGPUCB] and prng_state_learner live tests. test_transitions_populate_from_public_interface is flaky: it fails for 15 of 300 RandPool seeds, with identical results on the base commit. It is logged in docs/TODO.md.

🤖 Generated with Claude Code

https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt


Generated by Claude Code

Summary by Sourcery

Wire Minimax Phases 2–5 into user-facing workflows and correct their selection behavior so they can be measured behind opt-in flags.

New Features:

  • Wire Minimax Phases 2–5 into benchmark analysis, fuzzing, bandit scheduling, and corpus minimization behind opt-in flags.
  • Add benchmark arms for evaluating minimax operator scheduling, comparison-wall ordering, and corpus selection.

Bug Fixes:

  • Correct minimax operator exploration, robust corpus loss accounting, and risk-matrix handling of seed aggregation, missing data, and baseline arms.

Enhancements:

  • Expose minimax-robust pruning and admission as supported corpus-selection strategies with robust backup selection.

Documentation:

  • Document the new minimax flags, benchmark arms, and robust minimization workflow.

Tests:

  • Add regression coverage for flag wiring, operator minimax selection, comparison-wall ordering, robust corpus selection, and risk-matrix minimax analysis.

Chores:

  • Record follow-up A/B measurement work and an existing flaky transition test in project TODOs.

Dead code with no caller is now reachable and tested:
- P2: bench_paired analyse --risk-matrix prints the minimax-robust arm
  via EloTracker.select_minimax_scheduler. Matrix: mean regret over
  seeds (worst seed was ~1.0 for every arm), missing = 1.0, baseline kept.
- P3: fuzz --wall-order; condstmt_solve takes the head of
  solve_comparison_wall over the next 5 unsolved branches.
- P4: fuzz --op-minimax; bandit picks via select_op_minimax. Fixed root
  beam (searched first 4 ops in list order) and mean scoring (no
  exploration): now Thompson draws, shared with select_op.
- P5: fuzz --minimax-select (auto_minimize budget), minimize
  --minimax-robust. Fixed loss (seed size -> unique edges) and the
  pruning bound (seed count vs edge count); O(|edges|) per candidate.

Bench arms: elo-op-minimax, wall-order, minimax-select (+ minimize-2k).
Flags excluded from --hail-mary until measured.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
@sourcery-ai

sourcery-ai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Reviewer's Guide

This PR activates previously uncalled Minimax Phases 2–5 behind opt-in fuzzing, minimization, and benchmarking flags, while fixing their operator exploration, robust corpus-loss accounting, comparison-wall dispatch, and risk-matrix selection logic; extensive regression tests and documentation cover the new paths, which remain unmeasured pending A/B runs.

Sequence diagram for Minimax operator scheduling

sequenceDiagram
    participant Fuzzer
    participant Operators
    participant Scheduler as MonteCarloScheduler
    participant Target
    Fuzzer->>Operators: select_op(ops)
    Operators->>Scheduler: select_op_minimax(ops)
    Scheduler->>Scheduler: _thompson_vals(ops)
    Scheduler->>Target: Alpha-beta block response
    Target-->>Scheduler: Minimax value
    Scheduler-->>Operators: Selected operator
    Operators-->>Fuzzer: Apply operator
Loading

Sequence diagram for robust corpus minimization

sequenceDiagram
    participant CLI
    participant Minimizer
    participant Robust as RateDistortionCorpus
    participant Cover as _RobustCover
    CLI->>Minimizer: minimize_corpus(prune=MINIMAX_ROBUST)
    Minimizer->>Robust: minimax_robust_pruning(seed_edges)
    Robust->>Cover: _greedy_cover(...)
    Cover-->>Robust: Set-cover selection
    Robust->>Cover: _robust_fill(...)
    Cover-->>Robust: Backup seeds with lower risk
    Robust-->>Minimizer: Kept files and coverage fraction
    Minimizer-->>CLI: Pruning result
Loading

Flow diagram for Minimax comparison-wall ordering

flowchart TD
    Start[condstmt_solve] --> Flag{wall_order_enabled}
    Flag -- enabled --> Window[Select next 5 unsolved branches]
    Window --> Solver[Z3Solver.solve_comparison_wall]
    Solver --> Head["Choose order[0]"]
    Flag -- disabled --> Random[Choose random unsolved branch]
    Head --> Mutate[Apply operand-substitution mutation]
    Random --> Mutate
Loading

File-Level Changes

Change Details Files
Wire Minimax Phases 2–5 into opt-in CLI consumers and benchmark arms.
  • Add fuzz flags for wall ordering, operator minimax scheduling, and robust corpus selection.
  • Add minimize robust-pruning mode and risk-matrix analysis output with minimax arm selection.
  • Connect the flags to mutation context, fuzzer state, operator dispatch, auto-minimization, and cmin.
  • Add benchmark arms with explicit baselines, while leaving the features opt-in pending A/B measurement.
src/fuzzer_tool/cli/commands.py
src/fuzzer_tool/core/mutator_interface.py
src/fuzzer_tool/services/fuzzer.py
src/fuzzer_tool/services/minimize.py
src/fuzzer_tool/services/operators.py
tools/lib/bench_paired.py
Correct minimax operator scheduling to explore Thompson-sampled candidates through bounded alpha-beta search.
  • Share Thompson draws with the normal operator selector instead of using posterior means.
  • Rank the root candidate pool with heapq.nlargest rather than limiting search to list order.
  • Preserve bounded beam/depth traversal and route bandit selections through minimax when enabled.
src/fuzzer_tool/core/schedulers/op_monte_carlo.py
src/fuzzer_tool/services/operators.py
Implement robust corpus selection using unique-edge loss and incremental risk tracking.
  • Track per-seed uniquely owned edges, with risk ordered by maximum loss and the number of seeds at that loss.
  • Separate mandatory set-cover seeds from optional robust backups and stop when risk no longer strictly improves.
  • Use the robust selector for both auto-minimize budget admission and minimax cmin pruning with linear-per-candidate edge accounting.
src/fuzzer_tool/core/rate_distortion.py
src/fuzzer_tool/services/corpus_manager.py
src/fuzzer_tool/services/minimize.py
Make risk-matrix minimax selection target-oriented and conservative about missing evidence.
  • Average regret across seeds and maximize risk across targets instead of taking the worst seed per target.
  • Assign missing arm-target data a regret of 1.0 and retain the baseline arm in the matrix.
  • Print the selected minimax-robust arm from risk-matrix analysis.
tools/lib/bench_paired.py
Prioritize comparison-wall solving through a bounded minimax order when enabled.
  • Pass wall-order state through MutationContext without adding a fuzzer reference.
  • Evaluate only the next five unsolved comparisons, lazily initializing the Z3 wall solver and selecting the ordered head.
  • Retain random branch selection when the flag is disabled.
src/fuzzer_tool/core/mutator_interface.py
src/fuzzer_tool/services/operators.py
Add regression coverage for flag wiring, selection algorithms, and control/adversarial behavior.
  • Test CLI propagation, defaults, hail-mary behavior, benchmark arm composition, and minimize mode mapping.
  • Test Thompson exploration, robust corpus loss accounting, risk-matrix handling, and bounded wall ordering.
  • Extend mutation-context expectations for the new wall-order field.
tests/test_regression_minimax_flags.py
tests/test_regression_minimax_robust_corpus.py
tests/test_regression_op_minimax.py
tests/test_regression_risk_matrix_minimax.py
tests/test_regression_wall_order.py
tests/test_regression_mutator_interface.py
Document the new experimental functionality and outstanding measurement work.
  • Document flags, implementation behavior, expected costs, and robust minimization usage.
  • Record that benchmark arms require A/B evaluation and retain the known transition-test flake note.
  • Update the changelog with wiring, bug fixes, and benchmark status.
CHANGELOG.md
docs/DEEP_DIVE.md
docs/TODO.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

--op-minimax, --wall-order, --minimax-select join _HAIL_MARY_FLAGS.
hail-mary already sets mc_bandit (op-minimax's host) and a 5000-exec
minimize cycle (minimax-select's host).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fZHdKXWjWeWR5XqZC5YCt
@daedalus
daedalus marked this pull request as ready for review September 27, 2026 03:16
Copilot AI lite review requested due to automatic review settings September 27, 2026 03:16
@daedalus
daedalus merged commit bdf381d into master Sep 27, 2026
1 check passed

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!

Sourcery assessment

Needs a human reviewer. The new minimax pruning and corpus-admission paths can remove seeds from an in-place corpus, and reverting the code will not restore inputs already deleted; recovery requires an external backup or rebuilding the corpus. The impact is bounded to users who explicitly enable these opt-in flags, but the lost corpus data is not automatically repairable.


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Resolve the hail-mary opt-in conflict, preserve positional constructor compatibility, and correct risk-matrix aggregation for missing coverage and replicates.

Review effort: Lite
Findings: 2 Medium severity

Open (2)
What changed in this PR

Wires Minimax Phases 2–5 into benchmarking, fuzzing, scheduling, and corpus minimization behind flags.

Changes:

  • Adds Minimax flags and benchmark arms.
  • Fixes operator, corpus, and risk-matrix selection.
  • Adds regression tests and documentation.
File Description
tools/​lib/​bench_paired.py Adds benchmark arms and risk-matrix selection.
tests/​test_regression_wall_order.py Tests comparison-wall ordering.
tests/​test_regression_risk_matrix_minimax.py Tests risk-matrix selection.
tests/​test_regression_op_minimax.py Tests operator minimax scheduling.
tests/​test_regression_mutator_interface.py Tests context wiring.
tests/​test_regression_minimax_robust_corpus.py Tests robust corpus selection.
tests/​test_regression_minimax_flags.py Tests CLI flag wiring.
src/​fuzzer_tool/​services/​operators.py Integrates wall ordering and operator minimax.
src/​fuzzer_tool/​services/​minimize.py Adds minimax pruning modes.
src/​fuzzer_tool/​services/​fuzzer.py Adds Minimax configuration flags.
src/​fuzzer_tool/​services/​corpus_manager.py Applies robust corpus admission.
src/​fuzzer_tool/​core/​schedulers/​op_monte_carlo.py Implements Minimax operator scheduling.
src/​fuzzer_tool/​core/​rate_distortion.py Implements robust coverage selection.
src/​fuzzer_tool/​core/​mutator_interface.py Adds wall-order context state.
src/​fuzzer_tool/​cli/​commands.py Adds CLI flags and wiring.
docs/​TODO.md Records pending A/B measurements.
docs/​DEEP_DIVE.md Documents Minimax features.
CHANGELOG.md Records the new functionality.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +1933 to +1937
# Minimax Phases 3-5: op_minimax drives mc_bandit above; minimax_select
# acts in the hail-mary minimize cycle.
"op_minimax",
"wall_order",
"minimax_select",
Comment on lines +1269 to +1271
minimax_select=False,
op_minimax=False,
wall_order=False,
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants