Skip to content

feat: add cuckoo seed filter for pruned seed dedup - #36

Merged
daedalus merged 1 commit into
masterfrom
feat/cuckoo-seed-filter
Sep 28, 2026
Merged

daedalus merged 1 commit into
masterfrom
feat/cuckoo-seed-filter

Conversation

@daedalus

@daedalus daedalus commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

Add a CuckooFilter to track pruned seeds so their mutations are skipped at dedup time.

Changes

  • ** CLI flag** (default: off) — gates the feature
  • Startup: all seeds under are loaded into the filter
  • Mutation check: in , if the parent seed's hash is in the filter (was pruned), the mutation is skipped and original data is returned
  • During minimization: pruned seed hashes are added to the filter in
  • Re-admission: when a pruned seed is recovered in , its hash is removed from the filter
  • Filter sizing: capacity =

Tests

  • 8 new regression tests in
  • Updated in
  • All 141 affected tests pass

Verification

  • Pre-commit hooks pass (ruff, format, impactguard)
  • No regressions in existing cuckoo, bloom, operator registry tests

Summary by Sourcery

Add optional pruned-seed tracking to avoid reprocessing mutations from seeds removed during corpus minimization.

New Features:

  • Add an opt-in Cuckoo filter that tracks pruned seeds and skips mutations originating from them during deduplication.

Bug Fixes:

  • Remove recovered seeds from the pruned-seed filter so they can be admitted again.

Enhancements:

  • Load existing pruned seeds at startup and size the filter according to corpus size with a minimum capacity.

Tests:

  • Add regression coverage for startup loading, mutation skipping, normal mutation behavior, disabled operation, filter sizing, and pruned-seed tracking.

- Add --cuckoo-seed-filter CLI flag to gate the feature
- Initialize CuckooFilter for pruned seed tracking when flag is set
- Load pruned seeds into filter at startup (corpus/seeds/pruned/)
- Check filter in _dedup_mutate(): skip mutation if parent seed was pruned
- Add pruned seed hashes to filter during corpus minimization (_commit_minimize)
- Remove pruned seed hashes from filter when seeds are re-admitted (_recover_uncovered)
- Update _StubFuzzer with cuckoo_seed_filter attribute
- Add regression tests for the new feature

Fixes: tracker issue - ensure pruned seeds are not re-mutated after
corpus minimization, preventing wasted fuzzing effort on seeds already
deemed not valuable.
Copilot AI lite review requested due to automatic review settings September 28, 2026 14:48

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @daedalus, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 4 days and 13 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@sourcery-ai

sourcery-ai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Reviewer's Guide

Adds an opt-in CuckooFilter that is sized from the active corpus, populated from persisted and newly pruned seeds, cleared on re-admission, and consulted before mutation to avoid regenerating known-pruned seeds; regression tests cover the lifecycle and unchanged default behavior.

Sequence diagram for pruned seed filtering during mutation

sequenceDiagram
    participant Fuzzer
    participant CuckooFilter
    participant Mutator
    Fuzzer->>CuckooFilter: contains(_seed_key(data))
    alt parent seed was pruned
        CuckooFilter-->>Fuzzer: true
        Fuzzer-->>Fuzzer: return data
    else parent seed is not pruned
        CuckooFilter-->>Fuzzer: false
        Fuzzer->>Mutator: mutate(data)
        Mutator-->>Fuzzer: mutated data
    end
Loading

Sequence diagram for pruned seed filter lifecycle

sequenceDiagram
    participant Fuzzer
    participant Corpus
    participant CuckooFilter
    participant CorpusManager
    Fuzzer->>CuckooFilter: create capacity=max(10 * corpus size, 100000)
    Fuzzer->>Corpus: load pruned seeds from seeds/pruned/
    Corpus-->>Fuzzer: persisted pruned seed data
    Fuzzer->>CuckooFilter: add(_seed_key(data))
    CorpusManager->>CuckooFilter: add(_hash(seed))
    CorpusManager->>CuckooFilter: remove(seed_key(recovered seed))
Loading

File-Level Changes

Change Details Files
Add an opt-in CuckooFilter lifecycle for tracking pruned seed hashes.
  • Expose the --cuckoo-seed-filter CLI option and thread it into Fuzzer construction.
  • Allocate the filter at max(10 × active corpus size, 100,000) capacity.
  • Load hashes from corpus/seeds/pruned/ during startup.
  • Add hashes for seeds removed by minimization and remove hashes when seeds are re-admitted.
src/fuzzer_tool/cli/commands.py
src/fuzzer_tool/services/fuzzer.py
src/fuzzer_tool/services/corpus_manager.py
Use the pruned-seed filter to suppress redundant mutations.
  • Check the parent seed hash before invoking mutation or execution deduplication.
  • Return the original data and increment dedup hits when the hash is present.
  • Leave existing behavior unchanged when the feature is disabled or the seed is not filtered.
src/fuzzer_tool/services/fuzzer.py
Add regression coverage for filter initialization, sizing, gating, and mutation behavior.
  • Verify startup loading of one and multiple pruned seeds.
  • Verify active seeds are excluded, filtered mutations are skipped, and unfiltered mutations proceed.
  • Verify disabled mode and minimum capacity.
  • Update the bloom dedup test fixture with the optional filter attribute.
tests/test_regression_cuckoo_seed_filter.py
tests/test_bloom_exec_dedup.py

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical filter-recovery and symlink-safety issues, plus capacity and compatibility problems, remain unresolved.

Review effort: Lite
Findings: 2 High severity · 2 Medium severity · 1 Low severity

Open (5)
What changed in this PR

Adds optional Cuckoo-filter tracking for pruned seeds to avoid redundant mutations.

Changes:

  • Adds CLI and Fuzzer configuration.
  • Loads, tracks, and removes pruned seed hashes.
  • Adds regression coverage and updates test stubs.
File Summary
tests/​test_regression_cuckoo_seed_filter.py Adds feature regression tests.
tests/​test_bloom_exec_dedup.py Updates deduplication test support.
src/​fuzzer_tool/​services/​fuzzer.py Configures the filter and gates mutations.
src/​fuzzer_tool/​services/​corpus_manager.py Updates filter state during pruning and recovery.
src/​fuzzer_tool/​cli/​commands.py Adds the CLI option and help text.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +1759 to +1761
if f.cuckoo_seed_filter is not None:
h = self.seed_key(seed)
f.cuckoo_seed_filter.remove(h)
Comment on lines +4154 to +4157
if not fh.is_file():
continue
try:
data = fh.read_bytes()
# to the filter. When a seed is pruned during minimization, its hash
# is added. In _dedup_mutate(), if a mutation's hash is in the
# filter, the mutation is skipped (original data returned).
cuckoo_seed_filter=False,
Comment on lines +2040 to +2042
# Sized at 10x the corpus; minimum 100_000 to match
# the exec bloom default.
self.cuckoo_seed_filter = CuckooFilter(capacity=max(10 * len(self.corpus), 100_000))
Comment on lines +4045 to +4050
fuzz_parser.add_argument(
"--cuckoo-seed-filter",
action="store_true",
default=False,
help=(
"Track pruned seeds in a CuckooFilter so their mutations are "
@daedalus
daedalus merged commit f62c923 into master Sep 28, 2026
1 check passed
@daedalus
daedalus deleted the feat/cuckoo-seed-filter branch October 1, 2026 19:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants