Skip to content

docs(m6): raw benchmark data and README tables from release b782 - #199

Open
solderzzc wants to merge 2 commits into
mainfrom
docs/m6-b782-data
Open

solderzzc wants to merge 2 commits into
mainfrom
docs/m6-b782-data

Conversation

@solderzzc

@solderzzc solderzzc commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Commits the b782 benchmark raw data to docs/profiling/m6. That folder is linked as "benchmark script and raw JSONL", and the files there came from older builds, so they didn't match the b782 numbers being published (for example Qwen3.6 GPU decode at 548 tokens: 46.97 on main vs 48.35 on b782). Docs only.

Changes

  • qwen36_35b_a3b_4bit.{md,jsonl}: replaced with the full b782 run, GPU and --stream-experts at 548 / 2.3K / 9.8K / 40.8K tokens. It includes a note that the 40.8K needle check is unreliable in both modes (every miss had an OSPREY-NN code word, and GPU fails the same way on the same prompt), so the 40.8K rows are throughput only. The two 40.8K re-runs are included with "recheck": true.
  • gemma4_26b_a4b_4bit, gemma4_26b_a4b_8bit, qwen38_27b_4bit: only some lengths were re-run on b782, so each gets a b782 section on top and the existing rows stay, labelled "Earlier build (before b782)".
  • gemma4_26b_a4b_4bit_mtp_bf16.md: labelled as measured on an earlier build.
  • Each .md header names the build. New JSONL rows carry "build": "b782", and existing rows are unchanged.

Method

Measured on a Mac mini M6, 32 GB, macOS 27.0, using the official SwiftLM-b782-macos-arm64.tar.gz with scripts/profiling/m6_bench.py. That's 3 runs (1 at 32K and above) after one warm-up, temperature 0, reporting medians, with a unique nonce and a needle check per prompt, and a swap guard. Swap growth was 0 in every run.

README (second commit): the Mac mini M6 section now uses the same b782 numbers, labelled with the build and linked to the raw data. Rows not re-run on b782 are marked † (earlier build), and the 40.8K Qwen3.6 rows are marked throughput only. The M6 vs M1 Ultra line becomes 78% (48.4 vs 61.7) and notes that the M1 Ultra runs used temperature 0.6 with --repeat-penalty 1.1. The M1 Ultra table itself is unchanged. The --stream-experts crash-warning paragraph is left alone, since #198 also edits it.

AI usage: data collected and PR written by Claude Code (Claude Opus 5.5) in the M6 benchmarking session, with the repo owner's approval to open this PR.

🤖 Generated with Claude Code

solderzzc and others added 2 commits September 27, 2026 20:30
Re-measured on the official b782 binary with scripts/profiling/m6_bench.py:
- Qwen3.6-35B-A3B: full GPU and --stream-experts table at 548 to 40.8K
  tokens, replacing the pre-b782 run, with a note that the 40.8K needle check
  is unreliable in both modes (OSPREY-NN code words), so those rows are
  throughput only.
- Gemma 4 26B-A4B 4-bit (short prompt), Gemma 4 8-bit --stream-experts (533
  and 9.5K) and Qwen3.8-27B (548 and 9.8K): a b782 section on top; the
  existing rows stay, labelled as an earlier build.
- The MTP file is labelled as an earlier build.
New JSONL rows carry "build": "b782"; existing rows are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refresh the Mac mini M6 section to the b782 raw data in docs/profiling/m6:
Qwen3.6 GPU/SSD at all four lengths, Gemma 4 4-bit short prompt, Gemma 4
8-bit --stream-experts, Qwen3.8 at 548 and 9.8K. Rows not re-run on b782 are
marked † (earlier build). The 40.8K Qwen3.6 rows are marked throughput only
(needle check unreliable in both modes). The M6/M1 Ultra line becomes 78%
(48.4 vs 61.7) and notes the M1 Ultra runs used temperature 0.6 with
--repeat-penalty 1.1; the M1 Ultra table is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@solderzzc solderzzc changed the title docs(m6): raw benchmark data from release b782 docs(m6): raw benchmark data and README tables from release b782 Sep 28, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant