Conversation
Re-measured on the official b782 binary with scripts/profiling/m6_bench.py: - Qwen3.6-35B-A3B: full GPU and --stream-experts table at 548 to 40.8K tokens, replacing the pre-b782 run, with a note that the 40.8K needle check is unreliable in both modes (OSPREY-NN code words), so those rows are throughput only. - Gemma 4 26B-A4B 4-bit (short prompt), Gemma 4 8-bit --stream-experts (533 and 9.5K) and Qwen3.8-27B (548 and 9.8K): a b782 section on top; the existing rows stay, labelled as an earlier build. - The MTP file is labelled as an earlier build. New JSONL rows carry "build": "b782"; existing rows are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refresh the Mac mini M6 section to the b782 raw data in docs/profiling/m6: Qwen3.6 GPU/SSD at all four lengths, Gemma 4 4-bit short prompt, Gemma 4 8-bit --stream-experts, Qwen3.8 at 548 and 9.8K. Rows not re-run on b782 are marked † (earlier build). The 40.8K Qwen3.6 rows are marked throughput only (needle check unreliable in both modes). The M6/M1 Ultra line becomes 78% (48.4 vs 61.7) and notes the M1 Ultra runs used temperature 0.6 with --repeat-penalty 1.1; the M1 Ultra table is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Commits the b782 benchmark raw data to
docs/profiling/m6. That folder is linked as "benchmark script and raw JSONL", and the files there came from older builds, so they didn't match the b782 numbers being published (for example Qwen3.6 GPU decode at 548 tokens: 46.97 on main vs 48.35 on b782). Docs only.Changes
qwen36_35b_a3b_4bit.{md,jsonl}: replaced with the full b782 run, GPU and--stream-expertsat 548 / 2.3K / 9.8K / 40.8K tokens. It includes a note that the 40.8K needle check is unreliable in both modes (every miss had anOSPREY-NNcode word, and GPU fails the same way on the same prompt), so the 40.8K rows are throughput only. The two 40.8K re-runs are included with"recheck": true.gemma4_26b_a4b_4bit,gemma4_26b_a4b_8bit,qwen38_27b_4bit: only some lengths were re-run on b782, so each gets a b782 section on top and the existing rows stay, labelled "Earlier build (before b782)".gemma4_26b_a4b_4bit_mtp_bf16.md: labelled as measured on an earlier build..mdheader names the build. New JSONL rows carry"build": "b782", and existing rows are unchanged.Method
Measured on a Mac mini M6, 32 GB, macOS 27.0, using the official
SwiftLM-b782-macos-arm64.tar.gzwithscripts/profiling/m6_bench.py. That's 3 runs (1 at 32K and above) after one warm-up, temperature 0, reporting medians, with a unique nonce and a needle check per prompt, and a swap guard. Swap growth was 0 in every run.README (second commit): the Mac mini M6 section now uses the same b782 numbers, labelled with the build and linked to the raw data. Rows not re-run on b782 are marked † (earlier build), and the 40.8K Qwen3.6 rows are marked throughput only. The M6 vs M1 Ultra line becomes 78% (48.4 vs 61.7) and notes that the M1 Ultra runs used temperature 0.6 with
--repeat-penalty 1.1. The M1 Ultra table itself is unchanged. The--stream-expertscrash-warning paragraph is left alone, since #198 also edits it.AI usage: data collected and PR written by Claude Code (Claude Opus 5.5) in the M6 benchmarking session, with the repo owner's approval to open this PR.
🤖 Generated with Claude Code