Warn when --gpu-layers puts MoE layers on the CPU; mark #176 fixed in README - #182
Merged
Merged
Conversation
CPU-resident MoE layers run the quantized expert matmuls on one core (~0.4 tok/s prefill on Gemma 4 26B-A4B, see #176). Point users at --stream-experts instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Member
Author
|
Cross-review (M6 session): no blocking issues. OK to merge once CI is green. Reviewed the full diff (
Verified on a Mac mini M6 (32 GB), branch
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
--gpu-layersputs MoE layers on the CPU. The CPU runs the quantized expert matmuls on a single core, so it is very slow: ~0.4 tok/s prefill and ~3 s per token on Gemma 4 26B-A4B, on both M5 and M6. The warning recommends--stream-expertsinstead.--gpu-layerstimeout (--gpu-layers Ncrashes the server with a Metal GPU timeout on the first request #176) is fixed in chore: bump mlx-swift for compile scoped-device fix (#176) #177. The MoE speed caveat stays, and the note points 32 GB Macs to--stream-expertsfirst.Test plan
--gpu-layers 23prints7 MoE layers will run on the CPU ...🤖 Generated with Claude Code