Apple Silicon Benchmarks
Real speed measurements on Q4_K_M β tokens per second for popular models across all Apple Silicon chips.
π·πΊ Russian version: apple-silicon.ru.md
β Local models Β· Tools comparison
Sources: MLJourney, CraftRigs, MacYou, LLMCheck.
| Chip | 8B tok/s | 14B tok/s | 27β30B tok/s | 70B tok/s |
|---|---|---|---|---|
| M1 8 GB | 12β14 | |||
| M1 16 GB | 18β24 | 11β12 | ||
| M2 16 GB | 22β28 | 12β14 | ||
| M2 Pro 32 GB | 30β38 | 18β20 | β | |
| M2 Max 64 GB | 50β65 | 30β35 | β | 10β12 |
| M3 16 GB | 22β33 | 12β15 | ||
| M3 Pro 36 GB | 35β39 | 20β22 | 12β15 | |
| M3 Max 128 GB | 65β75 | 40β45 | 25β30 | 10β14 |
| M4 16 GB | 25β35 | 14β17 | ||
| M4 Pro 48 GB | 38β50 | 22β26 | 18β25 | 7β11 |
| M4 Max 64 GB | 65β85 | 40β50 | 28β35 | 12β18 |
| M4 Max 128 GB | 75β95 | 45β55 | 30β40 | 15β22 |
| M4 Ultra 192 GB | 100β140 | 60β80 | 40β60 | 25β35 |
= does not fit in RAM / constant swap (0.4β2.2 tok/s). All numbers are generation, not prefill. Prefill is typically 10β30Γ faster. For NVIDIA/Windows/Linux performance will be significantly higher (RTX 4090 does 80β120 tok/s on 8B).
Speed by task (M1 16 GB, Q4_K_M)
| Task | Model | tok/s | Note |
|---|---|---|---|
| Autocomplete (FIM) | Qwen 2.5 Coder 1.5B | 30+ | |
| Main coder | Qwen 2.5 Coder 7B | 22β25 | |
| Fast chat | Qwen 3.5 4B | 28β35 | |
| Balanced chat | Qwen 3.5 9B | 10β13 | |
| Reasoning | DeepSeek R1 Distill 14B | 8β10 |
MLX vs llama.cpp on M4 Pro
M4 Pro 24 GB, Qwen3-Coder-30B MoE (source: asiai.dev):
| Metric | LM Studio (MLX) | Ollama (llama.cpp) | Difference |
|---|---|---|---|
| Throughput | 102 tok/s | 70 tok/s | +46% |
| TTFT | 291 ms | 175 ms | Ollama faster |
| GPU Power | 12.4 W | 15.4 W | -20% |
| Process Memory | 21.4 GB | 41.6 GB | -49% |
MLX backend in Ollama
Since March 2026, Ollama can automatically switch to MLX backend on Mac with 32 GB+ RAM:
- Qwen 3.5-35B-A3B: 58 β 112 tok/s on M5 Max (+93%)
- Works with Qwen models. Llama/Mistral support coming.
| If you want | Go to |
|---|---|
| Choose a model for your Mac | models.md |
| Compare tools for max speed | tools.md |
| How much memory do you need | memory-and-context.md |
| Back to navigation | README.md |
In section: getting-started Β· running-models Β· models Β· catalog Β· quantization Β· memory-and-context Β· tools Β· advanced-setup Β· troubleshooting Β· apple-silicon
Related sections: Zero Level Β· AI Agents Β· Use Cases
Navigation: β Local Models Β· β Back to main Β· π·πΊ Π ΡΡΡΠΊΠΈΠΉ