Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 13 like-for-like pairs.
| Facet | Apple M5 Max | Apple M5 Pro |
|---|---|---|
| Model | gemma-3-1b-it, Gemma-4 12B IT (smart Q4_0, QAT-lossless), gemma-4-26B-A4B-it, gemma-4-E2B-it, gemma-4-E4B-it, gpt-oss-120b, gpt-oss-20b, Models Qwen Qwen3 0.6B, Muse-Glimmer-30B, NVIDIA-Nemotron-3-Nano-30B-A3B, Qwen3 0.6B Instruct, Qwen3-0.6B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B | gemma-3-1b-it, gemma-4-26B-A4B-it, gemma-4-E2B-it, gemma-4-E4B-it, Gemma-4-E4B-It, gpt-oss-20b, Gpt-Oss-20B, Llama 3.2 3B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Muse-Glimmer-30B, NVIDIA-Nemotron-3-Nano-30B-A3B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen3-4B-Thinking-2507, Qwen3-8B, Qwen3-8B-base-q2, Qwen3-8B-base-q3, Qwen3-8B-base-q5, Qwen3-8B-base-q6, Qwen3.5-2B, Qwen3.5-2B-Base, Qwen3.5-35B-A3B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B, Tinystories Lay8 Hs512 Hd8 33M, tinystories-lay8-hs512-hd8-33M |
| Runtime version | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10360 (48d22e295), llama.cpp b10901 (28ff09582) | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10809 (5266f24da), llama.cpp b9960 (a935fbffe) |
| Quantisation | passthrough_gguf, Q4, Q4_0, Q4_K_M, Q8, Q8_0 | BF16, F16, mxfp4, passthrough_gguf, Q2, Q2_K, Q3, Q3_K_M, Q4, Q4_0, Q4_1, Q4_K_M, Q5, Q5_K_M, Q6, Q6_K, Q8, Q8_0 |
| Decode workload | TG128 | TG128, TG128 @ 1 ctx |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Max relative to Apple M5 Pro.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 BaseRTQ4 | 179.2 | 100.1 | 1.79× | 4,942 | 2,499 | 1.98× | 6 / 1 |
Qwen3.8-27B BaseRTQ4 | 21.5 | 13.5 | 1.59× | 590 | 341 | 1.73× | 5 / 2 |
Qwen3-0.6B BaseRTQ4 | 703.4 | 512.0 | 1.37× | 33,557 | 20,603 | 1.63× | 2 / 3 |
gemma-3-1b-it BaseRTQ4 | 414.9 | 281.7 | 1.47× | 21,395 | 14,392 | 1.49× | 2 / 2 |
Qwen3-0.6B BaseRTQ8 | 549.5 | 349.8 | 1.57× | 33,088 | 20,366 | 1.62× | 1 / 3 |
Qwen3-0.6B llama.cppQ4_K_M | 439.1 | 358.1 | 1.23× | 24,365 | 14,509 | 1.68× | 1 / 3 |
gemma-4-E4B-it BaseRTQ4 | 121.1 | 82.3 | 1.47× | 8,156 | 4,646 | 1.76× | 1 / 2 |
gpt-oss-20b BaseRTQ4 | 169.0 | 110.5 | 1.53× | 2,179 | 1,156 | 1.89× | 1 / 2 |
Muse-Glimmer-30B BaseRTpassthrough_gguf | 15.7 | 15.7 | 1.00× | 830 | 480 | 1.73× | 1 / 2 |
gemma-4-26B-A4B-it BaseRTQ4 | 92.0 | 51.7 | 1.78× | 4,055 | 2,047 | 1.98× | 1 / 1 |
gemma-4-E2B-it BaseRTQ4 | 197.5 | 144.6 | 1.37× | 20,099 | 13,240 | 1.52× | 1 / 1 |
Qwen3-0.6B llama.cppQ8_0 | 379.0 | 283.0 | 1.34× | 24,983 | 14,942 | 1.67× | 1 / 1 |
Qwen3.6-27B BaseRTQ4 | 31.3 | 16.1 | 1.94× | 603 | 364 | 1.65× | 1 / 1 |
Top 25 of 27 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Max Metal | 698.5 TG128 | 32,977 PP512 | lukas | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | View Benchmark | |
Models Qwen Qwen3 0.6B llama.cpp b10360 (48d22e295)Q4_K_M | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 415.5 TG128 | 19,770 PP512 | lukas | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 414.2 TG128 | 23,019 PP512 | lukas | View Benchmark | |
Qwen3 0.6B Instruct llama.cpp b10360 (48d22e295)Q8_0 | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | View Benchmark | |
basecompute/gemma-4-E2B-it BaseRT 0.2.4Q4 | Apple M5 Max Metal | 197.5 TG128 | 20,099 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 184.4 TG128 | 4,984 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 183.7 TG128 | 4,941 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 180.8 TG128 | 4,951 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 177.6 TG128 | 4,527 PP512 | basecompute | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 177.0 TG128 | 4,821 PP512 | lukas | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Max Metal | 169.0 TG128 | 2,179 PP512 | lukas | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 121.1 TG128 | 8,156 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.6-35B-A3B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 119.6 TG128 | 1,540 PP512 | lukas | View Benchmark | |
basecompute/gpt-oss-120b BaseRT 0.2.4Q4 | Apple M5 Max Metal | 114.3 TG128 | 1,401 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 109.2 TG128 | 4,943 PP512 | lukas | View Benchmark | |
basecompute/gemma-4-26B-A4B-it BaseRT 0.2.4Q4 | Apple M5 Max Metal | 92.0 TG128 | 4,055 PP512 | lukas | View Benchmark | |
Gemma-4 12B IT (smart Q4_0, QAT-lossless) llama.cpp b10901 (28ff09582)Q4_0 | Apple M5 Max BLAS + Metal | 54.2 TG128 | 1,777 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 34.0 TG128 | 589 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.6-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 31.3 TG128 | 603 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 21.5 TG128 | 590 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 21.5 TG128 | 594 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 20.7 TG128 | 599 PP512 | lukas | View Benchmark |
Top 25 of 129 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
ivnle/tinystories-lay8-hs512-hd8-33M BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | View Benchmark | |
Tinystories Lay8 Hs512 Hd8 33M llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 311.0 TG128 | 18,474 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 270.4 TG128 | 13,717 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 229.1 TG128 | 11,673 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 218.2 TG128 | 2,645 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 217.7 TG128 | 2,641 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 210.3 TG128 | 7,268 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 210.3 TG128 | 15,061 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 206.3 TG128 | 10,528 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 205.3 TG128 | 12,067 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E2B-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 144.6 TG128 | 13,240 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 139.7 TG128 | 8,043 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 131.4 TG128 | 2,642 PP512 | arki05 | View Benchmark |