Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 42 like-for-like pairs.
| Facet | Apple M3 Ultra | Apple M5 Pro |
|---|---|---|
| Model | gemma-3-1b-it-Q4, gemma-3-1b-it-Q8, gemma-4-26B-A4B-it-Q4, gemma-4-26B-A4B-it-Q8, gemma-4-E2B-it-Q4, gemma-4-E2B-it-Q8, gemma-4-E4B-it-Q4, gemma-4-E4B-it-Q8, gpt-oss-120b-MXFP4, gpt-oss-20b-MXFP4, gpt-oss-20b-Q4, gpt-oss-20b-Q8, Llama-3.1-8B-Instruct-Q4, Llama-3.1-8B-Instruct-Q8, Llama-3.2-1B-Instruct-Q4, Llama-3.2-1B-Instruct-Q8, Llama-3.2-3B-Instruct-Q4, Llama-3.2-3B-Instruct-Q8, Mistral-7B-Instruct-v0.3-Q4, Mistral-7B-Instruct-v0.3-Q8, muse-glimmer-30B-kquant-17gb, muse-glimmer-30B-kquant-dynamic, NVIDIA-Nemotron-3-Nano-30B-A3B-Q4, NVIDIA-Nemotron-3-Nano-30B-A3B-Q8, Qwen3-0.6B-Q4, Qwen3-0.6B-Q8, Qwen3-1.7B-Q4, Qwen3-1.7B-Q8, Qwen3-30B-A3B-Instruct-2507-Q4, Qwen3-30B-A3B-Thinking-2507-Q4, Qwen3-30B-A3B-Thinking-2507-Q8, Qwen3-4B-Instruct-2507-Q4, Qwen3-4B-Instruct-2507-Q8, Qwen3-4B-Q4, Qwen3-4B-Q8, Qwen3-4B-Thinking-2507-Q4, Qwen3-4B-Thinking-2507-Q8, Qwen3-8B-Q4, Qwen3-8B-Q8, Qwen3.5-122B-A10B-Q4, Qwen3.5-2B-Base-Q4, Qwen3.5-2B-Base-Q8, Qwen3.5-2B-Q4, Qwen3.5-2B-Q8, Qwen3.5-35B-A3B-Q4, Qwen3.5-35B-A3B-Q8, Qwen3.6-27B-Q4, Qwen3.6-27B-Q8, Qwen3.6-35B-A3B-Q4, Qwen3.6-35B-A3B-Q8, Qwen3.8-27B-Q4, Qwen3.8-27B-Q4-mtp, Qwen3.8-27B-Q8 | gemma-3-1b-it, gemma-4-26B-A4B-it, gemma-4-E2B-it, gemma-4-E4B-it, Gemma-4-E4B-It, gpt-oss-20b, Gpt-Oss-20B, Llama 3.2 3B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Muse-Glimmer-30B, NVIDIA-Nemotron-3-Nano-30B-A3B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen3-4B-Thinking-2507, Qwen3-8B, Qwen3-8B-base-q2, Qwen3-8B-base-q3, Qwen3-8B-base-q5, Qwen3-8B-base-q6, Qwen3.5-2B, Qwen3.5-2B-Base, Qwen3.5-35B-A3B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B, Tinystories Lay8 Hs512 Hd8 33M, tinystories-lay8-hs512-hd8-33M |
| Runtime | BaseRT | BaseRT, llama.cpp |
| Runtime version | BaseRT 0.2.6 | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10809 (5266f24da), llama.cpp b9960 (a935fbffe) |
| Quantisation | mxfp4, passthrough_gguf, Q4, Q8 | BF16, F16, mxfp4, passthrough_gguf, Q2, Q2_K, Q3, Q3_K_M, Q4, Q4_0, Q4_1, Q4_K_M, Q5, Q5_K_M, Q6, Q6_K, Q8, Q8_0 |
| Backend | Metal | BLAS + Metal, Metal |
| Cooldown | off | off, on |
| Conditioning | warmup_only | idle_reset_then_warm, runtime_native_warmup, warmup_only |
| Decode workload | TG128 @ 1 ctx | TG128, TG128 @ 1 ctx |
| Harness schema | basert-benchmark-harness/1 | basert-benchmark-harness/1, computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M3 Ultra relative to Apple M5 Pro.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-8B BaseRTQ4 | 113.6 | 59.7 | 1.90× | 1,357 | 1,761 | 0.77× | 1 / 7 |
Llama-3.2-3B-Instruct BaseRTQ4 | 174.9 | 94.6 | 1.85× | 3,278 | 4,297 | 0.76× | 2 / 5 |
Llama-3.1-8B-Instruct BaseRTQ4 | 99.9 | 48.8 | 2.05× | 1,415 | 1,784 | 0.79× | 2 / 4 |
Llama-3.2-1B-Instruct BaseRTQ4 | 377.7 | 206.3 | 1.83× | 9,061 | 11,673 | 0.78× | 2 / 3 |
Mistral-7B-Instruct-v0.3 BaseRTQ4 | 104.1 | 51.7 | 2.01× | 1,414 | 1,789 | 0.79× | 2 / 2 |
Muse-Glimmer-30B BaseRTpassthrough_gguf | 29.2 | 15.7 | 1.86× | 402 | 480 | 0.84× | 2 / 2 |
Qwen3-0.6B BaseRTQ4 | 501.7 | 512.0 | 0.98× | 12,830 | 20,603 | 0.62× | 1 / 3 |
Qwen3-0.6B BaseRTQ8 | 440.2 | 349.8 | 1.26× | 12,494 | 20,366 | 0.61× | 1 / 3 |
Qwen3.8-27B BaseRTQ4 | 30.0 | 13.5 | 2.22× | 281 | 341 | 0.82× | 2 / 2 |
gemma-3-1b-it BaseRTQ4 | 328.0 | 281.7 | 1.16× | 9,903 | 14,392 | 0.69× | 1 / 2 |
gemma-4-E4B-it BaseRTQ4 | 94.6 | 82.3 | 1.15× | 3,610 | 4,646 | 0.78× | 1 / 2 |
gemma-4-E4B-it BaseRTQ8 | 75.4 | 51.2 | 1.47× | 3,499 | 4,476 | 0.78× | 1 / 2 |
gpt-oss-20b BaseRTQ4 | 185.1 | 110.5 | 1.67× | 2,554 | 1,156 | 2.21× | 1 / 2 |
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 BaseRTQ8 | 102.0 | 66.0 | 1.54× | 2,313 | 1,748 | 1.32× | 2 / 1 |
Qwen3-1.7B BaseRTQ4 | 307.4 | 222.9 | 1.38× | 5,420 | 7,654 | 0.71× | 1 / 2 |
Qwen3-1.7B BaseRTQ8 | 238.6 | 132.4 | 1.80× | 5,698 | 7,655 | 0.74× | 1 / 2 |
Qwen3-30B-A3B-Instruct-2507 BaseRTQ4 | 137.1 | 98.1 | 1.40× | 2,287 | 3,493 | 0.65× | 1 / 2 |
Qwen3-30B-A3B-Thinking-2507 BaseRTQ8 | 95.9 | 65.1 | 1.47× | 2,340 | 3,692 | 0.63× | 2 / 1 |
Qwen3-4B BaseRTQ4 | 165.7 | 105.6 | 1.57× | 2,461 | 3,378 | 0.73× | 1 / 2 |
Qwen3-4B-Instruct-2507 BaseRTQ4 | 153.7 | 89.9 | 1.71× | 2,529 | 3,345 | 0.76× | 1 / 2 |
Qwen3-4B-Instruct-2507 BaseRTQ8 | 118.3 | 63.5 | 1.86× | 2,557 | 3,402 | 0.75× | 1 / 2 |
Qwen3-8B BaseRTQ8 | 72.5 | 34.9 | 2.08× | 1,338 | 1,746 | 0.77× | 1 / 2 |
gemma-3-1b-it BaseRTQ8 | 257.8 | 210.3 | 1.23× | 9,953 | 15,061 | 0.66× | 1 / 1 |
gemma-4-26B-A4B-it BaseRTQ4 | 81.5 | 51.7 | 1.58× | 2,422 | 2,047 | 1.18× | 1 / 1 |
gemma-4-26B-A4B-it BaseRTQ8 | 75.3 | 51.4 | 1.46× | 2,258 | 3,253 | 0.69× | 1 / 1 |
gemma-4-E2B-it BaseRTQ4 | 137.4 | 144.6 | 0.95× | 8,996 | 13,240 | 0.68× | 1 / 1 |
gemma-4-E2B-it BaseRTQ8 | 116.6 | 96.0 | 1.21× | 8,610 | 12,646 | 0.68× | 1 / 1 |
gpt-oss-20b BaseRTQ8 | 175.8 | 100.6 | 1.75× | 2,627 | 1,065 | 2.47× | 1 / 1 |
gpt-oss-20b BaseRTmxfp4 | 163.9 | 103.6 | 1.58× | 2,590 | 1,013 | 2.56× | 1 / 1 |
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 BaseRTQ4 | 171.6 | 100.1 | 1.71× | 2,295 | 2,499 | 0.92× | 1 / 1 |
Qwen3-30B-A3B-Thinking-2507 BaseRTQ4 | 135.3 | 104.4 | 1.30× | 2,416 | 3,694 | 0.65× | 1 / 1 |
Qwen3-4B BaseRTQ8 | 116.8 | 64.7 | 1.81× | 2,441 | 3,412 | 0.72× | 1 / 1 |
Qwen3-4B-Thinking-2507 BaseRTQ4 | 155.6 | 91.6 | 1.70× | 2,582 | 3,363 | 0.77× | 1 / 1 |
Qwen3-4B-Thinking-2507 BaseRTQ8 | 118.2 | 64.7 | 1.83× | 2,559 | 3,414 | 0.75× | 1 / 1 |
Qwen3.5-2B BaseRTQ4 | 306.7 | 217.7 | 1.41× | 1,759 | 2,641 | 0.67× | 1 / 1 |
Qwen3.5-2B BaseRTQ8 | 222.9 | 130.7 | 1.70× | 1,756 | 2,638 | 0.67× | 1 / 1 |
Qwen3.5-2B-Base BaseRTQ4 | 293.8 | 218.2 | 1.35× | 1,756 | 2,645 | 0.66× | 1 / 1 |
Qwen3.5-2B-Base BaseRTQ8 | 222.4 | 131.4 | 1.69× | 1,696 | 2,642 | 0.64× | 1 / 1 |
Qwen3.5-35B-A3B BaseRTQ4 | 133.0 | 113.4 | 1.17× | 947 | 1,304 | 0.73× | 1 / 1 |
Qwen3.6-27B BaseRTQ4 | 36.8 | 16.1 | 2.28× | 281 | 364 | 0.77× | 1 / 1 |
Qwen3.6-27B BaseRTQ8 | 22.1 | 10.6 | 2.10× | 280 | 355 | 0.79× | 1 / 1 |
Qwen3.6-35B-A3B BaseRTQ4 | 133.0 | 113.2 | 1.17× | 922 | 1,275 | 0.72× | 1 / 1 |
Top 25 of 58 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 501.7 TG128 @ 1 ctx | 12,830 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 440.2 TG128 @ 1 ctx | 12,494 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 410.1 TG128 @ 1 ctx | 9,066 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 345.2 TG128 @ 1 ctx | 9,055 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 328.0 TG128 @ 1 ctx | 9,903 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 307.4 TG128 @ 1 ctx | 5,420 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 306.7 TG128 @ 1 ctx | 1,759 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Base-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 293.8 TG128 @ 1 ctx | 1,756 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 257.8 TG128 @ 1 ctx | 9,953 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 238.6 TG128 @ 1 ctx | 5,698 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 222.9 TG128 @ 1 ctx | 1,756 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Base-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 222.4 TG128 @ 1 ctx | 1,696 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 195.4 TG128 @ 1 ctx | 3,241 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 185.1 TG128 @ 1 ctx | 2,554 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 175.8 TG128 @ 1 ctx | 2,627 PP512 | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 171.6 TG128 @ 1 ctx | 2,295 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 165.7 TG128 @ 1 ctx | 2,461 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M3 Ultra Metal | 163.9 TG128 @ 1 ctx | 2,590 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Thinking-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 155.6 TG128 @ 1 ctx | 2,582 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 154.4 TG128 @ 1 ctx | 3,316 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 153.7 TG128 @ 1 ctx | 2,529 PP512 | basecompute | View Benchmark | |
gemma-4-E2B-it-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 137.4 TG128 @ 1 ctx | 8,996 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 137.1 TG128 @ 1 ctx | 2,287 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 135.3 TG128 @ 1 ctx | 2,416 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 133.0 TG128 @ 1 ctx | 922 PP512 | basecompute | View Benchmark |
Top 25 of 129 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
ivnle/tinystories-lay8-hs512-hd8-33M BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | View Benchmark | |
Tinystories Lay8 Hs512 Hd8 33M llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 311.0 TG128 | 18,474 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 270.4 TG128 | 13,717 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 229.1 TG128 | 11,673 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 218.2 TG128 | 2,645 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 217.7 TG128 | 2,641 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 210.3 TG128 | 7,268 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 210.3 TG128 | 15,061 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 206.3 TG128 | 10,528 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 205.3 TG128 | 12,067 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E2B-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 144.6 TG128 | 13,240 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 139.7 TG128 | 8,043 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 131.4 TG128 | 2,642 PP512 | arki05 | View Benchmark |