Leaderboard
Headline rankings compare compatible PP512 prefill and TG128 decode workloads. Every throughput value keeps its measured token length; nonstandard and older reports remain available in model, chip, and profile results.
Model
Runtime
Quantisation
Chip
Prefill size
Rank by
378 matching configurations
| # | Model / quant | Runtime | Chip / backend | User | Updated | Share | ||
|---|---|---|---|---|---|---|---|---|
| 1 | Tinystories Lay8 HS512 HD8 33MReported as ivnle/tinystories-lay8-hs512-hd8-33M · llama · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | ||
| 2 | Tinystories Lay8 HS512 HD8 33MReported as RichardErkhov/ivnle_-_tinystories-lay8-hs512-hd8-33M-gguf · llama Q4_0 | llama.cpp | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | ||
| 3 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | ||
| 4 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q4 | BaseRT | Apple M4 Max Metal | 583.3 TG128 @ 1 ctx | 9,998 PP512 | basecompute | ||
| 5 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8 Q8 | BaseRT | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | ||
| 6 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | ||
| 7 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 508.8 TG128 | 26,137 PP512 | arki05 | ||
| 8 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q4 | BaseRT | Apple M3 Ultra Metal | 501.7 TG128 @ 1 ctx | 12,830 PP512 | basecompute | ||
| 9 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 479.4 TG128 | 26,346 PP512 | arki05 | ||
| 10 | Qwen3 0.6BReported as Qwen3-0.6B-cuda-q4, Qwen3-0.6B-cuda-q4mix, Qwen3-0.6B-cuda-q8, Qwen/Qwen3-0.6B · qwen Q4 | BaseRT | NVIDIA GB10 CUDA | 458.8 TG128 @ 1 ctx | 52,750 PP512 | basecompute | ||
| 11 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q8 | BaseRT | Apple M4 Max Metal | 455.8 TG128 @ 1 ctx | 9,749 PP512 | basecompute | ||
| 12 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q8 | BaseRT | Apple M3 Ultra Metal | 440.2 TG128 @ 1 ctx | 12,494 PP512 | basecompute | ||
| 13 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | ||
| 14 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 415.5 TG128 | 23,019 PP512 | lukas | ||
| 15 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | Apple M3 Ultra Metal | 410.1 TG128 @ 1 ctx | 9,066 PP512 | basecompute | ||
| 16 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | ||
| 17 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | Apple M4 Max Metal | 371.7 TG128 @ 1 ctx | 6,058 PP512 | basecompute | ||
| 18 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT ROCm | 368.6 TG128 | 23,972 PP512 | arki05 | ||
| 19 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q4 | BaseRT | Apple M4 Max Metal | 361.2 TG128 @ 1 ctx | 7,448 PP512 | basecompute | ||
| 20 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,552 PP512 | arki05 | ||
| 21 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8 Q8 | BaseRT | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | ||
| 22 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q4 | BaseRT | Apple M3 Ultra Metal | 328.0 TG128 @ 1 ctx | 9,903 PP512 | basecompute | ||
| 23 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen Q4 | BaseRT | Apple M4 Max Metal | 318.4 TG128 @ 1 ctx | 4,146 PP512 | basecompute | ||
| 24 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | AMD Radeon RX 7900 XT ROCm | 313.7 TG128 | 24,691 PP512 | arki05 | ||
| 25 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen Q4 | BaseRT | Apple M3 Ultra Metal | 307.4 TG128 @ 1 ctx | 5,420 PP512 | basecompute | ||
| 26 | Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35 Q4 | BaseRT | Apple M3 Ultra Metal | 306.7 TG128 @ 1 ctx | 1,759 PP512 | basecompute | ||
| 27 | Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35 Q4 | BaseRT | Apple M4 Max Metal | 301.7 TG128 @ 1 ctx | 1,718 PP512 | basecompute | ||
| 28 | Gemma 3 1B ITReported as gemma-3-1b-it-cuda-q4mix, gemma-3-1b-it-cuda-q8, google/gemma-3-1b-it · gemma3 Q4 | BaseRT | NVIDIA GB10 CUDA | 295.9 TG128 @ 1 ctx | 33,841 PP512 | basecompute | ||
| 29 | Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35 Q4 | BaseRT | Apple M4 Max Metal | 295.1 TG128 @ 1 ctx | 1,713 PP512 | basecompute | ||
| 30 | Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35 Q4 | BaseRT | Apple M3 Ultra Metal | 293.8 TG128 @ 1 ctx | 1,756 PP512 | basecompute | ||
| 31 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | ||
| 32 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q8 | BaseRT | Apple M4 Pro Metal | 284.7 TG128 @ 1 ctx | 4,787 PP512 | basecompute | ||
| 33 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | ||
| 34 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q8 | BaseRT | Apple M4 Max Metal | 281.0 TG128 @ 1 ctx | 7,527 PP512 | basecompute | ||
| 35 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M1 Pro Metal | 268.8 TG128 @ 1 ctx | 2,847 PP512 | isu | ||
| 36 | Llama 3.2 1B InstructReported as Llama-3.2-1B-Instruct-cuda-q4mix, Llama-3.2-1B-Instruct-cuda-q8, meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | NVIDIA GB10 CUDA | 261.3 TG128 @ 1 ctx | 36,647 PP512 | basecompute | ||
| 37 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q8 | BaseRT | Apple M3 Ultra Metal | 257.8 TG128 @ 1 ctx | 9,953 PP512 | basecompute | ||
| 38 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | Apple M1 Max Metal | 242.1 TG128 @ 1 ctx | 3,372 PP512 | basecompute | ||
| 39 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen Q8 | BaseRT | Apple M3 Ultra Metal | 238.6 TG128 @ 1 ctx | 5,698 PP512 | basecompute | ||
| 40 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 |