Leaderboard

Headline rankings compare compatible PP512 prefill and TG128 decode workloads. Every throughput value keeps its measured token length; nonstandard and older reports remain available in model, chip, and profile results.

Model
Runtime
Quantisation
Chip
Prefill size
Rank by
378 matching configurations
#Model / quantRuntimeChip / backendUserUpdatedShare
1Tinystories Lay8 HS512 HD8 33MReported as ivnle/tinystories-lay8-hs512-hd8-33M · llama · default-q4
Q4
BaseRTApple M5 Pro
Metal
2,068.0
TG128
190,107
PP512
arki05Germany
2Tinystories Lay8 HS512 HD8 33MReported as RichardErkhov/ivnle_-_tinystories-lay8-hs512-hd8-33M-gguf · llama
Q4_0
llama.cppApple M5 Pro
BLAS + Metal
1,443.3
TG128
124,181
PP512
arki05Germany
3Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4
Q4
BaseRTApple M5 Max
Metal
708.3
TG128
34,136
PP512
basecomputeAustralia
4Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q4
BaseRTApple M4 Max
Metal
583.3
TG128 @ 1 ctx
9,998
PP512
basecomputeAustralia
5Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8
Q8
BaseRTApple M5 Max
Metal
549.5
TG128
33,088
PP512
basecomputeAustralia
6Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4
Q4
BaseRTApple M5 Pro
Metal
516.8
TG128
20,778
PP512
arki05Germany
7Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppAMD Radeon RX 7900 XT (RADV NAVI31)
Vulkan
508.8
TG128
26,137
PP512
arki05Germany
8Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q4
BaseRTApple M3 Ultra
Metal
501.7
TG128 @ 1 ctx
12,830
PP512
basecomputeAustralia
9Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppAMD Radeon RX 7900 XT (RADV NAVI31)
Vulkan
479.4
TG128
26,346
PP512
arki05Germany
10Qwen3 0.6BReported as Qwen3-0.6B-cuda-q4, Qwen3-0.6B-cuda-q4mix, Qwen3-0.6B-cuda-q8, Qwen/Qwen3-0.6B · qwen
Q4
BaseRTNVIDIA GB10
CUDA
458.8
TG128 @ 1 ctx
52,750
PP512
basecomputeAustralia
11Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q8
BaseRTApple M4 Max
Metal
455.8
TG128 @ 1 ctx
9,749
PP512
basecomputeAustralia
12Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q8
BaseRTApple M3 Ultra
Metal
440.2
TG128 @ 1 ctx
12,494
PP512
basecomputeAustralia
13Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppApple M5 Max
BLAS + Metal
439.1
TG128
24,365
PP512
basecomputeAustralia
14Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4
Q4
BaseRTApple M5 Max
Metal
415.5
TG128
23,019
PP512
lukasAustralia
15Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTApple M3 Ultra
Metal
410.1
TG128 @ 1 ctx
9,066
PP512
basecomputeAustralia
16Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppApple M5 Max
BLAS + Metal
379.0
TG128
24,983
PP512
basecomputeAustralia
17Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTApple M4 Max
Metal
371.7
TG128 @ 1 ctx
6,058
PP512
basecomputeAustralia
18Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppAMD Radeon RX 7900 XT
ROCm
368.6
TG128
23,972
PP512
arki05Germany
19Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q4
BaseRTApple M4 Max
Metal
361.2
TG128 @ 1 ctx
7,448
PP512
basecomputeAustralia
20Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppApple M5 Pro
BLAS + Metal
358.3
TG128
14,552
PP512
arki05Germany
21Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8
Q8
BaseRTApple M5 Pro
Metal
350.2
TG128
20,537
PP512
arki05Germany
22Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q4
BaseRTApple M3 Ultra
Metal
328.0
TG128 @ 1 ctx
9,903
PP512
basecomputeAustralia
23Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen
Q4
BaseRTApple M4 Max
Metal
318.4
TG128 @ 1 ctx
4,146
PP512
basecomputeAustralia
24Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppAMD Radeon RX 7900 XT
ROCm
313.7
TG128
24,691
PP512
arki05Germany
25Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen
Q4
BaseRTApple M3 Ultra
Metal
307.4
TG128 @ 1 ctx
5,420
PP512
basecomputeAustralia
26Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35
Q4
BaseRTApple M3 Ultra
Metal
306.7
TG128 @ 1 ctx
1,759
PP512
basecomputeAustralia
27Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35
Q4
BaseRTApple M4 Max
Metal
301.7
TG128 @ 1 ctx
1,718
PP512
basecomputeAustralia
28Gemma 3 1B ITReported as gemma-3-1b-it-cuda-q4mix, gemma-3-1b-it-cuda-q8, google/gemma-3-1b-it · gemma3
Q4
BaseRTNVIDIA GB10
CUDA
295.9
TG128 @ 1 ctx
33,841
PP512
basecomputeAustralia
29Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35
Q4
BaseRTApple M4 Max
Metal
295.1
TG128 @ 1 ctx
1,713
PP512
basecomputeAustralia
30Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35
Q4
BaseRTApple M3 Ultra
Metal
293.8
TG128 @ 1 ctx
1,756
PP512
basecomputeAustralia
31Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4
Q4
BaseRTApple M5 Pro
Metal
293.1
TG128
15,066
PP512
arki05Germany
32Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q8
BaseRTApple M4 Pro
Metal
284.7
TG128 @ 1 ctx
4,787
PP512
basecomputeAustralia
33Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppApple M5 Pro
BLAS + Metal
283.0
TG128
14,942
PP512
arki05Germany
34Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q8
BaseRTApple M4 Max
Metal
281.0
TG128 @ 1 ctx
7,527
PP512
basecomputeAustralia
35Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4
Q4
BaseRTApple M1 Pro
Metal
268.8
TG128 @ 1 ctx
2,847
PP512
isuAustralia
36Llama 3.2 1B InstructReported as Llama-3.2-1B-Instruct-cuda-q4mix, Llama-3.2-1B-Instruct-cuda-q8, meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTNVIDIA GB10
CUDA
261.3
TG128 @ 1 ctx
36,647
PP512
basecomputeAustralia
37Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q8
BaseRTApple M3 Ultra
Metal
257.8
TG128 @ 1 ctx
9,953
PP512
basecomputeAustralia
38Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTApple M1 Max
Metal
242.1
TG128 @ 1 ctx
3,372
PP512
basecomputeAustralia
39Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen
Q8
BaseRTApple M3 Ultra
Metal
238.6
TG128 @ 1 ctx
5,698
PP512
basecomputeAustralia
40Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen · default-q4
Q4
BaseRTApple M5 Pro
Metal
235.4
TG128
8,040
PP512
arki05Germany