Hardware4

Apple M5 Pro

122 runs
Best decode · TG128
2,068tok/s
Best prefill · PP512
190,107tok/s

AMD Radeon RX 7900 XT

41 runs
Best decode · TG128
368.6tok/s
Best prefill · PP512
24,691tok/s

AMD Radeon RX 7900 XT (RADV NAVI31)

41 runs
Best decode · TG128
508.8tok/s
Best prefill · PP512
26,346tok/s

Tesla T10/Tesla T10/Tesla T10/Tesla T10

7 runs
Best decode · TG128
157.9tok/s
Best prefill · PP512
3,390tok/s
Badges6
Record1
Pioneer1
Volume1
Coverage3
150 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
40 runs
Sustained contributor
Volume
Model explorer
27 models covered
Coverage
Hardware explorer
4 devices covered
Coverage
Multi-backend
BLAS,MTL + CUDA + ROCm + Vulkan + metal
Coverage
Submissions211
Sort by
Order
Prefill size
211 submissions · PP512 selected
Model / runtime / formatDeviceBackendReport
Ornith-1.5-9B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
104.8 tok/s
TG128
2,935 tok/s
PP512
Gpt-Oss-20B
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
157.9 tok/s
TG128
3,038 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
48.0 tok/s
TG128
1,174 tok/s
PP512
Meta Llama 3.1 8B Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
136.1 tok/s
TG128
3,390 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
124.7 tok/s
TG128
3,087 tok/s
PP512
Qwen3.8-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
46.0 tok/s
TG128
1,132 tok/s
PP512
Qwen3.8-27B
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Tesla T10/Tesla T10/Tesla T10/Tesla T10CUDA
48.1 tok/s
TG128
910 tok/s
PP512
ivnle/tinystories-lay8-hs512-hd8-33M
BaseRTQ4· default-q4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M5 ProMetal
2,068.0 tok/s
TG128
190,107 tok/s
PP512
Tinystories Lay8 Hs512 Hd8 33M
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
1,443.3 tok/s
TG128
124,181 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
181.3 tok/s
TG128
2,827 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
186.7 tok/s
TG128
2,868 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
32.5 tok/s
TG128
783 tok/s
PP512
Qwen3.6-35B-A3B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
124.1 tok/s
TG128
2,597 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
33.7 tok/s
TG128
782 tok/s
PP512
Qwen3-8B
llama.cppBF16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
45.6 tok/s
TG128
1,168 tok/s
PP512
Qwen3.6-27B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
36.8 tok/s
TG128
739 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
184.5 tok/s
TG128
2,474 tok/s
PP512
Gpt-Oss-20B
llama.cppF16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
146.8 tok/s
TG128
3,262 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
198.1 tok/s
TG128
2,837 tok/s
PP512
Gpt-Oss-20B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
195.6 tok/s
TG128
3,292 tok/s
PP512
Qwen3-8B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
67.4 tok/s
TG128
2,253 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
68.6 tok/s
TG128
2,286 tok/s
PP512
Qwen3-8B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
77.6 tok/s
TG128
2,669 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
95.6 tok/s
TG128
4,214 tok/s
PP512
Qwen3-8B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
88.6 tok/s
TG128
2,497 tok/s
PP512
Qwen3-8B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
95.6 tok/s
TG128
2,430 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
97.7 tok/s
TG128
2,512 tok/s
PP512
Qwen3-8B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
104.9 tok/s
TG128
2,640 tok/s
PP512
Qwen3-8B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
105.7 tok/s
TG128
2,629 tok/s
PP512
Qwen3-8B
llama.cppQ4_1
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
115.6 tok/s
TG128
2,902 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
115.3 tok/s
TG128
2,643 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
122.8 tok/s
TG128
4,161 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
118.5 tok/s
TG128
2,625 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
119.8 tok/s
TG128
2,756 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
126.6 tok/s
TG128
4,142 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
122.7 tok/s
TG128
2,743 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
126.7 tok/s
TG128
2,959 tok/s
PP512
Qwen3-8B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
127.3 tok/s
TG128
2,498 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
134.3 tok/s
TG128
4,824 tok/s
PP512
Qwen3-8B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
130.2 tok/s
TG128
2,459 tok/s
PP512
Qwen3-8B
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
146.2 tok/s
TG128
2,563 tok/s
PP512
Llama 3.2 3B Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
167.5 tok/s
TG128
5,846 tok/s
PP512
Qwen3-8B
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
150.7 tok/s
TG128
2,528 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
180.3 tok/s
TG128
4,793 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
185.1 tok/s
TG128
4,783 tok/s
PP512
Llama-3.2-3B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
223.4 tok/s
TG128
5,721 tok/s
PP512
Llama-3.2-3B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
229.7 tok/s
TG128
5,737 tok/s
PP512
Qwen3-0.6B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
479.4 tok/s
TG128
26,346 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
491.9 tok/s
TG128
25,964 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XT (RADV NAVI31)Vulkan
508.8 tok/s
TG128
26,137 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
128.6 tok/s
TG128
2,661 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
128.5 tok/s
TG128
2,820 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
29.3 tok/s
TG128
840 tok/s
PP512
Qwen3.6-35B-A3B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
100.3 tok/s
TG128
2,501 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
29.8 tok/s
TG128
841 tok/s
PP512
Qwen3-8B
llama.cppBF16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
43.5 tok/s
TG128
3,133 tok/s
PP512
Qwen3.6-27B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
29.7 tok/s
TG128
807 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
128.8 tok/s
TG128
2,596 tok/s
PP512
Gpt-Oss-20B
llama.cppF16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
128.8 tok/s
TG128
3,604 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
137.4 tok/s
TG128
2,386 tok/s
PP512
Gpt-Oss-20B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
160.1 tok/s
TG128
3,502 tok/s
PP512
Qwen3-8B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
64.0 tok/s
TG128
3,207 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
66.2 tok/s
TG128
3,292 tok/s
PP512
Qwen3-8B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
74.2 tok/s
TG128
3,241 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
84.3 tok/s
TG128
4,564 tok/s
PP512
Qwen3-8B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
80.1 tok/s
TG128
2,903 tok/s
PP512
Qwen3-8B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
84.0 tok/s
TG128
2,823 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
87.4 tok/s
TG128
2,851 tok/s
PP512
Qwen3-8B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
86.8 tok/s
TG128
2,869 tok/s
PP512
Qwen3-8B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
87.5 tok/s
TG128
2,837 tok/s
PP512
Qwen3-8B
llama.cppQ4_1
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
107.3 tok/s
TG128
2,894 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
93.2 tok/s
TG128
3,076 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
109.3 tok/s
TG128
4,369 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
93.7 tok/s
TG128
3,110 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
97.5 tok/s
TG128
3,150 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
111.0 tok/s
TG128
4,443 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
97.1 tok/s
TG128
3,172 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
117.4 tok/s
TG128
3,233 tok/s
PP512
Qwen3-8B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
91.6 tok/s
TG128
2,911 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
110.8 tok/s
TG128
5,581 tok/s
PP512
Qwen3-8B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
91.6 tok/s
TG128
2,864 tok/s
PP512
Qwen3-8B
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
105.5 tok/s
TG128
2,834 tok/s
PP512
Llama 3.2 3B Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
142.3 tok/s
TG128
7,210 tok/s
PP512
Qwen3-8B
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
112.3 tok/s
TG128
2,798 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
147.1 tok/s
TG128
5,337 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
147.8 tok/s
TG128
5,364 tok/s
PP512
Llama-3.2-3B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
171.9 tok/s
TG128
6,832 tok/s
PP512
Llama-3.2-3B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
172.0 tok/s
TG128
6,980 tok/s
PP512
Qwen3-0.6B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
313.7 tok/s
TG128
24,691 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
367.7 tok/s
TG128
23,718 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
AMD Radeon RX 7900 XTROCm
368.6 tok/s
TG128
23,972 tok/s
PP512
basecompute/Qwen3-30B-A3B-Instruct-2507
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
104.0 tok/s
TG128
3,680 tok/s
PP512
basecompute/gpt-oss-20b
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
110.1 tok/s
TG128
1,156 tok/s
PP512
basecompute/Llama-3.1-8B-Instruct
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
35.6 tok/s
TG128
1,774 tok/s
PP512
basecompute/Llama-3.1-8B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
63.2 tok/s
TG128
1,795 tok/s
PP512
basecompute/Qwen3-8B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
34.5 tok/s
TG128
1,744 tok/s
PP512
basecompute/Qwen3-8B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
60.7 tok/s
TG128
1,755 tok/s
PP512
basecompute/gemma-4-E4B-it
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
51.0 tok/s
TG128
4,462 tok/s
PP512
basecompute/gemma-4-E4B-it
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
81.9 tok/s
TG128
4,636 tok/s
PP512
basecompute/Qwen3-4B-Instruct-2507
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
62.8 tok/s
TG128
3,405 tok/s
PP512
basecompute/Qwen3-4B-Instruct-2507
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
90.4 tok/s
TG128
3,331 tok/s
PP512
basecompute/Llama-3.2-3B-Instruct
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
79.5 tok/s
TG128
4,343 tok/s
PP512
basecompute/Llama-3.2-3B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
106.8 tok/s
TG128
4,297 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
350.2 tok/s
TG128
20,537 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
516.8 tok/s
TG128
20,778 tok/s
PP512
Qwen3-8B-base-q6
BaseRTQ6
This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified.
Apple M5 ProMetal
42.1 tok/s
TG128
1,723 tok/s
PP512
Qwen3-8B-base-q5
BaseRTQ5
This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified.
Apple M5 ProMetal
48.4 tok/s
TG128
1,730 tok/s
PP512
Qwen3-8B-base-q3
BaseRTQ3
This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified.
Apple M5 ProMetal
62.0 tok/s
TG128
1,748 tok/s
PP512
Qwen3-8B-base-q2
BaseRTQ2
This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified.
Apple M5 ProMetal
81.8 tok/s
TG128
1,786 tok/s
PP512
Qwen3.6-35B-A3B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
61.5 tok/s
TG128
1,533 tok/s
PP512
Qwen3.6-35B-A3B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
61.5 tok/s
TG128
1,673 tok/s
PP512
Qwen3.6-35B-A3B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
67.5 tok/s
TG128
1,737 tok/s
PP512
Qwen3.6-27B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
11.9 tok/s
TG128
377 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
14.5 tok/s
TG128
374 tok/s
PP512
Qwen3.6-27B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
15.2 tok/s
TG128
376 tok/s
PP512
Qwen3.6-27B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
16.7 tok/s
TG128
373 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
78.9 tok/s
TG128
1,646 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
84.7 tok/s
TG128
1,644 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
95.2 tok/s
TG128
1,845 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
95.0 tok/s
TG128
1,831 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
97.3 tok/s
TG128
1,708 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
102.7 tok/s
TG128
1,893 tok/s
PP512
Gpt-Oss-20B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
96.3 tok/s
TG128
1,849 tok/s
PP512
Gpt-Oss-20B
llama.cppF16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
66.6 tok/s
TG128
1,852 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
29.6 tok/s
TG128
1,465 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
43.4 tok/s
TG128
1,421 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
55.5 tok/s
TG128
1,423 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
55.6 tok/s
TG128
1,424 tok/s
PP512
Llama-3.1-8B-Instruct
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
58.7 tok/s
TG128
1,538 tok/s
PP512
Qwen3-8B
llama.cppBF16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
18.9 tok/s
TG128
1,376 tok/s
PP512
Qwen3-8B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
29.2 tok/s
TG128
1,448 tok/s
PP512
Qwen3-8B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
34.4 tok/s
TG128
1,473 tok/s
PP512
Qwen3-8B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
39.4 tok/s
TG128
1,419 tok/s
PP512
Qwen3-8B
llama.cppQ6_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
42.6 tok/s
TG128
1,409 tok/s
PP512
Qwen3-8B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
47.6 tok/s
TG128
1,353 tok/s
PP512
Qwen3-8B
llama.cppQ5_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
47.9 tok/s
TG128
1,345 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
53.6 tok/s
TG128
1,396 tok/s
PP512
Qwen3-8B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
55.0 tok/s
TG128
1,401 tok/s
PP512
Qwen3-8B
llama.cppQ4_1
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
51.9 tok/s
TG128
1,535 tok/s
PP512
Qwen3-8B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
60.9 tok/s
TG128
1,414 tok/s
PP512
Qwen3-8B
llama.cppQ3_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
61.5 tok/s
TG128
1,418 tok/s
PP512
Qwen3-8B
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
69.6 tok/s
TG128
1,465 tok/s
PP512
Qwen3-8B
llama.cppQ2_K
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
72.8 tok/s
TG128
1,477 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
47.9 tok/s
TG128
2,346 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
69.4 tok/s
TG128
2,235 tok/s
PP512
Gemma-4-E4B-It
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
72.1 tok/s
TG128
2,263 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
60.8 tok/s
TG128
2,686 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
92.6 tok/s
TG128
2,566 tok/s
PP512
Qwen3-4B-Instruct-2507
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
94.4 tok/s
TG128
2,568 tok/s
PP512
Llama 3.2 3B Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
76.6 tok/s
TG128
3,516 tok/s
PP512
Llama-3.2-3B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
115.2 tok/s
TG128
3,315 tok/s
PP512
Llama-3.2-3B-Instruct
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
117.7 tok/s
TG128
3,318 tok/s
PP512
Qwen3-0.6B
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
283.0 tok/s
TG128
14,942 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
358.1 tok/s
TG128
14,344 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
358.3 tok/s
TG128
14,509 tok/s
PP512
Qwen3-0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProBLAS + Metal
356.2 tok/s
TG128
14,552 tok/s
PP512
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
66.0 tok/s
TG128
1,748 tok/s
PP512
basecompute/Qwen3-30B-A3B-Thinking-2507
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
65.1 tok/s
TG128
3,692 tok/s
PP512
basecompute/Qwen3.8-27B
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
10.6 tok/s
TG128
355 tok/s
PP512
basecompute/Qwen3.6-27B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
10.6 tok/s
TG128
355 tok/s
PP512
basecompute/gemma-4-26B-A4B-it
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
51.4 tok/s
TG128
3,253 tok/s
PP512
basecompute/Qwen3.5-2B-Base
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
131.4 tok/s
TG128
2,642 tok/s
PP512
basecompute/Qwen3.5-2B-Base
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
218.2 tok/s
TG128
2,645 tok/s
PP512
basecompute/Qwen3-4B-Thinking-2507
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
64.7 tok/s
TG128
3,414 tok/s
PP512
basecompute/Qwen3-4B-Thinking-2507
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
91.6 tok/s
TG128
3,363 tok/s
PP512
basecompute/Qwen3-30B-A3B-Thinking-2507
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
104.4 tok/s
TG128
3,694 tok/s
PP512
basecompute/gpt-oss-20b
BaseRTmxfp4· default-bf16
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
103.6 tok/s
TG128
1,013 tok/s
PP512
basecompute/gpt-oss-20b
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
100.6 tok/s
TG128
1,065 tok/s
PP512
basecompute/Qwen3.5-35B-A3B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
113.4 tok/s
TG128
1,304 tok/s
PP512
basecompute/Qwen3.6-35B-A3B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
113.2 tok/s
TG128
1,275 tok/s
PP512
basecompute/Muse-Glimmer-30B
BaseRTpassthrough_gguf· default-q4k-dynamic
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
14.4 tok/s
TG128
495 tok/s
PP512
basecompute/Muse-Glimmer-30B
BaseRTpassthrough_gguf· q4k-17gb
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
17.0 tok/s
TG128
466 tok/s
PP512
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
100.1 tok/s
TG128
2,499 tok/s
PP512
basecompute/gemma-4-26B-A4B-it
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
51.7 tok/s
TG128
2,047 tok/s
PP512
basecompute/Qwen3.8-27B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
16.5 tok/s
TG128
327 tok/s
PP512
basecompute/gemma-3-1b-it
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
270.4 tok/s
TG128
13,717 tok/s
PP512
basecompute/Llama-3.2-3B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
94.6 tok/s
TG128
3,845 tok/s
PP512
basecompute/Llama-3.2-1B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
206.3 tok/s
TG128
10,528 tok/s
PP512
basecompute/Qwen3-1.7B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
125.1 tok/s
TG128
7,267 tok/s
PP512
basecompute/Qwen3-1.7B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
210.3 tok/s
TG128
7,268 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
311.0 tok/s
TG128
18,474 tok/s
PP512
basecompute/Qwen3-30B-A3B-Instruct-2507
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
92.2 tok/s
TG128
3,305 tok/s
PP512
basecompute/Qwen3.6-27B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
16.1 tok/s
TG128
364 tok/s
PP512
basecompute/gpt-oss-20b
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
110.9 tok/s
TG128
1,155 tok/s
PP512
basecompute/Llama-3.1-8B-Instruct
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
35.5 tok/s
TG128
1,773 tok/s
PP512
basecompute/Llama-3.1-8B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
62.0 tok/s
TG128
1,795 tok/s
PP512
basecompute/Qwen3-8B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
35.2 tok/s
TG128
1,748 tok/s
PP512
basecompute/Qwen3-8B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
60.5 tok/s
TG128
1,763 tok/s
PP512
basecompute/Mistral-7B-Instruct-v0.3
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
37.4 tok/s
TG128
1,783 tok/s
PP512
basecompute/Mistral-7B-Instruct-v0.3
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
66.0 tok/s
TG128
1,795 tok/s
PP512
basecompute/gemma-4-E4B-it
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
51.3 tok/s
TG128
4,490 tok/s
PP512
basecompute/gemma-4-E4B-it
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
82.6 tok/s
TG128
4,656 tok/s
PP512
basecompute/gemma-4-E2B-it
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
96.0 tok/s
TG128
12,646 tok/s
PP512
basecompute/gemma-4-E2B-it
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
144.6 tok/s
TG128
13,240 tok/s
PP512
basecompute/Qwen3-4B-Instruct-2507
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
64.3 tok/s
TG128
3,398 tok/s
PP512
basecompute/Qwen3-4B-Instruct-2507
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
89.4 tok/s
TG128
3,359 tok/s
PP512
basecompute/Qwen3-4B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
64.7 tok/s
TG128
3,412 tok/s
PP512
basecompute/Llama-3.2-3B-Instruct
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
79.5 tok/s
TG128
4,344 tok/s
PP512
basecompute/gemma-3-1b-it
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
210.3 tok/s
TG128
15,061 tok/s
PP512
basecompute/Llama-3.2-1B-Instruct
BaseRTQ4· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
205.3 tok/s
TG128
12,067 tok/s
PP512
basecompute/Qwen3.5-2B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
130.7 tok/s
TG128
2,638 tok/s
PP512
basecompute/Qwen3.5-2B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
217.7 tok/s
TG128
2,641 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
512.0 tok/s
TG128
20,603 tok/s
PP512
basecompute/gemma-3-1b-it
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
293.1 tok/s
TG128
15,066 tok/s
PP512
basecompute/Llama-3.2-3B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
106.1 tok/s
TG128
4,291 tok/s
PP512
basecompute/Llama-3.2-1B-Instruct
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
229.1 tok/s
TG128
11,673 tok/s
PP512
basecompute/Qwen3-1.7B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
139.7 tok/s
TG128
8,043 tok/s
PP512
basecompute/Qwen3-1.7B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
235.4 tok/s
TG128
8,040 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
349.8 tok/s
TG128
20,366 tok/s
PP512
basecompute/Qwen3-4B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
103.9 tok/s
TG128
3,372 tok/s
PP512
basecompute/Qwen3-4B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 ProMetal
107.3 tok/s
TG128
3,384 tok/s
PP512