Hardware6

NVIDIA GB10

71 runs
Best decode · TG128 @ 1 ctx
458.8tok/s
Best prefill · PP512
52,750tok/s

Apple M3 Ultra

58 runs
Best decode · TG128 @ 1 ctx
501.7tok/s
Best prefill · PP512
12,830tok/s

Apple M1 Max

56 runs
Best decode · TG128 @ 1 ctx
242.1tok/s
Best prefill · PP512
4,521tok/s

Apple M4 Max

44 runs
Best decode · TG128 @ 1 ctx
583.3tok/s
Best prefill · PP512
9,998tok/s

Apple M4 Pro

33 runs
Best decode · TG128 @ 1 ctx
284.7tok/s
Best prefill · PP512
4,787tok/s

Apple M5 Max

7 runs
Best decode · TG128
708.3tok/s
Best prefill · PP512
34,136tok/s
Badges6
Record1
Pioneer1
Volume1
Coverage3
201 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
40 runs
Sustained contributor
Volume
Model explorer
28 models covered
Coverage
Hardware explorer
6 devices covered
Coverage
Multi-backend
BLAS,MTL + cuda + metal
Coverage
Submissions269
Sort by
Order
Prefill size
269 submissions · PP512 selected
Model / runtime / formatDeviceBackendReport
gemma-4-26B-A4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
73.8 tok/s
TG128 @ 1 ctx
1,679 tok/s
PP512
gpt-oss-20b-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
185.1 tok/s
TG128 @ 1 ctx
2,554 tok/s
PP512
gpt-oss-20b-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
175.8 tok/s
TG128 @ 1 ctx
2,627 tok/s
PP512
Qwen3.8-27B-Q4-mtp
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M3 UltraMetal
39.5 tok/s
TG128 @ 1 ctx
288 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
171.6 tok/s
TG128 @ 1 ctx
2,295 tok/s
PP512
muse-glimmer-30B-kquant-dynamic
BaseRTpassthrough_gguf
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
27.3 tok/s
TG128 @ 1 ctx
399 tok/s
PP512
muse-glimmer-30B-kquant-dynamic
BaseRTpassthrough_gguf
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
22.8 tok/s
TG128 @ 1 ctx
254 tok/s
PP512
Qwen3.6-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
139.7 tok/s
TG128 @ 1 ctx
886 tok/s
PP512
Qwen3.5-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
140.1 tok/s
TG128 @ 1 ctx
895 tok/s
PP512
muse-glimmer-30B-kquant-17gb
BaseRTpassthrough_gguf
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
27.2 tok/s
TG128 @ 1 ctx
266 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
169.1 tok/s
TG128 @ 1 ctx
1,632 tok/s
PP512
Llama-3.2-1B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
204.6 tok/s
TG128 @ 1 ctx
2,685 tok/s
PP512
Qwen3-1.7B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
193.1 tok/s
TG128 @ 1 ctx
1,828 tok/s
PP512
Qwen3.5-2B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
179.1 tok/s
TG128 @ 1 ctx
1,150 tok/s
PP512
Qwen3.5-2B-Base-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
176.8 tok/s
TG128 @ 1 ctx
1,150 tok/s
PP512
Qwen3.8-27B-Q4-mtp
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M4 MaxMetal
31.5 tok/s
TG128 @ 1 ctx
221 tok/s
PP512
Llama-3.2-1B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
179.3 tok/s
TG128 @ 1 ctx
2,692 tok/s
PP512
Qwen3-1.7B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
125.3 tok/s
TG128 @ 1 ctx
1,822 tok/s
PP512
Qwen3-4B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
90.3 tok/s
TG128 @ 1 ctx
734 tok/s
PP512
Qwen3.5-2B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
113.5 tok/s
TG128 @ 1 ctx
1,138 tok/s
PP512
Qwen3.5-2B-Base-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
113.9 tok/s
TG128 @ 1 ctx
1,149 tok/s
PP512
Llama-3.2-3B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
95.0 tok/s
TG128 @ 1 ctx
943 tok/s
PP512
Qwen3-4B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
77.9 tok/s
TG128 @ 1 ctx
727 tok/s
PP512
Qwen3-4B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
78.4 tok/s
TG128 @ 1 ctx
732 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
139.8 tok/s
TG128 @ 1 ctx
1,789 tok/s
PP512
gemma-4-E2B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
126.0 tok/s
TG128 @ 1 ctx
3,722 tok/s
PP512
Llama-3.2-3B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
71.2 tok/s
TG128 @ 1 ctx
943 tok/s
PP512
gemma-3-1b-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
328.0 tok/s
TG128 @ 1 ctx
9,903 tok/s
PP512
Llama-3.2-1B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
410.1 tok/s
TG128 @ 1 ctx
9,066 tok/s
PP512
gemma-3-1b-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
257.8 tok/s
TG128 @ 1 ctx
9,953 tok/s
PP512
Qwen3.5-2B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
306.7 tok/s
TG128 @ 1 ctx
1,759 tok/s
PP512
Qwen3-1.7B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
307.4 tok/s
TG128 @ 1 ctx
5,420 tok/s
PP512
Qwen3.5-2B-Base-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
293.8 tok/s
TG128 @ 1 ctx
1,756 tok/s
PP512
Llama-3.2-1B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
345.2 tok/s
TG128 @ 1 ctx
9,055 tok/s
PP512
Qwen3-1.7B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
238.6 tok/s
TG128 @ 1 ctx
5,698 tok/s
PP512
Qwen3-4B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
165.7 tok/s
TG128 @ 1 ctx
2,461 tok/s
PP512
Qwen3.5-2B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
222.9 tok/s
TG128 @ 1 ctx
1,756 tok/s
PP512
Qwen3.5-2B-Base-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
222.4 tok/s
TG128 @ 1 ctx
1,696 tok/s
PP512
Qwen3-4B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
56.3 tok/s
TG128 @ 1 ctx
726 tok/s
PP512
Llama-3.2-3B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
195.4 tok/s
TG128 @ 1 ctx
3,241 tok/s
PP512
gemma-3-1b-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
170.0 tok/s
TG128 @ 1 ctx
4,120 tok/s
PP512
Qwen3-4B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
153.7 tok/s
TG128 @ 1 ctx
2,529 tok/s
PP512
Llama-3.2-1B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
242.1 tok/s
TG128 @ 1 ctx
3,333 tok/s
PP512
Qwen3-4B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
155.6 tok/s
TG128 @ 1 ctx
2,582 tok/s
PP512
Qwen3-1.7B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
219.7 tok/s
TG128 @ 1 ctx
2,244 tok/s
PP512
gemma-4-E2B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
137.4 tok/s
TG128 @ 1 ctx
8,996 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
140.1 tok/s
TG128 @ 1 ctx
1,797 tok/s
PP512
Qwen3.5-2B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
192.9 tok/s
TG128 @ 1 ctx
429 tok/s
PP512
Llama-3.2-3B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
154.4 tok/s
TG128 @ 1 ctx
3,316 tok/s
PP512
Qwen3-4B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
116.8 tok/s
TG128 @ 1 ctx
2,441 tok/s
PP512
Qwen3-4B-Instruct-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
118.3 tok/s
TG128 @ 1 ctx
2,557 tok/s
PP512
Qwen3-4B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
118.2 tok/s
TG128 @ 1 ctx
2,559 tok/s
PP512
Qwen3.5-2B-Base-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
190.0 tok/s
TG128 @ 1 ctx
439 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
129.3 tok/s
TG128 @ 1 ctx
1,440 tok/s
PP512
gemma-4-E4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
94.6 tok/s
TG128 @ 1 ctx
3,610 tok/s
PP512
Qwen3-8B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
113.6 tok/s
TG128 @ 1 ctx
1,357 tok/s
PP512
Qwen3-1.7B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
153.9 tok/s
TG128 @ 1 ctx
2,223 tok/s
PP512
Llama-3.2-1B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
229.4 tok/s
TG128 @ 1 ctx
3,372 tok/s
PP512
Llama-3.1-8B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
124.4 tok/s
TG128 @ 1 ctx
1,438 tok/s
PP512
Qwen3-4B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
101.9 tok/s
TG128 @ 1 ctx
927 tok/s
PP512
Qwen3-4B-Instruct-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
56.2 tok/s
TG128 @ 1 ctx
731 tok/s
PP512
gemma-4-E2B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
116.6 tok/s
TG128 @ 1 ctx
8,610 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
78.9 tok/s
TG128 @ 1 ctx
1,388 tok/s
PP512
Qwen3.5-2B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
141.5 tok/s
TG128 @ 1 ctx
400 tok/s
PP512
Llama-3.1-8B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
75.3 tok/s
TG128 @ 1 ctx
1,392 tok/s
PP512
gemma-4-E4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
75.4 tok/s
TG128 @ 1 ctx
3,499 tok/s
PP512
Qwen3.5-2B-Base-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
140.7 tok/s
TG128 @ 1 ctx
438 tok/s
PP512
Qwen3-8B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
72.5 tok/s
TG128 @ 1 ctx
1,338 tok/s
PP512
gpt-oss-20b-MXFP4
BaseRTmxfp4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
163.9 tok/s
TG128 @ 1 ctx
2,590 tok/s
PP512
Llama-3.2-3B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
110.7 tok/s
TG128 @ 1 ctx
1,229 tok/s
PP512
Qwen3.6-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
36.8 tok/s
TG128 @ 1 ctx
281 tok/s
PP512
Qwen3-4B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
91.3 tok/s
TG128 @ 1 ctx
922 tok/s
PP512
gemma-4-26B-A4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
81.5 tok/s
TG128 @ 1 ctx
1,664 tok/s
PP512
Qwen3-4B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
56.3 tok/s
TG128 @ 1 ctx
731 tok/s
PP512
Qwen3-4B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
93.2 tok/s
TG128 @ 1 ctx
921 tok/s
PP512
gemma-4-26B-A4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
81.5 tok/s
TG128 @ 1 ctx
2,422 tok/s
PP512
gemma-4-E2B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
120.8 tok/s
TG128 @ 1 ctx
4,521 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
40.6 tok/s
TG128 @ 1 ctx
509 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
135.3 tok/s
TG128 @ 1 ctx
2,416 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
137.1 tok/s
TG128 @ 1 ctx
2,287 tok/s
PP512
gemma-4-E2B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
91.1 tok/s
TG128 @ 1 ctx
4,442 tok/s
PP512
muse-glimmer-30B-kquant-17gb
BaseRTpassthrough_gguf
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
31.1 tok/s
TG128 @ 1 ctx
405 tok/s
PP512
Llama-3.1-8B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
63.7 tok/s
TG128 @ 1 ctx
509 tok/s
PP512
gemma-4-E4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
69.1 tok/s
TG128 @ 1 ctx
1,094 tok/s
PP512
Qwen3-4B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
71.5 tok/s
TG128 @ 1 ctx
925 tok/s
PP512
Qwen3.5-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
133.0 tok/s
TG128 @ 1 ctx
947 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
59.0 tok/s
TG128 @ 1 ctx
394 tok/s
PP512
Qwen3.6-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
133.0 tok/s
TG128 @ 1 ctx
922 tok/s
PP512
Qwen3-4B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
71.3 tok/s
TG128 @ 1 ctx
920 tok/s
PP512
gemma-4-26B-A4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
75.3 tok/s
TG128 @ 1 ctx
2,258 tok/s
PP512
Llama-3.2-3B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
88.4 tok/s
TG128 @ 1 ctx
1,240 tok/s
PP512
gemma-3-1b-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
178.8 tok/s
TG128 @ 1 ctx
3,337 tok/s
PP512
Qwen3-4B-Instruct-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
72.1 tok/s
TG128 @ 1 ctx
925 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
66.4 tok/s
TG128 @ 1 ctx
511 tok/s
PP512
Qwen3.6-27B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
22.1 tok/s
TG128 @ 1 ctx
280 tok/s
PP512
gemma-4-E4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
75.9 tok/s
TG128 @ 1 ctx
1,394 tok/s
PP512
Qwen3-8B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
52.6 tok/s
TG128 @ 1 ctx
394 tok/s
PP512
Qwen3-8B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
60.2 tok/s
TG128 @ 1 ctx
507 tok/s
PP512
Qwen3.8-27B-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
22.1 tok/s
TG128 @ 1 ctx
280 tok/s
PP512
Qwen3.6-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
30.3 tok/s
TG128 @ 1 ctx
221 tok/s
PP512
Llama-3.1-8B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
38.6 tok/s
TG128 @ 1 ctx
509 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
95.9 tok/s
TG128 @ 1 ctx
2,289 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
95.8 tok/s
TG128 @ 1 ctx
2,391 tok/s
PP512
gemma-4-E4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
52.0 tok/s
TG128 @ 1 ctx
1,368 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
102.1 tok/s
TG128 @ 1 ctx
2,311 tok/s
PP512
Qwen3-8B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
38.1 tok/s
TG128 @ 1 ctx
500 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
101.9 tok/s
TG128 @ 1 ctx
2,315 tok/s
PP512
Llama-3.1-8B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
56.1 tok/s
TG128 @ 1 ctx
394 tok/s
PP512
Qwen3.5-35B-A3B-Q8
BaseRTQ8
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M3 UltraMetal
109.6 tok/s
TG128 @ 1 ctx
927 tok/s
PP512
gpt-oss-20b-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
82.3 tok/s
TG128 @ 1 ctx
881 tok/s
PP512
Qwen3.5-35B-A3B-Q8
BaseRTQ8
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M3 UltraMetal
109.1 tok/s
TG128 @ 1 ctx
939 tok/s
PP512
gpt-oss-20b-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
76.7 tok/s
TG128 @ 1 ctx
880 tok/s
PP512
Qwen3.6-35B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
110.8 tok/s
TG128 @ 1 ctx
910 tok/s
PP512
gpt-oss-20b-MXFP4
BaseRTmxfp4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
76.6 tok/s
TG128 @ 1 ctx
890 tok/s
PP512
Qwen3.6-35B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
109.5 tok/s
TG128 @ 1 ctx
913 tok/s
PP512
Qwen3.6-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
17.9 tok/s
TG128 @ 1 ctx
80 tok/s
PP512
gpt-oss-120b-MXFP4
BaseRTmxfp4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M3 UltraMetal
114.0 tok/s
TG128 @ 1 ctx
1,700 tok/s
PP512
gpt-oss-20b-MXFP4
BaseRTmxfp4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
146.7 tok/s
TG128 @ 1 ctx
1,715 tok/s
PP512
gemma-4-E2B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
83.9 tok/s
TG128 @ 1 ctx
3,628 tok/s
PP512
Qwen3.5-122B-A10B-Q4
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M3 UltraMetal
66.4 tok/s
TG128 @ 1 ctx
536 tok/s
PP512
Qwen3.5-122B-A10B-Q4
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M3 UltraMetal
66.3 tok/s
TG128 @ 1 ctx
532 tok/s
PP512
gemma-4-26B-A4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
51.5 tok/s
TG128 @ 1 ctx
871 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
33.8 tok/s
TG128 @ 1 ctx
393 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
87.2 tok/s
TG128 @ 1 ctx
901 tok/s
PP512
gpt-oss-20b-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
146.9 tok/s
TG128 @ 1 ctx
1,700 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
87.7 tok/s
TG128 @ 1 ctx
901 tok/s
PP512
Qwen3.8-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
38.0 tok/s
TG128 @ 1 ctx
282 tok/s
PP512
Qwen3.8-27B-Q4-mtp
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M1 MaxMetal
18.2 tok/s
TG128 @ 1 ctx
82 tok/s
PP512
gpt-oss-20b-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
160.5 tok/s
TG128 @ 1 ctx
1,709 tok/s
PP512
Llama-3.1-8B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
32.1 tok/s
TG128 @ 1 ctx
392 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
110.3 tok/s
TG128 @ 1 ctx
870 tok/s
PP512
muse-glimmer-30B-kquant-17gb
BaseRTpassthrough_gguf
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
17.3 tok/s
TG128 @ 1 ctx
140 tok/s
PP512
Qwen3-8B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
59.6 tok/s
TG128 @ 1 ctx
942 tok/s
PP512
Qwen3.5-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
93.7 tok/s
TG128 @ 1 ctx
280 tok/s
PP512
Qwen3.6-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
94.2 tok/s
TG128 @ 1 ctx
278 tok/s
PP512
gemma-4-E4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
44.5 tok/s
TG128 @ 1 ctx
1,081 tok/s
PP512
gemma-4-E4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
77.4 tok/s
TG128 @ 1 ctx
2,481 tok/s
PP512
muse-glimmer-30B-kquant-dynamic
BaseRTpassthrough_gguf
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
14.3 tok/s
TG128 @ 1 ctx
136 tok/s
PP512
Llama-3.1-8B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
60.8 tok/s
TG128 @ 1 ctx
947 tok/s
PP512
gemma-4-26B-A4B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
45.5 tok/s
TG128 @ 1 ctx
877 tok/s
PP512
Qwen3.6-27B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
11.0 tok/s
TG128 @ 1 ctx
83 tok/s
PP512
Qwen3-8B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
31.6 tok/s
TG128 @ 1 ctx
392 tok/s
PP512
Qwen3.5-122B-A10B-Q4
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
30.6 tok/s
TG128 @ 1 ctx
799 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
64.1 tok/s
TG128 @ 1 ctx
946 tok/s
PP512
gpt-oss-120b-MXFP4
BaseRTmxfp4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
41.4 tok/s
TG128 @ 1 ctx
2,622 tok/s
PP512
Qwen3.8-27B-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
11.0 tok/s
TG128 @ 1 ctx
82 tok/s
PP512
gemma-4-E2B-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
132.2 tok/s
TG128 @ 1 ctx
7,448 tok/s
PP512
Qwen3.5-122B-A10B-Q4
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
31.2 tok/s
TG128 @ 1 ctx
795 tok/s
PP512
gpt-oss-120b-MXFP4
BaseRTmxfp4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
41.7 tok/s
TG128 @ 1 ctx
2,626 tok/s
PP512
Llama-3.1-8B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
105.5 tok/s
TG128 @ 1 ctx
946 tok/s
PP512
gpt-oss-120b-Q8
BaseRTQ8
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
35.8 tok/s
TG128 @ 1 ctx
2,638 tok/s
PP512
gpt-oss-20b-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
79.3 tok/s
TG128 @ 1 ctx
732 tok/s
PP512
Qwen3-8B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
99.2 tok/s
TG128 @ 1 ctx
943 tok/s
PP512
Qwen3.8-27B-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
11.0 tok/s
TG128 @ 1 ctx
82 tok/s
PP512
gemma-4-E4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
109.3 tok/s
TG128 @ 1 ctx
2,552 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
110.3 tok/s
TG128 @ 1 ctx
947 tok/s
PP512
gpt-oss-120b-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
39.7 tok/s
TG128 @ 1 ctx
2,619 tok/s
PP512
gpt-oss-120b-Q8
BaseRTQ8
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
35.5 tok/s
TG128 @ 1 ctx
2,278 tok/s
PP512
Qwen3-4B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
104.3 tok/s
TG128 @ 1 ctx
1,714 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
64.6 tok/s
TG128 @ 1 ctx
893 tok/s
PP512
gpt-oss-20b-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
73.7 tok/s
TG128 @ 1 ctx
731 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
63.7 tok/s
TG128 @ 1 ctx
895 tok/s
PP512
Qwen3-4B-Instruct-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
104.3 tok/s
TG128 @ 1 ctx
1,723 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
61.1 tok/s
TG128 @ 1 ctx
857 tok/s
PP512
Qwen3.6-35B-A3B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
55.4 tok/s
TG128 @ 1 ctx
2,181 tok/s
PP512
Qwen3.5-35B-A3B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
54.7 tok/s
TG128 @ 1 ctx
2,447 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
60.6 tok/s
TG128 @ 1 ctx
858 tok/s
PP512
Qwen3-4B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
105.4 tok/s
TG128 @ 1 ctx
1,712 tok/s
PP512
Qwen3.5-35B-A3B-Q8
BaseRTQ8
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M1 MaxMetal
72.0 tok/s
TG128 @ 1 ctx
250 tok/s
PP512
Qwen3.6-35B-A3B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
54.4 tok/s
TG128 @ 1 ctx
2,188 tok/s
PP512
Qwen3.5-35B-A3B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
54.2 tok/s
TG128 @ 1 ctx
2,428 tok/s
PP512
Llama-3.2-3B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
133.1 tok/s
TG128 @ 1 ctx
2,218 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
61.2 tok/s
TG128 @ 1 ctx
7,253 tok/s
PP512
gemma-4-E2B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
175.3 tok/s
TG128 @ 1 ctx
7,670 tok/s
PP512
Qwen3.5-35B-A3B-Q8
BaseRTQ8
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
Apple M1 MaxMetal
72.2 tok/s
TG128 @ 1 ctx
251 tok/s
PP512
Qwen3.8-27B-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
7.5 tok/s
TG128 @ 1 ctx
1,127 tok/s
PP512
Qwen3.6-27B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
8.4 tok/s
TG128 @ 1 ctx
1,146 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
49.4 tok/s
TG128 @ 1 ctx
7,315 tok/s
PP512
Qwen3-4B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
122.7 tok/s
TG128 @ 1 ctx
1,721 tok/s
PP512
Qwen3.8-27B-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
7.6 tok/s
TG128 @ 1 ctx
1,152 tok/s
PP512
Qwen3.6-27B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
7.5 tok/s
TG128 @ 1 ctx
1,141 tok/s
PP512
Qwen3-4B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
121.9 tok/s
TG128 @ 1 ctx
1,711 tok/s
PP512
Qwen3.6-35B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
70.8 tok/s
TG128 @ 1 ctx
247 tok/s
PP512
gpt-oss-20b-MXFP4
BaseRTmxfp4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
72.7 tok/s
TG128 @ 1 ctx
734 tok/s
PP512
gemma-4-26B-A4B-it-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
40.4 tok/s
TG128 @ 1 ctx
6,045 tok/s
PP512
Llama-3.2-3B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
178.8 tok/s
TG128 @ 1 ctx
2,202 tok/s
PP512
Qwen3.6-35B-A3B-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
72.8 tok/s
TG128 @ 1 ctx
2,279 tok/s
PP512
Qwen3.5-2B-Base-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
208.0 tok/s
TG128 @ 1 ctx
1,727 tok/s
PP512
Qwen3.5-35B-A3B-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
92.4 tok/s
TG128 @ 1 ctx
2,521 tok/s
PP512
Qwen3.6-35B-A3B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
72.2 tok/s
TG128 @ 1 ctx
232 tok/s
PP512
Qwen3.6-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
84.3 tok/s
TG128 @ 1 ctx
2,028 tok/s
PP512
Qwen3.6-27B-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
12.3 tok/s
TG128 @ 1 ctx
1,127 tok/s
PP512
Qwen3.5-35B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
84.4 tok/s
TG128 @ 1 ctx
2,093 tok/s
PP512
Qwen3.5-2B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
205.4 tok/s
TG128 @ 1 ctx
1,721 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
99.0 tok/s
TG128 @ 1 ctx
7,439 tok/s
PP512
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
99.8 tok/s
TG128 @ 1 ctx
3,862 tok/s
PP512
Qwen3.8-27B-Q4-mtp
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
13.1 tok/s
TG128 @ 1 ctx
1,136 tok/s
PP512
Qwen3.8-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
13.1 tok/s
TG128 @ 1 ctx
1,131 tok/s
PP512
Qwen3-4B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
157.5 tok/s
TG128 @ 1 ctx
1,725 tok/s
PP512
Qwen3-30B-A3B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
80.7 tok/s
TG128 @ 1 ctx
4,715 tok/s
PP512
Qwen3.6-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
13.0 tok/s
TG128 @ 1 ctx
344 tok/s
PP512
Qwen3-30B-A3B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
81.1 tok/s
TG128 @ 1 ctx
4,738 tok/s
PP512
Qwen3.8-27B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
18.2 tok/s
TG128 @ 1 ctx
82 tok/s
PP512
Qwen3-1.7B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
224.1 tok/s
TG128 @ 1 ctx
4,025 tok/s
PP512
gemma-4-26B-A4B-it-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
45.3 tok/s
TG128 @ 1 ctx
6,190 tok/s
PP512
gemma-4-26B-A4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
44.0 tok/s
TG128 @ 1 ctx
6,179 tok/s
PP512
gpt-oss-20b-MXFP4
BaseRTmxfp4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
60.2 tok/s
TG128 @ 1 ctx
4,424 tok/s
PP512
Llama-3.2-1B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
320.5 tok/s
TG128 @ 1 ctx
6,058 tok/s
PP512
gpt-oss-20b-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
51.4 tok/s
TG128 @ 1 ctx
4,421 tok/s
PP512
gpt-oss-20b-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
57.3 tok/s
TG128 @ 1 ctx
4,292 tok/s
PP512
Qwen3.5-2B-Base-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
295.1 tok/s
TG128 @ 1 ctx
1,713 tok/s
PP512
Llama-3.1-8B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
25.4 tok/s
TG128 @ 1 ctx
7,299 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
27.0 tok/s
TG128 @ 1 ctx
7,510 tok/s
PP512
gemma-4-E2B-it-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
87.7 tok/s
TG128 @ 1 ctx
19,246 tok/s
PP512
Qwen3.5-2B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
301.7 tok/s
TG128 @ 1 ctx
1,718 tok/s
PP512
Llama-3.1-8B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
53.1 tok/s
TG128 @ 1 ctx
7,567 tok/s
PP512
Qwen3-8B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
50.3 tok/s
TG128 @ 1 ctx
1,452 tok/s
PP512
gemma-4-E4B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
71.3 tok/s
TG128 @ 1 ctx
3,674 tok/s
PP512
Mistral-7B-Instruct-v0.3-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
53.5 tok/s
TG128 @ 1 ctx
7,554 tok/s
PP512
gemma-4-E2B-it-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
79.0 tok/s
TG128 @ 1 ctx
18,372 tok/s
PP512
Qwen3-1.7B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
318.4 tok/s
TG128 @ 1 ctx
4,146 tok/s
PP512
Qwen3-4B-Instruct-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
54.8 tok/s
TG128 @ 1 ctx
10,970 tok/s
PP512
Qwen3-4B-Thinking-2507-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
54.7 tok/s
TG128 @ 1 ctx
11,473 tok/s
PP512
Llama-3.2-3B-Instruct-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
69.4 tok/s
TG128 @ 1 ctx
14,988 tok/s
PP512
Llama-3.2-3B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
68.7 tok/s
TG128 @ 1 ctx
14,871 tok/s
PP512
gemma-4-E2B-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
139.8 tok/s
TG128 @ 1 ctx
12,036 tok/s
PP512
Llama-3.2-1B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
371.7 tok/s
TG128 @ 1 ctx
5,963 tok/s
PP512
Qwen3-4B-Thinking-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
73.2 tok/s
TG128 @ 1 ctx
11,020 tok/s
PP512
Qwen3-4B-Instruct-2507-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
73.7 tok/s
TG128 @ 1 ctx
11,246 tok/s
PP512
Llama-3.2-3B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
89.7 tok/s
TG128 @ 1 ctx
14,073 tok/s
PP512
Qwen3-4B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
91.3 tok/s
TG128 @ 1 ctx
2,822 tok/s
PP512
Qwen3.5-2B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
113.5 tok/s
TG128 @ 1 ctx
6,754 tok/s
PP512
gemma-3-1b-it-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
281.0 tok/s
TG128 @ 1 ctx
7,527 tok/s
PP512
Llama-3.2-3B-Instruct-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
110.5 tok/s
TG128 @ 1 ctx
14,651 tok/s
PP512
Qwen3.5-2B-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
151.5 tok/s
TG128 @ 1 ctx
6,837 tok/s
PP512
Llama-3.2-1B-Instruct-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
173.7 tok/s
TG128 @ 1 ctx
34,296 tok/s
PP512
Llama-3.2-1B-Instruct-Q8
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
172.8 tok/s
TG128 @ 1 ctx
35,941 tok/s
PP512
Qwen3.5-2B-Base-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
181.9 tok/s
TG128 @ 1 ctx
4,191 tok/s
PP512
Qwen3.5-2B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
181.6 tok/s
TG128 @ 1 ctx
4,185 tok/s
PP512
Qwen3-1.7B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
202.5 tok/s
TG128 @ 1 ctx
7,394 tok/s
PP512
gemma-3-1b-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
361.2 tok/s
TG128 @ 1 ctx
7,448 tok/s
PP512
Llama-3.2-1B-Instruct-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
195.3 tok/s
TG128 @ 1 ctx
33,397 tok/s
PP512
gemma-3-1b-it-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
193.2 tok/s
TG128 @ 1 ctx
33,641 tok/s
PP512
Llama-3.2-1B-Instruct-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
261.3 tok/s
TG128 @ 1 ctx
36,647 tok/s
PP512
gemma-3-1b-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
224.9 tok/s
TG128 @ 1 ctx
3,430 tok/s
PP512
gemma-3-1b-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
220.0 tok/s
TG128 @ 1 ctx
4,102 tok/s
PP512
gemma-3-1b-it-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
295.9 tok/s
TG128 @ 1 ctx
14,350 tok/s
PP512
gemma-3-1b-it-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
268.0 tok/s
TG128 @ 1 ctx
33,841 tok/s
PP512
Qwen3-0.6B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
440.2 tok/s
TG128 @ 1 ctx
12,494 tok/s
PP512
Qwen3-0.6B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M3 UltraMetal
501.7 tok/s
TG128 @ 1 ctx
12,830 tok/s
PP512
Qwen3-0.6B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
284.7 tok/s
TG128 @ 1 ctx
4,787 tok/s
PP512
Qwen3-0.6B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
455.8 tok/s
TG128 @ 1 ctx
9,749 tok/s
PP512
Qwen3-0.6B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 ProMetal
208.6 tok/s
TG128 @ 1 ctx
3,571 tok/s
PP512
Qwen3-0.6B-Q8
BaseRTQ8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
74.3 tok/s
TG128 @ 1 ctx
2,225 tok/s
PP512
Qwen3-0.6B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
458.8 tok/s
TG128 @ 1 ctx
20,970 tok/s
PP512
Qwen3-0.6B-cuda-q4
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
456.1 tok/s
TG128 @ 1 ctx
52,750 tok/s
PP512
Qwen3-0.6B-cuda-q4mix
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
413.2 tok/s
TG128 @ 1 ctx
49,865 tok/s
PP512
Qwen3-0.6B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M4 MaxMetal
583.3 tok/s
TG128 @ 1 ctx
9,998 tok/s
PP512
Qwen3-0.6B-Q4
BaseRTQ4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M1 MaxMetal
100.8 tok/s
TG128 @ 1 ctx
1,851 tok/s
PP512
Qwen3-0.6B-cuda-q8
BaseRTQ4
Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled.
NVIDIA GB10CUDA
315.8 tok/s
TG128 @ 1 ctx
51,373 tok/s
PP512
basecompute/gpt-oss-120b
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GB10CUDA
40.3 tok/s
TG128
2,679 tok/s
PP512
Qwen/Qwen3.6-27B
BaseRTQ4· dense-q4mix-cuda
This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified.
NVIDIA GB10CUDA
12.4 tok/s
TG128
1,140 tok/s
PP512
basecompute/Qwen3.6-27B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxMetal
31.3 tok/s
TG128
603 tok/s
PP512
Models Qwen Qwen3 0.6B
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxBLAS + Metal
439.1 tok/s
TG128
24,365 tok/s
PP512
Qwen3 0.6B Instruct
llama.cppQ8_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxBLAS + Metal
379.0 tok/s
TG128
24,983 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ8· default-q8
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxMetal
549.5 tok/s
TG128
33,088 tok/s
PP512
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxMetal
177.6 tok/s
TG128
4,527 tok/s
PP512
basecompute/Qwen3-0.6B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxMetal
708.3 tok/s
TG128
34,136 tok/s
PP512
basecompute/Qwen3.8-27B
BaseRTQ4· default-q4
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
Apple M5 MaxMetal
34.0 tok/s
TG128
589 tok/s
PP512