Hardware6
NVIDIA GB10
71 runsBest decode · TG128 @ 1 ctx
458.8tok/s
Best prefill · PP512
52,750tok/s
Apple M3 Ultra
58 runsBest decode · TG128 @ 1 ctx
501.7tok/s
Best prefill · PP512
12,830tok/s
Apple M1 Max
56 runsBest decode · TG128 @ 1 ctx
242.1tok/s
Best prefill · PP512
4,521tok/s
Apple M4 Max
44 runsBest decode · TG128 @ 1 ctx
583.3tok/s
Best prefill · PP512
9,998tok/s
Apple M4 Pro
33 runsBest decode · TG128 @ 1 ctx
284.7tok/s
Best prefill · PP512
4,787tok/s
Apple M5 Max
7 runsBest decode · TG128
708.3tok/s
Best prefill · PP512
34,136tok/s
Badges6
Record1
Pioneer1
Volume1
Coverage3
201 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
40 runs
Sustained contributor
Volume
Model explorer
28 models covered
Coverage
Hardware explorer
6 devices covered
Coverage
Multi-backend
BLAS,MTL + cuda + metal
Coverage
Submissions269
Sort by
Order
Prefill size
269 submissions · PP512 selected
| Model / runtime / format | Device | Backend | Report | |||
|---|---|---|---|---|---|---|
gemma-4-26B-A4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 73.8 tok/s TG128 @ 1 ctx | 1,679 tok/s PP512 | ||
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 185.1 tok/s TG128 @ 1 ctx | 2,554 tok/s PP512 | ||
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 175.8 tok/s TG128 @ 1 ctx | 2,627 tok/s PP512 | ||
Qwen3.8-27B-Q4-mtp BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M3 Ultra | Metal | 39.5 tok/s TG128 @ 1 ctx | 288 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 171.6 tok/s TG128 @ 1 ctx | 2,295 tok/s PP512 | ||
muse-glimmer-30B-kquant-dynamic BaseRTpassthrough_gguf The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 27.3 tok/s TG128 @ 1 ctx | 399 tok/s PP512 | ||
muse-glimmer-30B-kquant-dynamic BaseRTpassthrough_gguf The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 22.8 tok/s TG128 @ 1 ctx | 254 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 139.7 tok/s TG128 @ 1 ctx | 886 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 140.1 tok/s TG128 @ 1 ctx | 895 tok/s PP512 | ||
muse-glimmer-30B-kquant-17gb BaseRTpassthrough_gguf The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 27.2 tok/s TG128 @ 1 ctx | 266 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 169.1 tok/s TG128 @ 1 ctx | 1,632 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 204.6 tok/s TG128 @ 1 ctx | 2,685 tok/s PP512 | ||
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 193.1 tok/s TG128 @ 1 ctx | 1,828 tok/s PP512 | ||
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 179.1 tok/s TG128 @ 1 ctx | 1,150 tok/s PP512 | ||
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 176.8 tok/s TG128 @ 1 ctx | 1,150 tok/s PP512 | ||
Qwen3.8-27B-Q4-mtp BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M4 Max | Metal | 31.5 tok/s TG128 @ 1 ctx | 221 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 179.3 tok/s TG128 @ 1 ctx | 2,692 tok/s PP512 | ||
Qwen3-1.7B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 125.3 tok/s TG128 @ 1 ctx | 1,822 tok/s PP512 | ||
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 90.3 tok/s TG128 @ 1 ctx | 734 tok/s PP512 | ||
Qwen3.5-2B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 113.5 tok/s TG128 @ 1 ctx | 1,138 tok/s PP512 | ||
Qwen3.5-2B-Base-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 113.9 tok/s TG128 @ 1 ctx | 1,149 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 95.0 tok/s TG128 @ 1 ctx | 943 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 77.9 tok/s TG128 @ 1 ctx | 727 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 78.4 tok/s TG128 @ 1 ctx | 732 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 139.8 tok/s TG128 @ 1 ctx | 1,789 tok/s PP512 | ||
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 126.0 tok/s TG128 @ 1 ctx | 3,722 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 71.2 tok/s TG128 @ 1 ctx | 943 tok/s PP512 | ||
gemma-3-1b-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 328.0 tok/s TG128 @ 1 ctx | 9,903 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 410.1 tok/s TG128 @ 1 ctx | 9,066 tok/s PP512 | ||
gemma-3-1b-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 257.8 tok/s TG128 @ 1 ctx | 9,953 tok/s PP512 | ||
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 306.7 tok/s TG128 @ 1 ctx | 1,759 tok/s PP512 | ||
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 307.4 tok/s TG128 @ 1 ctx | 5,420 tok/s PP512 | ||
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 293.8 tok/s TG128 @ 1 ctx | 1,756 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 345.2 tok/s TG128 @ 1 ctx | 9,055 tok/s PP512 | ||
Qwen3-1.7B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 238.6 tok/s TG128 @ 1 ctx | 5,698 tok/s PP512 | ||
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 165.7 tok/s TG128 @ 1 ctx | 2,461 tok/s PP512 | ||
Qwen3.5-2B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 222.9 tok/s TG128 @ 1 ctx | 1,756 tok/s PP512 | ||
Qwen3.5-2B-Base-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 222.4 tok/s TG128 @ 1 ctx | 1,696 tok/s PP512 | ||
Qwen3-4B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 56.3 tok/s TG128 @ 1 ctx | 726 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 195.4 tok/s TG128 @ 1 ctx | 3,241 tok/s PP512 | ||
gemma-3-1b-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 170.0 tok/s TG128 @ 1 ctx | 4,120 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 153.7 tok/s TG128 @ 1 ctx | 2,529 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 242.1 tok/s TG128 @ 1 ctx | 3,333 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 155.6 tok/s TG128 @ 1 ctx | 2,582 tok/s PP512 | ||
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 219.7 tok/s TG128 @ 1 ctx | 2,244 tok/s PP512 | ||
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 137.4 tok/s TG128 @ 1 ctx | 8,996 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 140.1 tok/s TG128 @ 1 ctx | 1,797 tok/s PP512 | ||
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 192.9 tok/s TG128 @ 1 ctx | 429 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 154.4 tok/s TG128 @ 1 ctx | 3,316 tok/s PP512 | ||
Qwen3-4B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 116.8 tok/s TG128 @ 1 ctx | 2,441 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 118.3 tok/s TG128 @ 1 ctx | 2,557 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 118.2 tok/s TG128 @ 1 ctx | 2,559 tok/s PP512 | ||
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 190.0 tok/s TG128 @ 1 ctx | 439 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 129.3 tok/s TG128 @ 1 ctx | 1,440 tok/s PP512 | ||
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 94.6 tok/s TG128 @ 1 ctx | 3,610 tok/s PP512 | ||
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 113.6 tok/s TG128 @ 1 ctx | 1,357 tok/s PP512 | ||
Qwen3-1.7B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 153.9 tok/s TG128 @ 1 ctx | 2,223 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 229.4 tok/s TG128 @ 1 ctx | 3,372 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 124.4 tok/s TG128 @ 1 ctx | 1,438 tok/s PP512 | ||
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 101.9 tok/s TG128 @ 1 ctx | 927 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 56.2 tok/s TG128 @ 1 ctx | 731 tok/s PP512 | ||
gemma-4-E2B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 116.6 tok/s TG128 @ 1 ctx | 8,610 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 78.9 tok/s TG128 @ 1 ctx | 1,388 tok/s PP512 | ||
Qwen3.5-2B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 141.5 tok/s TG128 @ 1 ctx | 400 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 75.3 tok/s TG128 @ 1 ctx | 1,392 tok/s PP512 | ||
gemma-4-E4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 75.4 tok/s TG128 @ 1 ctx | 3,499 tok/s PP512 | ||
Qwen3.5-2B-Base-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 140.7 tok/s TG128 @ 1 ctx | 438 tok/s PP512 | ||
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 72.5 tok/s TG128 @ 1 ctx | 1,338 tok/s PP512 | ||
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 163.9 tok/s TG128 @ 1 ctx | 2,590 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 110.7 tok/s TG128 @ 1 ctx | 1,229 tok/s PP512 | ||
Qwen3.6-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 36.8 tok/s TG128 @ 1 ctx | 281 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 91.3 tok/s TG128 @ 1 ctx | 922 tok/s PP512 | ||
gemma-4-26B-A4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 81.5 tok/s TG128 @ 1 ctx | 1,664 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 56.3 tok/s TG128 @ 1 ctx | 731 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 93.2 tok/s TG128 @ 1 ctx | 921 tok/s PP512 | ||
gemma-4-26B-A4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 81.5 tok/s TG128 @ 1 ctx | 2,422 tok/s PP512 | ||
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 120.8 tok/s TG128 @ 1 ctx | 4,521 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 40.6 tok/s TG128 @ 1 ctx | 509 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 135.3 tok/s TG128 @ 1 ctx | 2,416 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 137.1 tok/s TG128 @ 1 ctx | 2,287 tok/s PP512 | ||
gemma-4-E2B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 91.1 tok/s TG128 @ 1 ctx | 4,442 tok/s PP512 | ||
muse-glimmer-30B-kquant-17gb BaseRTpassthrough_gguf The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 31.1 tok/s TG128 @ 1 ctx | 405 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 63.7 tok/s TG128 @ 1 ctx | 509 tok/s PP512 | ||
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 69.1 tok/s TG128 @ 1 ctx | 1,094 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 71.5 tok/s TG128 @ 1 ctx | 925 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 133.0 tok/s TG128 @ 1 ctx | 947 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 59.0 tok/s TG128 @ 1 ctx | 394 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 133.0 tok/s TG128 @ 1 ctx | 922 tok/s PP512 | ||
Qwen3-4B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 71.3 tok/s TG128 @ 1 ctx | 920 tok/s PP512 | ||
gemma-4-26B-A4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 75.3 tok/s TG128 @ 1 ctx | 2,258 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 88.4 tok/s TG128 @ 1 ctx | 1,240 tok/s PP512 | ||
gemma-3-1b-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 178.8 tok/s TG128 @ 1 ctx | 3,337 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 72.1 tok/s TG128 @ 1 ctx | 925 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 66.4 tok/s TG128 @ 1 ctx | 511 tok/s PP512 | ||
Qwen3.6-27B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 22.1 tok/s TG128 @ 1 ctx | 280 tok/s PP512 | ||
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 75.9 tok/s TG128 @ 1 ctx | 1,394 tok/s PP512 | ||
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 52.6 tok/s TG128 @ 1 ctx | 394 tok/s PP512 | ||
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 60.2 tok/s TG128 @ 1 ctx | 507 tok/s PP512 | ||
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 22.1 tok/s TG128 @ 1 ctx | 280 tok/s PP512 | ||
Qwen3.6-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 30.3 tok/s TG128 @ 1 ctx | 221 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 38.6 tok/s TG128 @ 1 ctx | 509 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 95.9 tok/s TG128 @ 1 ctx | 2,289 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 95.8 tok/s TG128 @ 1 ctx | 2,391 tok/s PP512 | ||
gemma-4-E4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 52.0 tok/s TG128 @ 1 ctx | 1,368 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 102.1 tok/s TG128 @ 1 ctx | 2,311 tok/s PP512 | ||
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 38.1 tok/s TG128 @ 1 ctx | 500 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 101.9 tok/s TG128 @ 1 ctx | 2,315 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 56.1 tok/s TG128 @ 1 ctx | 394 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M3 Ultra | Metal | 109.6 tok/s TG128 @ 1 ctx | 927 tok/s PP512 | ||
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 82.3 tok/s TG128 @ 1 ctx | 881 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M3 Ultra | Metal | 109.1 tok/s TG128 @ 1 ctx | 939 tok/s PP512 | ||
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 76.7 tok/s TG128 @ 1 ctx | 880 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 110.8 tok/s TG128 @ 1 ctx | 910 tok/s PP512 | ||
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 76.6 tok/s TG128 @ 1 ctx | 890 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 109.5 tok/s TG128 @ 1 ctx | 913 tok/s PP512 | ||
Qwen3.6-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 17.9 tok/s TG128 @ 1 ctx | 80 tok/s PP512 | ||
gpt-oss-120b-MXFP4 BaseRTmxfp4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M3 Ultra | Metal | 114.0 tok/s TG128 @ 1 ctx | 1,700 tok/s PP512 | ||
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 146.7 tok/s TG128 @ 1 ctx | 1,715 tok/s PP512 | ||
gemma-4-E2B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 83.9 tok/s TG128 @ 1 ctx | 3,628 tok/s PP512 | ||
Qwen3.5-122B-A10B-Q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M3 Ultra | Metal | 66.4 tok/s TG128 @ 1 ctx | 536 tok/s PP512 | ||
Qwen3.5-122B-A10B-Q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M3 Ultra | Metal | 66.3 tok/s TG128 @ 1 ctx | 532 tok/s PP512 | ||
gemma-4-26B-A4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 51.5 tok/s TG128 @ 1 ctx | 871 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 33.8 tok/s TG128 @ 1 ctx | 393 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 87.2 tok/s TG128 @ 1 ctx | 901 tok/s PP512 | ||
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 146.9 tok/s TG128 @ 1 ctx | 1,700 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 87.7 tok/s TG128 @ 1 ctx | 901 tok/s PP512 | ||
Qwen3.8-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 38.0 tok/s TG128 @ 1 ctx | 282 tok/s PP512 | ||
Qwen3.8-27B-Q4-mtp BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M1 Max | Metal | 18.2 tok/s TG128 @ 1 ctx | 82 tok/s PP512 | ||
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 160.5 tok/s TG128 @ 1 ctx | 1,709 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 32.1 tok/s TG128 @ 1 ctx | 392 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 110.3 tok/s TG128 @ 1 ctx | 870 tok/s PP512 | ||
muse-glimmer-30B-kquant-17gb BaseRTpassthrough_gguf The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 17.3 tok/s TG128 @ 1 ctx | 140 tok/s PP512 | ||
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 59.6 tok/s TG128 @ 1 ctx | 942 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 93.7 tok/s TG128 @ 1 ctx | 280 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 94.2 tok/s TG128 @ 1 ctx | 278 tok/s PP512 | ||
gemma-4-E4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 44.5 tok/s TG128 @ 1 ctx | 1,081 tok/s PP512 | ||
gemma-4-E4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 77.4 tok/s TG128 @ 1 ctx | 2,481 tok/s PP512 | ||
muse-glimmer-30B-kquant-dynamic BaseRTpassthrough_gguf The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 14.3 tok/s TG128 @ 1 ctx | 136 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 60.8 tok/s TG128 @ 1 ctx | 947 tok/s PP512 | ||
gemma-4-26B-A4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 45.5 tok/s TG128 @ 1 ctx | 877 tok/s PP512 | ||
Qwen3.6-27B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 11.0 tok/s TG128 @ 1 ctx | 83 tok/s PP512 | ||
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 31.6 tok/s TG128 @ 1 ctx | 392 tok/s PP512 | ||
Qwen3.5-122B-A10B-Q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 30.6 tok/s TG128 @ 1 ctx | 799 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 64.1 tok/s TG128 @ 1 ctx | 946 tok/s PP512 | ||
gpt-oss-120b-MXFP4 BaseRTmxfp4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 41.4 tok/s TG128 @ 1 ctx | 2,622 tok/s PP512 | ||
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 11.0 tok/s TG128 @ 1 ctx | 82 tok/s PP512 | ||
gemma-4-E2B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 132.2 tok/s TG128 @ 1 ctx | 7,448 tok/s PP512 | ||
Qwen3.5-122B-A10B-Q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 31.2 tok/s TG128 @ 1 ctx | 795 tok/s PP512 | ||
gpt-oss-120b-MXFP4 BaseRTmxfp4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 41.7 tok/s TG128 @ 1 ctx | 2,626 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 105.5 tok/s TG128 @ 1 ctx | 946 tok/s PP512 | ||
gpt-oss-120b-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 35.8 tok/s TG128 @ 1 ctx | 2,638 tok/s PP512 | ||
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 79.3 tok/s TG128 @ 1 ctx | 732 tok/s PP512 | ||
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 99.2 tok/s TG128 @ 1 ctx | 943 tok/s PP512 | ||
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 11.0 tok/s TG128 @ 1 ctx | 82 tok/s PP512 | ||
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 109.3 tok/s TG128 @ 1 ctx | 2,552 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 110.3 tok/s TG128 @ 1 ctx | 947 tok/s PP512 | ||
gpt-oss-120b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 39.7 tok/s TG128 @ 1 ctx | 2,619 tok/s PP512 | ||
gpt-oss-120b-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 35.5 tok/s TG128 @ 1 ctx | 2,278 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 104.3 tok/s TG128 @ 1 ctx | 1,714 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 64.6 tok/s TG128 @ 1 ctx | 893 tok/s PP512 | ||
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 73.7 tok/s TG128 @ 1 ctx | 731 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 63.7 tok/s TG128 @ 1 ctx | 895 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 104.3 tok/s TG128 @ 1 ctx | 1,723 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 61.1 tok/s TG128 @ 1 ctx | 857 tok/s PP512 | ||
Qwen3.6-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 55.4 tok/s TG128 @ 1 ctx | 2,181 tok/s PP512 | ||
Qwen3.5-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 54.7 tok/s TG128 @ 1 ctx | 2,447 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 60.6 tok/s TG128 @ 1 ctx | 858 tok/s PP512 | ||
Qwen3-4B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 105.4 tok/s TG128 @ 1 ctx | 1,712 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M1 Max | Metal | 72.0 tok/s TG128 @ 1 ctx | 250 tok/s PP512 | ||
Qwen3.6-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 54.4 tok/s TG128 @ 1 ctx | 2,188 tok/s PP512 | ||
Qwen3.5-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 54.2 tok/s TG128 @ 1 ctx | 2,428 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 133.1 tok/s TG128 @ 1 ctx | 2,218 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 61.2 tok/s TG128 @ 1 ctx | 7,253 tok/s PP512 | ||
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 175.3 tok/s TG128 @ 1 ctx | 7,670 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M1 Max | Metal | 72.2 tok/s TG128 @ 1 ctx | 251 tok/s PP512 | ||
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 7.5 tok/s TG128 @ 1 ctx | 1,127 tok/s PP512 | ||
Qwen3.6-27B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 8.4 tok/s TG128 @ 1 ctx | 1,146 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 49.4 tok/s TG128 @ 1 ctx | 7,315 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 122.7 tok/s TG128 @ 1 ctx | 1,721 tok/s PP512 | ||
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 7.6 tok/s TG128 @ 1 ctx | 1,152 tok/s PP512 | ||
Qwen3.6-27B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 7.5 tok/s TG128 @ 1 ctx | 1,141 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 121.9 tok/s TG128 @ 1 ctx | 1,711 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 70.8 tok/s TG128 @ 1 ctx | 247 tok/s PP512 | ||
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 72.7 tok/s TG128 @ 1 ctx | 734 tok/s PP512 | ||
gemma-4-26B-A4B-it-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 40.4 tok/s TG128 @ 1 ctx | 6,045 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 178.8 tok/s TG128 @ 1 ctx | 2,202 tok/s PP512 | ||
Qwen3.6-35B-A3B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 72.8 tok/s TG128 @ 1 ctx | 2,279 tok/s PP512 | ||
Qwen3.5-2B-Base-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 208.0 tok/s TG128 @ 1 ctx | 1,727 tok/s PP512 | ||
Qwen3.5-35B-A3B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 92.4 tok/s TG128 @ 1 ctx | 2,521 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 72.2 tok/s TG128 @ 1 ctx | 232 tok/s PP512 | ||
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 84.3 tok/s TG128 @ 1 ctx | 2,028 tok/s PP512 | ||
Qwen3.6-27B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 12.3 tok/s TG128 @ 1 ctx | 1,127 tok/s PP512 | ||
Qwen3.5-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 84.4 tok/s TG128 @ 1 ctx | 2,093 tok/s PP512 | ||
Qwen3.5-2B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 205.4 tok/s TG128 @ 1 ctx | 1,721 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 99.0 tok/s TG128 @ 1 ctx | 7,439 tok/s PP512 | ||
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 99.8 tok/s TG128 @ 1 ctx | 3,862 tok/s PP512 | ||
Qwen3.8-27B-Q4-mtp BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 13.1 tok/s TG128 @ 1 ctx | 1,136 tok/s PP512 | ||
Qwen3.8-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 13.1 tok/s TG128 @ 1 ctx | 1,131 tok/s PP512 | ||
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 157.5 tok/s TG128 @ 1 ctx | 1,725 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 80.7 tok/s TG128 @ 1 ctx | 4,715 tok/s PP512 | ||
Qwen3.6-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 13.0 tok/s TG128 @ 1 ctx | 344 tok/s PP512 | ||
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 81.1 tok/s TG128 @ 1 ctx | 4,738 tok/s PP512 | ||
Qwen3.8-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 18.2 tok/s TG128 @ 1 ctx | 82 tok/s PP512 | ||
Qwen3-1.7B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 224.1 tok/s TG128 @ 1 ctx | 4,025 tok/s PP512 | ||
gemma-4-26B-A4B-it-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 45.3 tok/s TG128 @ 1 ctx | 6,190 tok/s PP512 | ||
gemma-4-26B-A4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 44.0 tok/s TG128 @ 1 ctx | 6,179 tok/s PP512 | ||
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 60.2 tok/s TG128 @ 1 ctx | 4,424 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 320.5 tok/s TG128 @ 1 ctx | 6,058 tok/s PP512 | ||
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 51.4 tok/s TG128 @ 1 ctx | 4,421 tok/s PP512 | ||
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 57.3 tok/s TG128 @ 1 ctx | 4,292 tok/s PP512 | ||
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 295.1 tok/s TG128 @ 1 ctx | 1,713 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 25.4 tok/s TG128 @ 1 ctx | 7,299 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 27.0 tok/s TG128 @ 1 ctx | 7,510 tok/s PP512 | ||
gemma-4-E2B-it-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 87.7 tok/s TG128 @ 1 ctx | 19,246 tok/s PP512 | ||
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 301.7 tok/s TG128 @ 1 ctx | 1,718 tok/s PP512 | ||
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 53.1 tok/s TG128 @ 1 ctx | 7,567 tok/s PP512 | ||
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 50.3 tok/s TG128 @ 1 ctx | 1,452 tok/s PP512 | ||
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 71.3 tok/s TG128 @ 1 ctx | 3,674 tok/s PP512 | ||
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 53.5 tok/s TG128 @ 1 ctx | 7,554 tok/s PP512 | ||
gemma-4-E2B-it-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 79.0 tok/s TG128 @ 1 ctx | 18,372 tok/s PP512 | ||
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 318.4 tok/s TG128 @ 1 ctx | 4,146 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 54.8 tok/s TG128 @ 1 ctx | 10,970 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 54.7 tok/s TG128 @ 1 ctx | 11,473 tok/s PP512 | ||
Llama-3.2-3B-Instruct-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 69.4 tok/s TG128 @ 1 ctx | 14,988 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 68.7 tok/s TG128 @ 1 ctx | 14,871 tok/s PP512 | ||
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 139.8 tok/s TG128 @ 1 ctx | 12,036 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 371.7 tok/s TG128 @ 1 ctx | 5,963 tok/s PP512 | ||
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 73.2 tok/s TG128 @ 1 ctx | 11,020 tok/s PP512 | ||
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 73.7 tok/s TG128 @ 1 ctx | 11,246 tok/s PP512 | ||
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 89.7 tok/s TG128 @ 1 ctx | 14,073 tok/s PP512 | ||
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 91.3 tok/s TG128 @ 1 ctx | 2,822 tok/s PP512 | ||
Qwen3.5-2B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 113.5 tok/s TG128 @ 1 ctx | 6,754 tok/s PP512 | ||
gemma-3-1b-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 281.0 tok/s TG128 @ 1 ctx | 7,527 tok/s PP512 | ||
Llama-3.2-3B-Instruct-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 110.5 tok/s TG128 @ 1 ctx | 14,651 tok/s PP512 | ||
Qwen3.5-2B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 151.5 tok/s TG128 @ 1 ctx | 6,837 tok/s PP512 | ||
Llama-3.2-1B-Instruct-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 173.7 tok/s TG128 @ 1 ctx | 34,296 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 172.8 tok/s TG128 @ 1 ctx | 35,941 tok/s PP512 | ||
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 181.9 tok/s TG128 @ 1 ctx | 4,191 tok/s PP512 | ||
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 181.6 tok/s TG128 @ 1 ctx | 4,185 tok/s PP512 | ||
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 202.5 tok/s TG128 @ 1 ctx | 7,394 tok/s PP512 | ||
gemma-3-1b-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 361.2 tok/s TG128 @ 1 ctx | 7,448 tok/s PP512 | ||
Llama-3.2-1B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 195.3 tok/s TG128 @ 1 ctx | 33,397 tok/s PP512 | ||
gemma-3-1b-it-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 193.2 tok/s TG128 @ 1 ctx | 33,641 tok/s PP512 | ||
Llama-3.2-1B-Instruct-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 261.3 tok/s TG128 @ 1 ctx | 36,647 tok/s PP512 | ||
gemma-3-1b-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 224.9 tok/s TG128 @ 1 ctx | 3,430 tok/s PP512 | ||
gemma-3-1b-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 220.0 tok/s TG128 @ 1 ctx | 4,102 tok/s PP512 | ||
gemma-3-1b-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 295.9 tok/s TG128 @ 1 ctx | 14,350 tok/s PP512 | ||
gemma-3-1b-it-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 268.0 tok/s TG128 @ 1 ctx | 33,841 tok/s PP512 | ||
Qwen3-0.6B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 440.2 tok/s TG128 @ 1 ctx | 12,494 tok/s PP512 | ||
Qwen3-0.6B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra | Metal | 501.7 tok/s TG128 @ 1 ctx | 12,830 tok/s PP512 | ||
Qwen3-0.6B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 284.7 tok/s TG128 @ 1 ctx | 4,787 tok/s PP512 | ||
Qwen3-0.6B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 455.8 tok/s TG128 @ 1 ctx | 9,749 tok/s PP512 | ||
Qwen3-0.6B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro | Metal | 208.6 tok/s TG128 @ 1 ctx | 3,571 tok/s PP512 | ||
Qwen3-0.6B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 74.3 tok/s TG128 @ 1 ctx | 2,225 tok/s PP512 | ||
Qwen3-0.6B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 458.8 tok/s TG128 @ 1 ctx | 20,970 tok/s PP512 | ||
Qwen3-0.6B-cuda-q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 456.1 tok/s TG128 @ 1 ctx | 52,750 tok/s PP512 | ||
Qwen3-0.6B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 413.2 tok/s TG128 @ 1 ctx | 49,865 tok/s PP512 | ||
Qwen3-0.6B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max | Metal | 583.3 tok/s TG128 @ 1 ctx | 9,998 tok/s PP512 | ||
Qwen3-0.6B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max | Metal | 100.8 tok/s TG128 @ 1 ctx | 1,851 tok/s PP512 | ||
Qwen3-0.6B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 | CUDA | 315.8 tok/s TG128 @ 1 ctx | 51,373 tok/s PP512 | ||
basecompute/gpt-oss-120b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 | CUDA | 40.3 tok/s TG128 | 2,679 tok/s PP512 | ||
Qwen/Qwen3.6-27B BaseRTQ4· dense-q4mix-cuda This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified. | NVIDIA GB10 | CUDA | 12.4 tok/s TG128 | 1,140 tok/s PP512 | ||
basecompute/Qwen3.6-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 31.3 tok/s TG128 | 603 tok/s PP512 | ||
Models Qwen Qwen3 0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | BLAS + Metal | 439.1 tok/s TG128 | 24,365 tok/s PP512 | ||
Qwen3 0.6B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | BLAS + Metal | 379.0 tok/s TG128 | 24,983 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 549.5 tok/s TG128 | 33,088 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 177.6 tok/s TG128 | 4,527 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 708.3 tok/s TG128 | 34,136 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 34.0 tok/s TG128 | 589 tok/s PP512 |