Hardware4
Apple M5 Pro
122 runsBest decode · TG128
2,068tok/s
Best prefill · PP512
190,107tok/s
AMD Radeon RX 7900 XT
41 runsBest decode · TG128
368.6tok/s
Best prefill · PP512
24,691tok/s
AMD Radeon RX 7900 XT (RADV NAVI31)
41 runsBest decode · TG128
508.8tok/s
Best prefill · PP512
26,346tok/s
Tesla T10/Tesla T10/Tesla T10/Tesla T10
7 runsBest decode · TG128
157.9tok/s
Best prefill · PP512
3,390tok/s
Badges6
Record1
Pioneer1
Volume1
Coverage3
150 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
40 runs
Sustained contributor
Volume
Model explorer
27 models covered
Coverage
Hardware explorer
4 devices covered
Coverage
Multi-backend
BLAS,MTL + CUDA + ROCm + Vulkan + metal
Coverage
Submissions211
Sort by
Order
Prefill size
211 submissions · PP512 selected
| Model / runtime / format | Device | Backend | Report | |||
|---|---|---|---|---|---|---|
Ornith-1.5-9B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 104.8 tok/s TG128 | 2,935 tok/s PP512 | ||
Gpt-Oss-20B llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 157.9 tok/s TG128 | 3,038 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 48.0 tok/s TG128 | 1,174 tok/s PP512 | ||
Meta Llama 3.1 8B Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 136.1 tok/s TG128 | 3,390 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 124.7 tok/s TG128 | 3,087 tok/s PP512 | ||
Qwen3.8-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 46.0 tok/s TG128 | 1,132 tok/s PP512 | ||
Qwen3.8-27B llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | CUDA | 48.1 tok/s TG128 | 910 tok/s PP512 | ||
ivnle/tinystories-lay8-hs512-hd8-33M BaseRTQ4· default-q4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | Apple M5 Pro | Metal | 2,068.0 tok/s TG128 | 190,107 tok/s PP512 | ||
Tinystories Lay8 Hs512 Hd8 33M llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 1,443.3 tok/s TG128 | 124,181 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 181.3 tok/s TG128 | 2,827 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 186.7 tok/s TG128 | 2,868 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 32.5 tok/s TG128 | 783 tok/s PP512 | ||
Qwen3.6-35B-A3B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 124.1 tok/s TG128 | 2,597 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 33.7 tok/s TG128 | 782 tok/s PP512 | ||
Qwen3-8B llama.cppBF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 45.6 tok/s TG128 | 1,168 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 36.8 tok/s TG128 | 739 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 184.5 tok/s TG128 | 2,474 tok/s PP512 | ||
Gpt-Oss-20B llama.cppF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 146.8 tok/s TG128 | 3,262 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 198.1 tok/s TG128 | 2,837 tok/s PP512 | ||
Gpt-Oss-20B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 195.6 tok/s TG128 | 3,292 tok/s PP512 | ||
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 67.4 tok/s TG128 | 2,253 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 68.6 tok/s TG128 | 2,286 tok/s PP512 | ||
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 77.6 tok/s TG128 | 2,669 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 95.6 tok/s TG128 | 4,214 tok/s PP512 | ||
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 88.6 tok/s TG128 | 2,497 tok/s PP512 | ||
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 95.6 tok/s TG128 | 2,430 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 97.7 tok/s TG128 | 2,512 tok/s PP512 | ||
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 104.9 tok/s TG128 | 2,640 tok/s PP512 | ||
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 105.7 tok/s TG128 | 2,629 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_1 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 115.6 tok/s TG128 | 2,902 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 115.3 tok/s TG128 | 2,643 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 122.8 tok/s TG128 | 4,161 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 118.5 tok/s TG128 | 2,625 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 119.8 tok/s TG128 | 2,756 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 126.6 tok/s TG128 | 4,142 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 122.7 tok/s TG128 | 2,743 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 126.7 tok/s TG128 | 2,959 tok/s PP512 | ||
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 127.3 tok/s TG128 | 2,498 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 134.3 tok/s TG128 | 4,824 tok/s PP512 | ||
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 130.2 tok/s TG128 | 2,459 tok/s PP512 | ||
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 146.2 tok/s TG128 | 2,563 tok/s PP512 | ||
Llama 3.2 3B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 167.5 tok/s TG128 | 5,846 tok/s PP512 | ||
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 150.7 tok/s TG128 | 2,528 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 180.3 tok/s TG128 | 4,793 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 185.1 tok/s TG128 | 4,783 tok/s PP512 | ||
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 223.4 tok/s TG128 | 5,721 tok/s PP512 | ||
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 229.7 tok/s TG128 | 5,737 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 479.4 tok/s TG128 | 26,346 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 491.9 tok/s TG128 | 25,964 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) | Vulkan | 508.8 tok/s TG128 | 26,137 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 128.6 tok/s TG128 | 2,661 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 128.5 tok/s TG128 | 2,820 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 29.3 tok/s TG128 | 840 tok/s PP512 | ||
Qwen3.6-35B-A3B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 100.3 tok/s TG128 | 2,501 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 29.8 tok/s TG128 | 841 tok/s PP512 | ||
Qwen3-8B llama.cppBF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 43.5 tok/s TG128 | 3,133 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 29.7 tok/s TG128 | 807 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 128.8 tok/s TG128 | 2,596 tok/s PP512 | ||
Gpt-Oss-20B llama.cppF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 128.8 tok/s TG128 | 3,604 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 137.4 tok/s TG128 | 2,386 tok/s PP512 | ||
Gpt-Oss-20B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 160.1 tok/s TG128 | 3,502 tok/s PP512 | ||
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 64.0 tok/s TG128 | 3,207 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 66.2 tok/s TG128 | 3,292 tok/s PP512 | ||
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 74.2 tok/s TG128 | 3,241 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 84.3 tok/s TG128 | 4,564 tok/s PP512 | ||
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 80.1 tok/s TG128 | 2,903 tok/s PP512 | ||
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 84.0 tok/s TG128 | 2,823 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 87.4 tok/s TG128 | 2,851 tok/s PP512 | ||
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 86.8 tok/s TG128 | 2,869 tok/s PP512 | ||
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 87.5 tok/s TG128 | 2,837 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_1 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 107.3 tok/s TG128 | 2,894 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 93.2 tok/s TG128 | 3,076 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 109.3 tok/s TG128 | 4,369 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 93.7 tok/s TG128 | 3,110 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 97.5 tok/s TG128 | 3,150 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 111.0 tok/s TG128 | 4,443 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 97.1 tok/s TG128 | 3,172 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 117.4 tok/s TG128 | 3,233 tok/s PP512 | ||
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 91.6 tok/s TG128 | 2,911 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 110.8 tok/s TG128 | 5,581 tok/s PP512 | ||
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 91.6 tok/s TG128 | 2,864 tok/s PP512 | ||
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 105.5 tok/s TG128 | 2,834 tok/s PP512 | ||
Llama 3.2 3B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 142.3 tok/s TG128 | 7,210 tok/s PP512 | ||
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 112.3 tok/s TG128 | 2,798 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 147.1 tok/s TG128 | 5,337 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 147.8 tok/s TG128 | 5,364 tok/s PP512 | ||
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 171.9 tok/s TG128 | 6,832 tok/s PP512 | ||
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 172.0 tok/s TG128 | 6,980 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 313.7 tok/s TG128 | 24,691 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 367.7 tok/s TG128 | 23,718 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT | ROCm | 368.6 tok/s TG128 | 23,972 tok/s PP512 | ||
basecompute/Qwen3-30B-A3B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 104.0 tok/s TG128 | 3,680 tok/s PP512 | ||
basecompute/gpt-oss-20b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 110.1 tok/s TG128 | 1,156 tok/s PP512 | ||
basecompute/Llama-3.1-8B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 35.6 tok/s TG128 | 1,774 tok/s PP512 | ||
basecompute/Llama-3.1-8B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 63.2 tok/s TG128 | 1,795 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 34.5 tok/s TG128 | 1,744 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 60.7 tok/s TG128 | 1,755 tok/s PP512 | ||
basecompute/gemma-4-E4B-it BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 51.0 tok/s TG128 | 4,462 tok/s PP512 | ||
basecompute/gemma-4-E4B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 81.9 tok/s TG128 | 4,636 tok/s PP512 | ||
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 62.8 tok/s TG128 | 3,405 tok/s PP512 | ||
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 90.4 tok/s TG128 | 3,331 tok/s PP512 | ||
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 79.5 tok/s TG128 | 4,343 tok/s PP512 | ||
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 106.8 tok/s TG128 | 4,297 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 350.2 tok/s TG128 | 20,537 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 516.8 tok/s TG128 | 20,778 tok/s PP512 | ||
Qwen3-8B-base-q6 BaseRTQ6 This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified. | Apple M5 Pro | Metal | 42.1 tok/s TG128 | 1,723 tok/s PP512 | ||
Qwen3-8B-base-q5 BaseRTQ5 This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified. | Apple M5 Pro | Metal | 48.4 tok/s TG128 | 1,730 tok/s PP512 | ||
Qwen3-8B-base-q3 BaseRTQ3 This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified. | Apple M5 Pro | Metal | 62.0 tok/s TG128 | 1,748 tok/s PP512 | ||
Qwen3-8B-base-q2 BaseRTQ2 This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified. | Apple M5 Pro | Metal | 81.8 tok/s TG128 | 1,786 tok/s PP512 | ||
Qwen3.6-35B-A3B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 61.5 tok/s TG128 | 1,533 tok/s PP512 | ||
Qwen3.6-35B-A3B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 61.5 tok/s TG128 | 1,673 tok/s PP512 | ||
Qwen3.6-35B-A3B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 67.5 tok/s TG128 | 1,737 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 11.9 tok/s TG128 | 377 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 14.5 tok/s TG128 | 374 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 15.2 tok/s TG128 | 376 tok/s PP512 | ||
Qwen3.6-27B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 16.7 tok/s TG128 | 373 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 78.9 tok/s TG128 | 1,646 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 84.7 tok/s TG128 | 1,644 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 95.2 tok/s TG128 | 1,845 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 95.0 tok/s TG128 | 1,831 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 97.3 tok/s TG128 | 1,708 tok/s PP512 | ||
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 102.7 tok/s TG128 | 1,893 tok/s PP512 | ||
Gpt-Oss-20B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 96.3 tok/s TG128 | 1,849 tok/s PP512 | ||
Gpt-Oss-20B llama.cppF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 66.6 tok/s TG128 | 1,852 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 29.6 tok/s TG128 | 1,465 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 43.4 tok/s TG128 | 1,421 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 55.5 tok/s TG128 | 1,423 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 55.6 tok/s TG128 | 1,424 tok/s PP512 | ||
Llama-3.1-8B-Instruct llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 58.7 tok/s TG128 | 1,538 tok/s PP512 | ||
Qwen3-8B llama.cppBF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 18.9 tok/s TG128 | 1,376 tok/s PP512 | ||
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 29.2 tok/s TG128 | 1,448 tok/s PP512 | ||
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 34.4 tok/s TG128 | 1,473 tok/s PP512 | ||
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 39.4 tok/s TG128 | 1,419 tok/s PP512 | ||
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 42.6 tok/s TG128 | 1,409 tok/s PP512 | ||
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 47.6 tok/s TG128 | 1,353 tok/s PP512 | ||
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 47.9 tok/s TG128 | 1,345 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 53.6 tok/s TG128 | 1,396 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 55.0 tok/s TG128 | 1,401 tok/s PP512 | ||
Qwen3-8B llama.cppQ4_1 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 51.9 tok/s TG128 | 1,535 tok/s PP512 | ||
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 60.9 tok/s TG128 | 1,414 tok/s PP512 | ||
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 61.5 tok/s TG128 | 1,418 tok/s PP512 | ||
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 69.6 tok/s TG128 | 1,465 tok/s PP512 | ||
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 72.8 tok/s TG128 | 1,477 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 47.9 tok/s TG128 | 2,346 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 69.4 tok/s TG128 | 2,235 tok/s PP512 | ||
Gemma-4-E4B-It llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 72.1 tok/s TG128 | 2,263 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 60.8 tok/s TG128 | 2,686 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 92.6 tok/s TG128 | 2,566 tok/s PP512 | ||
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 94.4 tok/s TG128 | 2,568 tok/s PP512 | ||
Llama 3.2 3B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 76.6 tok/s TG128 | 3,516 tok/s PP512 | ||
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 115.2 tok/s TG128 | 3,315 tok/s PP512 | ||
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 117.7 tok/s TG128 | 3,318 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 283.0 tok/s TG128 | 14,942 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 358.1 tok/s TG128 | 14,344 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 358.3 tok/s TG128 | 14,509 tok/s PP512 | ||
Qwen3-0.6B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 356.2 tok/s TG128 | 14,552 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 66.0 tok/s TG128 | 1,748 tok/s PP512 | ||
basecompute/Qwen3-30B-A3B-Thinking-2507 BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 65.1 tok/s TG128 | 3,692 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 10.6 tok/s TG128 | 355 tok/s PP512 | ||
basecompute/Qwen3.6-27B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 10.6 tok/s TG128 | 355 tok/s PP512 | ||
basecompute/gemma-4-26B-A4B-it BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 51.4 tok/s TG128 | 3,253 tok/s PP512 | ||
basecompute/Qwen3.5-2B-Base BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 131.4 tok/s TG128 | 2,642 tok/s PP512 | ||
basecompute/Qwen3.5-2B-Base BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 218.2 tok/s TG128 | 2,645 tok/s PP512 | ||
basecompute/Qwen3-4B-Thinking-2507 BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 64.7 tok/s TG128 | 3,414 tok/s PP512 | ||
basecompute/Qwen3-4B-Thinking-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 91.6 tok/s TG128 | 3,363 tok/s PP512 | ||
basecompute/Qwen3-30B-A3B-Thinking-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 104.4 tok/s TG128 | 3,694 tok/s PP512 | ||
basecompute/gpt-oss-20b BaseRTmxfp4· default-bf16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 103.6 tok/s TG128 | 1,013 tok/s PP512 | ||
basecompute/gpt-oss-20b BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 100.6 tok/s TG128 | 1,065 tok/s PP512 | ||
basecompute/Qwen3.5-35B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 113.4 tok/s TG128 | 1,304 tok/s PP512 | ||
basecompute/Qwen3.6-35B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 113.2 tok/s TG128 | 1,275 tok/s PP512 | ||
basecompute/Muse-Glimmer-30B BaseRTpassthrough_gguf· default-q4k-dynamic The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 14.4 tok/s TG128 | 495 tok/s PP512 | ||
basecompute/Muse-Glimmer-30B BaseRTpassthrough_gguf· q4k-17gb The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 17.0 tok/s TG128 | 466 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 100.1 tok/s TG128 | 2,499 tok/s PP512 | ||
basecompute/gemma-4-26B-A4B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 51.7 tok/s TG128 | 2,047 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 16.5 tok/s TG128 | 327 tok/s PP512 | ||
basecompute/gemma-3-1b-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 270.4 tok/s TG128 | 13,717 tok/s PP512 | ||
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 94.6 tok/s TG128 | 3,845 tok/s PP512 | ||
basecompute/Llama-3.2-1B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 206.3 tok/s TG128 | 10,528 tok/s PP512 | ||
basecompute/Qwen3-1.7B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 125.1 tok/s TG128 | 7,267 tok/s PP512 | ||
basecompute/Qwen3-1.7B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 210.3 tok/s TG128 | 7,268 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 311.0 tok/s TG128 | 18,474 tok/s PP512 | ||
basecompute/Qwen3-30B-A3B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 92.2 tok/s TG128 | 3,305 tok/s PP512 | ||
basecompute/Qwen3.6-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 16.1 tok/s TG128 | 364 tok/s PP512 | ||
basecompute/gpt-oss-20b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 110.9 tok/s TG128 | 1,155 tok/s PP512 | ||
basecompute/Llama-3.1-8B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 35.5 tok/s TG128 | 1,773 tok/s PP512 | ||
basecompute/Llama-3.1-8B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 62.0 tok/s TG128 | 1,795 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 35.2 tok/s TG128 | 1,748 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 60.5 tok/s TG128 | 1,763 tok/s PP512 | ||
basecompute/Mistral-7B-Instruct-v0.3 BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 37.4 tok/s TG128 | 1,783 tok/s PP512 | ||
basecompute/Mistral-7B-Instruct-v0.3 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 66.0 tok/s TG128 | 1,795 tok/s PP512 | ||
basecompute/gemma-4-E4B-it BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 51.3 tok/s TG128 | 4,490 tok/s PP512 | ||
basecompute/gemma-4-E4B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 82.6 tok/s TG128 | 4,656 tok/s PP512 | ||
basecompute/gemma-4-E2B-it BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 96.0 tok/s TG128 | 12,646 tok/s PP512 | ||
basecompute/gemma-4-E2B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 144.6 tok/s TG128 | 13,240 tok/s PP512 | ||
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 64.3 tok/s TG128 | 3,398 tok/s PP512 | ||
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 89.4 tok/s TG128 | 3,359 tok/s PP512 | ||
basecompute/Qwen3-4B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 64.7 tok/s TG128 | 3,412 tok/s PP512 | ||
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 79.5 tok/s TG128 | 4,344 tok/s PP512 | ||
basecompute/gemma-3-1b-it BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 210.3 tok/s TG128 | 15,061 tok/s PP512 | ||
basecompute/Llama-3.2-1B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 205.3 tok/s TG128 | 12,067 tok/s PP512 | ||
basecompute/Qwen3.5-2B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 130.7 tok/s TG128 | 2,638 tok/s PP512 | ||
basecompute/Qwen3.5-2B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 217.7 tok/s TG128 | 2,641 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 512.0 tok/s TG128 | 20,603 tok/s PP512 | ||
basecompute/gemma-3-1b-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 293.1 tok/s TG128 | 15,066 tok/s PP512 | ||
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 106.1 tok/s TG128 | 4,291 tok/s PP512 | ||
basecompute/Llama-3.2-1B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 229.1 tok/s TG128 | 11,673 tok/s PP512 | ||
basecompute/Qwen3-1.7B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 139.7 tok/s TG128 | 8,043 tok/s PP512 | ||
basecompute/Qwen3-1.7B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 235.4 tok/s TG128 | 8,040 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 349.8 tok/s TG128 | 20,366 tok/s PP512 | ||
basecompute/Qwen3-4B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 103.9 tok/s TG128 | 3,372 tok/s PP512 | ||
basecompute/Qwen3-4B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 107.3 tok/s TG128 | 3,384 tok/s PP512 |