Hardware2
Apple M5 Max
20 runsBest decode · TG128
698.5tok/s
Best prefill · PP512
32,977tok/s
Apple M1 Pro
2 runsBest decode · TG128
151.6tok/s
Best prefill · PP512
2,620tok/s
Badges6
Record1
Pioneer1
Volume1
Coverage3
12 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
20 runs
Regular contributor
Volume
Model explorer
13 models covered
Coverage
Hardware explorer
2 devices covered
Coverage
Multi-backend
MTL,BLAS + metal
Coverage
Submissions22
Sort by
Order
Prefill size
22 submissions · PP512 selected
| Model / runtime / format | Device | Backend | Report | |||
|---|---|---|---|---|---|---|
basecompute/gemma-3-1b-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 414.2 tok/s TG128 | 23,019 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 698.5 tok/s TG128 | 32,977 tok/s PP512 | ||
basecompute/gemma-3-1b-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 415.5 tok/s TG128 | 19,770 tok/s PP512 | ||
basecompute/gemma-4-E4B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 121.1 tok/s TG128 | 8,156 tok/s PP512 | ||
basecompute/gemma-4-E2B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 197.5 tok/s TG128 | 20,099 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 21.5 tok/s TG128 | 594 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 20.7 tok/s TG128 | 599 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 21.5 tok/s TG128 | 590 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 177.0 tok/s TG128 | 4,821 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 180.8 tok/s TG128 | 4,951 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 183.7 tok/s TG128 | 4,941 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 109.2 tok/s TG128 | 4,943 tok/s PP512 | ||
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 184.4 tok/s TG128 | 4,984 tok/s PP512 | ||
basecompute/gpt-oss-120b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 114.3 tok/s TG128 | 1,401 tok/s PP512 | ||
basecompute/gpt-oss-20b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 169.0 tok/s TG128 | 2,179 tok/s PP512 | ||
basecompute/Muse-Glimmer-30B BaseRTpassthrough_gguf· default-q4k-dynamic The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 15.7 tok/s TG128 | 830 tok/s PP512 | ||
Gemma-4 12B IT (smart Q4_0, QAT-lossless) llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | BLAS + Metal | 54.2 tok/s TG128 | 1,777 tok/s PP512 | ||
basecompute/gemma-4-26B-A4B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 92.0 tok/s TG128 | 4,055 tok/s PP512 | ||
basecompute/Qwen3.6-35B-A3B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 119.6 tok/s TG128 | 1,540 tok/s PP512 | ||
basecompute/Qwen3.8-27B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max | Metal | 19.7 tok/s TG128 | 589 tok/s PP512 | ||
Qwen 0.6b Coder llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Pro | BLAS + Metal | 151.6 tok/s TG128 | 2,620 tok/s PP512 | ||
basecompute/gemma-4-E4B-it BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Pro | Metal | 45.9 tok/s TG128 | 565 tok/s PP512 |