Hardware3
Apple M5 Pro
7 runsBest decode · TG128 @ 1 ctx
481.5tok/s
Best prefill · PP512
19,595tok/s
AMD Ryzen 5 7600 6-Core Processor
1 runBest decode · TG128
16.2tok/s
Best prefill · PP512
78tok/s
Apple M1 Pro
1 runBest decode · TG128 @ 1 ctx
268.8tok/s
Best prefill · PP512
2,847tok/s
Badges6
Record1
Pioneer1
Volume1
Coverage3
3 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
5 runs
Contributor
Volume
Model explorer
4 models covered
Coverage
Hardware explorer
3 devices covered
Coverage
Multi-backend
BLAS,MTL + CPU + metal
Coverage
Submissions9
Sort by
Order
Prefill size
9 submissions · PP512 selected
| Model / runtime / format | Device | Backend | Report | |||
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Pro | Metal | 268.8 tok/s TG128 @ 1 ctx | 2,847 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 59.7 tok/s TG128 @ 1 ctx | 1,762 tok/s PP512 | ||
basecompute/Qwen3-0.6B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 481.5 tok/s TG128 @ 1 ctx | 19,595 tok/s PP512 | ||
Qwen3.5-4B llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Ryzen 5 7600 6-Core Processor | CPU | 16.2 tok/s TG128 | 78 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 59.7 tok/s TG128 | 1,761 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 58.2 tok/s TG128 | 1,758 tok/s PP512 | ||
Qwen3-4B llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | BLAS + Metal | 93.6 tok/s TG128 | 2,705 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 60.6 tok/s TG128 | 1,763 tok/s PP512 | ||
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro | Metal | 56.9 tok/s TG128 | 1,707 tok/s PP512 |