Community submitted benchmarks for Local AI.
Every result comes from a signed report produced on community hardware. Browse by model, quantisation and chip.
Install & start · macOS / Linux
curl -LsSf https://computearena.ai/install.sh | sh -s launchRequires curl, tar and Python 3. Review and confirm the installation; the interactive menu opens in this terminal. No sudo required.
Latest runs
- Nemotron 3 Nano 30B A3BReported as nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 · nemotron_h_moebasecomputeBaseRTQ8Apple M4 Max94.9decode1,638prefill
- BaseRTQ8Apple M4 Max17.5decode222prefill
- BaseRTQ4Apple M4 Max17.7decode221prefill
- BaseRTQ8Apple M4 Max104.6decode878prefill
- BaseRTQ4Apple M4 Max31.5decode221prefill
- BaseRTQ8Apple M4 Max95.9decode1,795prefill
How it works
01
Install ComputeArena
Get the ComputeArena CLI for macOS or Linux, then choose BaseRT or llama.cpp as your runtime.
02
Run offline
No account is required. Signed reports remain on your machine until you submit them.
03
Submit when ready
Sign in, review the data in your saved reports, and submit the benchmarks you want to share publicly.
Leaderboard
Model × quantisation × chip
Headline rankings compare compatible PP512 prefill and TG128 decode workloads. The measured workload is shown beside every result.
Model
Runtime
Quantisation
Chip
Prefill size
Rank by
380 matching configurations
| # | Model / quant | Runtime | Chip / backend | User | Updated | Share | ||
|---|---|---|---|---|---|---|---|---|
| 1 | Tinystories Lay8 HS512 HD8 33MReported as ivnle/tinystories-lay8-hs512-hd8-33M · llama · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | ||
| 2 | Tinystories Lay8 HS512 HD8 33MReported as RichardErkhov/ivnle_-_tinystories-lay8-hs512-hd8-33M-gguf · llama Q4_0 | llama.cpp | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | ||
| 3 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | ||
| 4 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q4 | BaseRT | Apple M4 Max Metal | 583.3 TG128 @ 1 ctx | 9,998 PP512 | basecompute | ||
| 5 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8 Q8 | BaseRT | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | ||
| 6 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | ||
| 7 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 508.8 TG128 | 26,137 PP512 | arki05 | ||
| 8 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q4 | BaseRT | Apple M3 Ultra Metal | 501.7 TG128 @ 1 ctx | 12,830 PP512 | basecompute | ||
| 9 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 479.4 TG128 | 26,346 PP512 | arki05 | ||
| 10 | Qwen3 0.6BReported as Qwen3-0.6B-cuda-q4, Qwen3-0.6B-cuda-q4mix, Qwen3-0.6B-cuda-q8, Qwen/Qwen3-0.6B · qwen Q4 | BaseRT | NVIDIA GB10 CUDA | 458.8 TG128 @ 1 ctx | 52,750 PP512 | basecompute | ||
| 11 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q8 | BaseRT | Apple M4 Max Metal | 455.8 TG128 @ 1 ctx | 9,749 PP512 | basecompute | ||
| 12 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q8 | BaseRT | Apple M3 Ultra Metal | 440.2 TG128 @ 1 ctx | 12,494 PP512 | basecompute | ||
| 13 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | ||
| 14 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 415.5 TG128 | 23,019 PP512 | lukas | ||
| 15 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | Apple M3 Ultra Metal | 410.1 TG128 @ 1 ctx | 9,066 PP512 | basecompute | ||
| 16 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | ||
| 17 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | Apple M4 Max Metal | 371.7 TG128 @ 1 ctx | 6,058 PP512 | basecompute | ||
| 18 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT ROCm | 368.6 TG128 | 23,972 PP512 | arki05 | ||
| 19 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q4 | BaseRT | Apple M4 Max Metal | 361.2 TG128 @ 1 ctx | 7,448 PP512 | basecompute | ||
| 20 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,552 PP512 | arki05 | ||
| 21 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8 Q8 | BaseRT | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | ||
| 22 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q4 | BaseRT | Apple M3 Ultra Metal | 328.0 TG128 @ 1 ctx | 9,903 PP512 | basecompute | ||
| 23 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen Q4 | BaseRT | Apple M4 Max Metal | 318.4 TG128 @ 1 ctx | 4,146 PP512 | basecompute | ||
| 24 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | AMD Radeon RX 7900 XT ROCm | 313.7 TG128 | 24,691 PP512 | arki05 | ||
| 25 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen Q4 | BaseRT | Apple M3 Ultra Metal | 307.4 TG128 @ 1 ctx | 5,420 PP512 | basecompute | ||
| 26 | Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35 Q4 | BaseRT | Apple M3 Ultra Metal | 306.7 TG128 @ 1 ctx | 1,759 PP512 | basecompute | ||
| 27 | Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35 Q4 | BaseRT | Apple M4 Max Metal | 301.7 TG128 @ 1 ctx | 1,718 PP512 | basecompute | ||
| 28 | Gemma 3 1B ITReported as gemma-3-1b-it-cuda-q4mix, gemma-3-1b-it-cuda-q8, google/gemma-3-1b-it · gemma3 Q4 | BaseRT | NVIDIA GB10 CUDA | 295.9 TG128 @ 1 ctx | 33,841 PP512 | basecompute | ||
| 29 | Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35 Q4 | BaseRT | Apple M4 Max Metal | 295.1 TG128 @ 1 ctx | 1,713 PP512 | basecompute | ||
| 30 | Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35 Q4 | BaseRT | Apple M3 Ultra Metal | 293.8 TG128 @ 1 ctx | 1,756 PP512 | basecompute | ||
| 31 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | ||
| 32 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen Q8 | BaseRT | Apple M4 Pro Metal | 284.7 TG128 @ 1 ctx | 4,787 PP512 | basecompute | ||
| 33 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | ||
| 34 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q8 | BaseRT | Apple M4 Max Metal | 281.0 TG128 @ 1 ctx | 7,527 PP512 | basecompute | ||
| 35 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M1 Pro Metal | 268.8 TG128 @ 1 ctx | 2,847 PP512 | isu | ||
| 36 | Llama 3.2 1B InstructReported as Llama-3.2-1B-Instruct-cuda-q4mix, Llama-3.2-1B-Instruct-cuda-q8, meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | NVIDIA GB10 CUDA | 261.3 TG128 @ 1 ctx | 36,647 PP512 | basecompute | ||
| 37 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 Q8 | BaseRT | Apple M3 Ultra Metal | 257.8 TG128 @ 1 ctx | 9,953 PP512 | basecompute | ||
| 38 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama Q4 | BaseRT | Apple M1 Max Metal | 242.1 TG128 @ 1 ctx | 3,372 PP512 | basecompute | ||
| 39 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen Q8 | BaseRT | Apple M3 Ultra Metal | 238.6 TG128 @ 1 ctx | 5,698 PP512 | basecompute | ||
| 40 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 |