Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gemma-3-1b-it, gemma-4-26B-A4B-it, gemma-4-E2B-it, gemma-4-E4B-it, gpt-oss-120b, gpt-oss-20b, Muse-Glimmer-30B, NVIDIA-Nemotron-3-Nano-30B-A3B, Qwen3-0.6B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B | Gemma-4 12B IT (smart Q4_0, QAT-lossless), Models Qwen Qwen3 0.6B, Qwen3 0.6B Instruct |
| Backend | Metal | BLAS + Metal |
| Cooldown | off, on | off |
| Conditioning | idle_reset_then_warm, warmup_only | runtime_native_warmup |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-0.6B Apple M5 Max | 698.5 | 409.0 | 1.71× | 33,088 | 24,674 | 1.34× | 3 / 2 |
Top 24 of 24 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Max Metal | 698.5 TG128 | 32,977 PP512 | lukas | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 415.5 TG128 | 19,770 PP512 | lukas | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 414.2 TG128 | 23,019 PP512 | lukas | View Benchmark | |
basecompute/gemma-4-E2B-it BaseRT 0.2.4Q4 | Apple M5 Max Metal | 197.5 TG128 | 20,099 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 184.4 TG128 | 4,984 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 183.7 TG128 | 4,941 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 180.8 TG128 | 4,951 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 177.6 TG128 | 4,527 PP512 | basecompute | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 177.0 TG128 | 4,821 PP512 | lukas | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Max Metal | 169.0 TG128 | 2,179 PP512 | lukas | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 121.1 TG128 | 8,156 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.6-35B-A3B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 119.6 TG128 | 1,540 PP512 | lukas | View Benchmark | |
basecompute/gpt-oss-120b BaseRT 0.2.4Q4 | Apple M5 Max Metal | 114.3 TG128 | 1,401 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 109.2 TG128 | 4,943 PP512 | lukas | View Benchmark | |
basecompute/gemma-4-26B-A4B-it BaseRT 0.2.4Q4 | Apple M5 Max Metal | 92.0 TG128 | 4,055 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 34.0 TG128 | 589 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.6-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 31.3 TG128 | 603 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 21.5 TG128 | 590 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 21.5 TG128 | 594 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 20.7 TG128 | 599 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 19.7 TG128 | 589 PP512 | lukas | View Benchmark | |
basecompute/Muse-Glimmer-30B BaseRT 0.2.4passthrough_gguf | Apple M5 Max Metal | 15.7 TG128 | 830 PP512 | lukas | View Benchmark |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Models Qwen Qwen3 0.6B llama.cpp b10360 (48d22e295)Q4_K_M | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | View Benchmark | |
Qwen3 0.6B Instruct llama.cpp b10360 (48d22e295)Q8_0 | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | View Benchmark | |
Gemma-4 12B IT (smart Q4_0, QAT-lossless) llama.cpp b10901 (28ff09582)Q4_0 | Apple M5 Max BLAS + Metal | 54.2 TG128 | 1,777 PP512 | lukas | View Benchmark |