Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gemma-3-1b-it-Q4, gemma-3-1b-it-Q8, gemma-4-26B-A4B-it-Q4, gemma-4-26B-A4B-it-Q8, gemma-4-E2B-it-Q4, gemma-4-E2B-it-Q8, gemma-4-E4B-it-Q4, gemma-4-E4B-it-Q8, gpt-oss-20b-MXFP4, gpt-oss-20b-Q4, gpt-oss-20b-Q8, Llama-3.1-8B-Instruct-Q4, Llama-3.1-8B-Instruct-Q8, Llama-3.2-1B-Instruct-Q4, Llama-3.2-1B-Instruct-Q8, Llama-3.2-3B-Instruct-Q4, Llama-3.2-3B-Instruct-Q8, Mistral-7B-Instruct-v0.3-Q4, Mistral-7B-Instruct-v0.3-Q8, muse-glimmer-30B-kquant-17gb, muse-glimmer-30B-kquant-dynamic, NVIDIA-Nemotron-3-Nano-30B-A3B-Q4, Qwen3-0.6B-Q4, Qwen3-0.6B-Q8, Qwen3-1.7B-Q4, Qwen3-1.7B-Q8, Qwen3-30B-A3B-Instruct-2507-Q4, Qwen3-30B-A3B-Thinking-2507-Q4, Qwen3-4B-Instruct-2507-Q4, Qwen3-4B-Instruct-2507-Q8, Qwen3-4B-Q4, Qwen3-4B-Q8, Qwen3-4B-Thinking-2507-Q4, Qwen3-4B-Thinking-2507-Q8, Qwen3-8B-Q4, Qwen3-8B-Q8, Qwen3.5-2B-Base-Q4, Qwen3.5-2B-Base-Q8, Qwen3.5-2B-Q4, Qwen3.5-2B-Q8, Qwen3.5-35B-A3B-Q4, Qwen3.6-27B-Q4, Qwen3.6-35B-A3B-Q4, Qwen3.8-27B-Q4, Qwen3.8-27B-Q4-mtp | Qwen3.8-27B |
| Backend | Metal | BLAS + Metal |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3.8-27B Apple M4 Max | 31.6 | 22.9 | 1.38× | 222 | 243 | 0.91× | 1 / 1 |
Top 25 of 45 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 583.3 TG128 @ 1 ctx | 9,998 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 455.8 TG128 @ 1 ctx | 9,749 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 371.7 TG128 @ 1 ctx | 5,963 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 361.2 TG128 @ 1 ctx | 7,448 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 320.5 TG128 @ 1 ctx | 6,058 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 318.4 TG128 @ 1 ctx | 4,146 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 301.7 TG128 @ 1 ctx | 1,718 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Base-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 295.1 TG128 @ 1 ctx | 1,713 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 281.0 TG128 @ 1 ctx | 7,527 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 224.1 TG128 @ 1 ctx | 4,025 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Base-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 208.0 TG128 @ 1 ctx | 1,727 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 205.4 TG128 @ 1 ctx | 1,721 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 178.8 TG128 @ 1 ctx | 2,202 PP512 | basecompute | View Benchmark | |
gemma-4-E2B-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 175.3 TG128 @ 1 ctx | 7,670 PP512 | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 169.1 TG128 @ 1 ctx | 1,632 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 160.5 TG128 @ 1 ctx | 1,709 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 157.5 TG128 @ 1 ctx | 1,725 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 146.9 TG128 @ 1 ctx | 1,700 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M4 Max Metal | 146.7 TG128 @ 1 ctx | 1,715 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 140.1 TG128 @ 1 ctx | 1,797 PP512 | basecompute | View Benchmark | |
Qwen3.5-35B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 140.1 TG128 @ 1 ctx | 895 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 139.8 TG128 @ 1 ctx | 1,789 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 139.7 TG128 @ 1 ctx | 886 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 133.1 TG128 @ 1 ctx | 2,218 PP512 | basecompute | View Benchmark | |
gemma-4-E2B-it-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 132.2 TG128 @ 1 ctx | 7,448 PP512 | basecompute | View Benchmark |
Top 1 of 1 run by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.8-27B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M4 Max BLAS + Metal | 22.9 TG128 | 243 PP512 | fabian | View Benchmark |