Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Ratio of side medians · no like-for-like pairs.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gemma-3-1b-it-Q4, gemma-3-1b-it-Q8, gemma-4-E2B-it-Q4, gemma-4-E2B-it-Q8, gemma-4-E4B-it-Q4, gemma-4-E4B-it-Q8, gpt-oss-20b-MXFP4, gpt-oss-20b-Q4, gpt-oss-20b-Q8, Llama-3.1-8B-Instruct-Q4, Llama-3.1-8B-Instruct-Q8, Llama-3.2-1B-Instruct-Q4, Llama-3.2-1B-Instruct-Q8, Llama-3.2-3B-Instruct-Q4, Llama-3.2-3B-Instruct-Q8, Mistral-7B-Instruct-v0.3-Q4, Mistral-7B-Instruct-v0.3-Q8, Qwen3-0.6B-Q4, Qwen3-0.6B-Q8, Qwen3-1.7B-Q4, Qwen3-1.7B-Q8, Qwen3-4B-Instruct-2507-Q4, Qwen3-4B-Instruct-2507-Q8, Qwen3-4B-Q4, Qwen3-4B-Q8, Qwen3-4B-Thinking-2507-Q4, Qwen3-4B-Thinking-2507-Q8, Qwen3-8B-Q4, Qwen3-8B-Q8, Qwen3.5-2B-Base-Q4, Qwen3.5-2B-Base-Q8, Qwen3.5-2B-Q4, Qwen3.5-2B-Q8 | Gemma-4 26B-A4B IT (smart Q4_0, QAT-lossless) |
| Backend | Metal | BLAS + Metal |
| Cooldown | off | on |
| Conditioning | warmup_only | idle_reset_before_workload_process |
| Decode workload | TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
No model, quantisation, and stack combination has runs on both runtimes. Widen the filters, or submit a run that mirrors the other side.
Top 25 of 33 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 284.7 TG128 @ 1 ctx | 4,787 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 224.9 TG128 @ 1 ctx | 3,430 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 208.6 TG128 @ 1 ctx | 3,571 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 204.6 TG128 @ 1 ctx | 2,685 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 193.1 TG128 @ 1 ctx | 1,828 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 179.3 TG128 @ 1 ctx | 2,692 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 179.1 TG128 @ 1 ctx | 1,150 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 178.8 TG128 @ 1 ctx | 3,337 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Base-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 176.8 TG128 @ 1 ctx | 1,150 PP512 | basecompute | View Benchmark | |
gemma-4-E2B-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 126.0 TG128 @ 1 ctx | 3,722 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 125.3 TG128 @ 1 ctx | 1,822 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Base-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 113.9 TG128 @ 1 ctx | 1,149 PP512 | basecompute | View Benchmark | |
Qwen3.5-2B-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 113.5 TG128 @ 1 ctx | 1,138 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 95.0 TG128 @ 1 ctx | 943 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 90.3 TG128 @ 1 ctx | 734 PP512 | basecompute | View Benchmark | |
gemma-4-E2B-it-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 83.9 TG128 @ 1 ctx | 3,628 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 79.3 TG128 @ 1 ctx | 732 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Thinking-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 78.4 TG128 @ 1 ctx | 732 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 77.9 TG128 @ 1 ctx | 727 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 73.7 TG128 @ 1 ctx | 731 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M4 Pro Metal | 72.7 TG128 @ 1 ctx | 734 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 71.2 TG128 @ 1 ctx | 943 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 69.1 TG128 @ 1 ctx | 1,094 PP512 | basecompute | View Benchmark | |
Mistral-7B-Instruct-v0.3-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 59.0 TG128 @ 1 ctx | 394 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Thinking-2507-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 56.3 TG128 @ 1 ctx | 731 PP512 | basecompute | View Benchmark |
Top 1 of 1 run by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Gemma-4 26B-A4B IT (smart Q4_0, QAT-lossless) llama.cpp b9650 (d8a3f523c)Q4_0 | Apple M4 Pro BLAS + Metal | 74.7 TG128 | 757 PP512 | invocation | View Benchmark |