Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | Qwen3-4B, Qwen3-4B-Q4, Qwen3-4B-Q8 | Qwen3-4B |
| Chip | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M4 Pro, Apple M5 Pro, NVIDIA GB10 | Apple M5 Pro |
| Backend | CUDA, Metal | BLAS + Metal |
| Cooldown | off, on | off |
| Conditioning | idle_reset_then_warm, warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-4B Apple M5 Pro | 103.9 | 93.6 | 1.11× | 3,384 | 2,705 | 1.25× | 3 / 1 |
Top 12 of 12 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 165.7 TG128 @ 1 ctx | 2,461 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 157.5 TG128 @ 1 ctx | 1,725 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 116.8 TG128 @ 1 ctx | 2,441 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-4B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 107.3 TG128 | 3,384 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 105.4 TG128 @ 1 ctx | 1,712 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-4B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 103.9 TG128 | 3,372 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 101.9 TG128 @ 1 ctx | 927 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 91.3 TG128 @ 1 ctx | 2,822 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 90.3 TG128 @ 1 ctx | 734 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 71.3 TG128 @ 1 ctx | 920 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-4B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 64.7 TG128 | 3,412 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 56.3 TG128 @ 1 ctx | 726 PP512 | basecompute | View Benchmark |
Top 1 of 1 run by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-4B llama.cpp b9960 (a935fbffe)Q4_0 | Apple M5 Pro BLAS + Metal | 93.6 TG128 | 2,705 PP512 | isu | View Benchmark |