Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Ratio of side medians · no like-for-like pairs.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gemma-4-E4B-it, Qwen3-0.6B | Qwen 0.6b Coder |
| Backend | Metal | BLAS + Metal |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
No model, quantisation, and stack combination has runs on both runtimes. Widen the filters, or submit a run that mirrors the other side.
Top 2 of 2 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRT 0.2.6Q4 | Apple M1 Pro Metal | 268.8 TG128 @ 1 ctx | 2,847 PP512 | isu | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q4 | Apple M1 Pro Metal | 45.9 TG128 | 565 PP512 | lukas | View Benchmark |
Top 1 of 1 run by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen 0.6b Coder llama.cpp b10894 (d344123fe)Q2_K | Apple M1 Pro BLAS + Metal | 151.6 TG128 | 2,620 PP512 | lukas | View Benchmark |