Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | Qwen3.8-27B, Qwen3.8-27B-Q4, Qwen3.8-27B-Q8 | Qwen3.8-27B |
| Chip | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M5 Max, Apple M5 Pro, NVIDIA GB10 | Apple M4 Max, NVIDIA H100 80GB HBM3 × 8, Tesla T10/Tesla T10/Tesla T10/Tesla T10 |
| Backend | CUDA, Metal | BLAS + Metal, CUDA |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3.8-27B Apple M4 Max | 31.6 | 22.9 | 1.38× | 222 | 243 | 0.91× | 1 / 1 |
Top 16 of 16 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.8-27B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 38.0 TG128 @ 1 ctx | 282 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 34.0 TG128 | 589 PP512 | basecompute | View Benchmark | |
Qwen3.8-27B-Q4 BaseRT 0.2.4Q4 | Apple M4 Max Metal | 31.6 TG128 | 222 PP512 | fabian | View Benchmark | |
Qwen3.8-27B-Q8 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 22.1 TG128 @ 1 ctx | 280 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 21.5 TG128 | 590 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 21.5 TG128 | 594 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 20.7 TG128 | 599 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 19.7 TG128 | 589 PP512 | lukas | View Benchmark | |
Qwen3.8-27B-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 18.2 TG128 @ 1 ctx | 82 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 16.5 TG128 | 327 PP512 | arki05 | View Benchmark | |
Qwen3.8-27B-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 13.1 TG128 @ 1 ctx | 1,131 PP512 | basecompute | View Benchmark | |
Qwen3.8-27B-Q8 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 11.0 TG128 @ 1 ctx | 82 PP512 | basecompute | View Benchmark | |
Qwen3.8-27B-Q8 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 11.0 TG128 @ 1 ctx | 82 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.8-27B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 10.6 TG128 | 355 PP512 | arki05 | View Benchmark | |
Qwen3.8-27B-Q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 7.6 TG128 @ 1 ctx | 1,152 PP512 | basecompute | View Benchmark | |
Qwen3.8-27B-Q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 7.5 TG128 @ 1 ctx | 1,127 PP512 | basecompute | View Benchmark |
Top 4 of 4 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.8-27B llama.cpp b10902 (df03399b8)Q4_1 | NVIDIA H100 80GB HBM3 × 8 CUDA | 95.3 TG128 | 2,513 PP512 | sarthak247 | View Benchmark | |
Qwen3.8-27B llama.cpp b1 (e64c0ea)Q4_0 | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 48.1 TG128 | 910 PP512 | arki05 | View Benchmark | |
Qwen3.8-27B llama.cpp b1 (e64c0ea)Q4_K_M | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 46.0 TG128 | 1,132 PP512 | arki05 | View Benchmark | |
Qwen3.8-27B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M4 Max BLAS + Metal | 22.9 TG128 | 243 PP512 | fabian | View Benchmark |