Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | Qwen3.6-35B-A3B, Qwen3.6-35B-A3B-cuda-q4mix, Qwen3.6-35B-A3B-cuda-q8, Qwen3.6-35B-A3B-Q4, Qwen3.6-35B-A3B-Q8 | Qwen3.6-35B-A3B |
| Chip | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M5 Max, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro |
| Backend | CUDA, Metal | BLAS + Metal, ROCm, Vulkan |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3.6-35B-A3B Apple M5 Pro | 113.2 | 61.5 | 1.84× | 1,275 | 1,673 | 0.76× | 1 / 3 |
Top 13 of 13 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.6-35B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 139.7 TG128 @ 1 ctx | 886 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 133.0 TG128 @ 1 ctx | 922 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3.6-35B-A3B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 119.6 TG128 | 1,540 PP512 | lukas | View Benchmark | |
basecompute/Qwen3.6-35B-A3B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 113.2 TG128 | 1,275 PP512 | arki05 | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 110.8 TG128 @ 1 ctx | 910 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 109.5 TG128 @ 1 ctx | 913 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 94.2 TG128 @ 1 ctx | 278 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 84.3 TG128 @ 1 ctx | 2,028 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 72.8 TG128 @ 1 ctx | 2,279 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 72.2 TG128 @ 1 ctx | 232 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 70.8 TG128 @ 1 ctx | 247 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 55.4 TG128 @ 1 ctx | 2,181 PP512 | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 54.4 TG128 @ 1 ctx | 2,188 PP512 | basecompute | View Benchmark |
Top 5 of 5 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.6-35B-A3B llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 124.1 TG128 | 2,597 PP512 | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT ROCm | 100.3 TG128 | 2,501 PP512 | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cpp b10809 (5266f24da)Q3_K_M | Apple M5 Pro BLAS + Metal | 67.5 TG128 | 1,737 PP512 | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cpp b10809 (5266f24da)Q5_K_M | Apple M5 Pro BLAS + Metal | 61.5 TG128 | 1,533 PP512 | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 61.5 TG128 | 1,673 PP512 | arki05 | View Benchmark |