Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | Apple M5 Pro | NVIDIA GB10 |
|---|---|---|
| Model | Qwen3.6-27B | Qwen3.6-27B, Qwen3.6-27B-cuda-q4mix, Qwen3.6-27B-cuda-q8, Qwen3.6-27B-Q4 |
| Runtime | BaseRT, llama.cpp | BaseRT |
| Runtime version | BaseRT 0.2.4, llama.cpp b10809 (5266f24da) | BaseRT 0.2.4, BaseRT 0.2.6 |
| Quantisation | Q3_K_M, Q4, Q4_K_M, Q6_K, Q8 | Q4 |
| Backend | BLAS + Metal, Metal | CUDA |
| Conditioning | runtime_native_warmup, warmup_only | warmup_only |
| Decode workload | TG128 | TG128, TG128 @ 1 ctx |
| Harness schema | basert-benchmark-harness/1, computearena-measurements/1 | basert-benchmark-harness/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Pro relative to NVIDIA GB10.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3.6-27B-cuda-q8 BaseRTQ4 | 16.1 | 12.3 | 1.31× | 364 | 1,140 | 0.32× | 1 / 5 |
Top 6 of 6 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.6-27B llama.cpp b10809 (5266f24da)Q3_K_M | Apple M5 Pro BLAS + Metal | 16.7 TG128 | 373 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.6-27B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 16.1 TG128 | 364 PP512 | arki05 | View Benchmark | |
Qwen3.6-27B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 15.2 TG128 | 376 PP512 | arki05 | View Benchmark | |
Qwen3.6-27B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 14.5 TG128 | 374 PP512 | arki05 | View Benchmark | |
Qwen3.6-27B llama.cpp b10809 (5266f24da)Q6_K | Apple M5 Pro BLAS + Metal | 11.9 TG128 | 377 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.6-27B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 10.6 TG128 | 355 PP512 | arki05 | View Benchmark |
Top 5 of 5 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3.6-27B-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 13.0 TG128 @ 1 ctx | 344 PP512 | basecompute | View Benchmark | |
Qwen/Qwen3.6-27B BaseRT 0.2.4Q4 | NVIDIA GB10 CUDA | 12.4 TG128 | 1,140 PP512 | basecompute | View Benchmark | |
Qwen3.6-27B-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 12.3 TG128 @ 1 ctx | 1,127 PP512 | basecompute | View Benchmark | |
Qwen3.6-27B-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 8.4 TG128 @ 1 ctx | 1,146 PP512 | basecompute | View Benchmark | |
Qwen3.6-27B-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 7.5 TG128 @ 1 ctx | 1,141 PP512 | basecompute | View Benchmark |