Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | llama.cpp | BaseRT |
|---|---|---|
| Model | Qwen3-30B-A3B-Instruct-2507 | Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Instruct-2507-cuda-q4mix, Qwen3-30B-A3B-Instruct-2507-cuda-q8, Qwen3-30B-A3B-Instruct-2507-Q4 |
| Chip | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M5 Pro, NVIDIA GB10 |
| Backend | BLAS + Metal, ROCm, Vulkan | CUDA, Metal |
| Conditioning | runtime_native_warmup | warmup_only |
| Decode workload | TG128 | TG128, TG128 @ 1 ctx |
| Harness schema | computearena-measurements/1 | basert-benchmark-harness/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read llama.cpp relative to BaseRT.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-30B-A3B-Instruct-2507 Apple M5 Pro | 95.1 | 98.1 | 0.97× | 1,769 | 3,493 | 0.51× | 6 / 2 |
Top 14 of 14 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 198.1 TG128 | 2,837 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 186.7 TG128 | 2,868 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 184.5 TG128 | 2,474 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 181.3 TG128 | 2,827 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT ROCm | 137.4 TG128 | 2,386 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT ROCm | 128.8 TG128 | 2,596 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 128.6 TG128 | 2,661 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 128.5 TG128 | 2,820 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q2_K | Apple M5 Pro BLAS + Metal | 102.7 TG128 | 1,893 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q3_K_M | Apple M5 Pro BLAS + Metal | 97.3 TG128 | 1,708 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 95.2 TG128 | 1,845 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 95.0 TG128 | 1,831 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q5_K_M | Apple M5 Pro BLAS + Metal | 84.7 TG128 | 1,644 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q6_K | Apple M5 Pro BLAS + Metal | 78.9 TG128 | 1,646 PP512 | arki05 | View Benchmark |
Top 9 of 9 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 139.8 TG128 @ 1 ctx | 1,789 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 137.1 TG128 @ 1 ctx | 2,287 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-30B-A3B-Instruct-2507 BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 104.0 TG128 | 3,680 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 99.0 TG128 @ 1 ctx | 7,439 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-30B-A3B-Instruct-2507 BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 92.2 TG128 | 3,305 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 87.7 TG128 @ 1 ctx | 901 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 80.7 TG128 @ 1 ctx | 4,715 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 61.2 TG128 @ 1 ctx | 7,253 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 49.4 TG128 @ 1 ctx | 7,315 PP512 | basecompute | View Benchmark |