Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | Qwen3-4B-Instruct-2507, Qwen3-4B-Instruct-2507-Q4, Qwen3-4B-Instruct-2507-Q8 | Qwen3-4B-Instruct-2507 |
| Chip | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M4 Pro, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro |
| Backend | CUDA, Metal | BLAS + Metal, ROCm, Vulkan |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-4B-Instruct-2507 Apple M5 Pro | 76.9 | 92.6 | 0.83× | 3,379 | 2,568 | 1.32× | 4 / 3 |
Top 14 of 14 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 153.7 TG128 @ 1 ctx | 2,529 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 121.9 TG128 @ 1 ctx | 1,711 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 118.3 TG128 @ 1 ctx | 2,557 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 104.3 TG128 @ 1 ctx | 1,723 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 91.3 TG128 @ 1 ctx | 922 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 90.4 TG128 | 3,331 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 89.4 TG128 | 3,359 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 77.9 TG128 @ 1 ctx | 727 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 73.7 TG128 @ 1 ctx | 11,246 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 72.1 TG128 @ 1 ctx | 925 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 64.3 TG128 | 3,398 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 62.8 TG128 | 3,405 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 56.2 TG128 @ 1 ctx | 731 PP512 | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRT 0.2.6Q8 | NVIDIA GB10 CUDA | 54.8 TG128 @ 1 ctx | 10,970 PP512 | basecompute | View Benchmark |
Top 9 of 9 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 185.1 TG128 | 4,783 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 180.3 TG128 | 4,793 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 147.8 TG128 | 5,364 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 147.1 TG128 | 5,337 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 134.3 TG128 | 4,824 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 110.8 TG128 | 5,581 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 94.4 TG128 | 2,568 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 92.6 TG128 | 2,566 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 60.8 TG128 | 2,686 PP512 | arki05 | View Benchmark |