Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 2 like-for-like pairs.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | Qwen3-0.6B, Qwen3-0.6B-cuda-q4, Qwen3-0.6B-cuda-q4mix, Qwen3-0.6B-cuda-q8, Qwen3-0.6B-Q4, Qwen3-0.6B-Q8 | Models Qwen Qwen3 0.6B, Qwen3 0.6B Instruct, Qwen3-0.6B |
| Chip | Apple M1 Max, Apple M1 Pro, Apple M3 Ultra, Apple M4 Max, Apple M4 Pro, Apple M5, Apple M5 Max, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Max, Apple M5 Pro |
| Backend | CUDA, Metal | BLAS + Metal, ROCm, Vulkan |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-0.6B Apple M5 Pro | 415.8 | 357.2 | 1.16× | 20,451 | 14,530 | 1.41× | 6 / 4 |
Qwen3-0.6B Apple M5 Max | 698.5 | 409.0 | 1.71× | 33,088 | 24,674 | 1.34× | 3 / 2 |
Top 23 of 23 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Max Metal | 698.5 TG128 | 32,977 PP512 | lukas | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 583.3 TG128 @ 1 ctx | 9,998 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 501.7 TG128 @ 1 ctx | 12,830 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 458.8 TG128 @ 1 ctx | 20,970 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-cuda-q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 456.1 TG128 @ 1 ctx | 52,750 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 455.8 TG128 @ 1 ctx | 9,749 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 440.2 TG128 @ 1 ctx | 12,494 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 413.2 TG128 @ 1 ctx | 49,865 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 315.8 TG128 @ 1 ctx | 51,373 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 311.0 TG128 | 18,474 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 284.7 TG128 @ 1 ctx | 4,787 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.6Q4 | Apple M1 Pro Metal | 268.8 TG128 @ 1 ctx | 2,847 PP512 | isu | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 208.6 TG128 @ 1 ctx | 3,571 PP512 | basecompute | View Benchmark | |
Qwen/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Metal | 172.7 TG128 | 4,997 PP512 | skogul97 | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 100.8 TG128 @ 1 ctx | 1,851 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 74.3 TG128 @ 1 ctx | 2,225 PP512 | basecompute | View Benchmark |
Top 12 of 12 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 508.8 TG128 | 26,137 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 491.9 TG128 | 25,964 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 479.4 TG128 | 26,346 PP512 | arki05 | View Benchmark | |
Models Qwen Qwen3 0.6B llama.cpp b10360 (48d22e295)Q4_K_M | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | View Benchmark | |
Qwen3 0.6B Instruct llama.cpp b10360 (48d22e295)Q8_0 | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 368.6 TG128 | 23,972 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 367.7 TG128 | 23,718 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 313.7 TG128 | 24,691 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark |