Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 14 like-for-like pairs.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gemma-3-1b-it, gemma-3-1b-it-cuda-q4mix, gemma-3-1b-it-cuda-q8, gemma-3-1b-it-Q4, gemma-3-1b-it-Q8, gemma-4-26B-A4B-it, gemma-4-26B-A4B-it-cuda-q4mix, gemma-4-26B-A4B-it-cuda-q8, gemma-4-26B-A4B-it-Q4, gemma-4-26B-A4B-it-Q8, gemma-4-E2B-it, gemma-4-E2B-it-cuda-q4mix, gemma-4-E2B-it-cuda-q8, gemma-4-E2B-it-Q4, gemma-4-E2B-it-Q8, gemma-4-E4B-it, gemma-4-E4B-it-Q4, gemma-4-E4B-it-Q8, gpt-oss-120b, gpt-oss-120b-MXFP4, gpt-oss-120b-Q4, gpt-oss-120b-Q8, gpt-oss-20b, gpt-oss-20b-MXFP4, gpt-oss-20b-Q4, gpt-oss-20b-Q8, Llama-3.1-8B-Instruct, Llama-3.1-8B-Instruct-Q4, Llama-3.1-8B-Instruct-Q8, Llama-3.2-1B-Instruct, Llama-3.2-1B-Instruct-cuda-q4mix, Llama-3.2-1B-Instruct-cuda-q8, Llama-3.2-1B-Instruct-Q4, Llama-3.2-1B-Instruct-Q8, Llama-3.2-3B-Instruct, Llama-3.2-3B-Instruct-cuda-q4mix, Llama-3.2-3B-Instruct-cuda-q8, Llama-3.2-3B-Instruct-Q4, Llama-3.2-3B-Instruct-Q8, Mistral-7B-Instruct-v0.3, Mistral-7B-Instruct-v0.3-Q4, Mistral-7B-Instruct-v0.3-Q8, Muse-Glimmer-30B, muse-glimmer-30B-kquant-17gb, muse-glimmer-30B-kquant-dynamic, NVIDIA-Nemotron-3-Nano-30B-A3B, NVIDIA-Nemotron-3-Nano-30B-A3B-Q4, NVIDIA-Nemotron-3-Nano-30B-A3B-Q8, Ornith-1.5-9B, Qwen3-0.6B, Qwen3-0.6B-cuda-q4, Qwen3-0.6B-cuda-q4mix, Qwen3-0.6B-cuda-q8, Qwen3-0.6B-Q4, Qwen3-0.6B-Q8, Qwen3-1.7B, Qwen3-1.7B-Q4, Qwen3-1.7B-Q8, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Instruct-2507-cuda-q4mix, Qwen3-30B-A3B-Instruct-2507-cuda-q8, Qwen3-30B-A3B-Instruct-2507-Q4, Qwen3-30B-A3B-Thinking-2507, Qwen3-30B-A3B-Thinking-2507-Q4, Qwen3-30B-A3B-Thinking-2507-Q8, Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen3-4B-Instruct-2507-Q4, Qwen3-4B-Instruct-2507-Q8, Qwen3-4B-Q4, Qwen3-4B-Q8, Qwen3-4B-Thinking-2507, Qwen3-4B-Thinking-2507-Q4, Qwen3-4B-Thinking-2507-Q8, Qwen3-8B, Qwen3-8B-base-q2, Qwen3-8B-base-q3, Qwen3-8B-base-q5, Qwen3-8B-base-q6, Qwen3-8B-Q4, Qwen3-8B-Q8, Qwen3.5-122B-A10B-Q4, Qwen3.5-2B, Qwen3.5-2B-Base, Qwen3.5-2B-Base-Q4, Qwen3.5-2B-Base-Q8, Qwen3.5-2B-cuda-q4mix, Qwen3.5-2B-cuda-q8, Qwen3.5-2B-Q4, Qwen3.5-2B-Q8, Qwen3.5-35B-A3B, Qwen3.5-35B-A3B-cuda-q4mix, Qwen3.5-35B-A3B-cuda-q8, Qwen3.5-35B-A3B-Q4, Qwen3.5-35B-A3B-Q8, Qwen3.6-27B, Qwen3.6-27B-cuda-q4mix, Qwen3.6-27B-cuda-q8, Qwen3.6-27B-Q4, Qwen3.6-27B-Q8, Qwen3.6-35B-A3B, Qwen3.6-35B-A3B-cuda-q4mix, Qwen3.6-35B-A3B-cuda-q8, Qwen3.6-35B-A3B-Q4, Qwen3.6-35B-A3B-Q8, Qwen3.8-27B, Qwen3.8-27B-Q4, Qwen3.8-27B-Q4-mtp, Qwen3.8-27B-Q8, tinystories-lay8-hs512-hd8-33M | Gemma-4 12B IT (smart Q4_0, QAT-lossless), Gemma-4 26B-A4B IT (smart Q4_0, QAT-lossless), Gemma-4-E4B-It, Glm-4.6V-Flash, Gpt-Oss-20B, Llama 3.2 3B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Meta Llama 3.1 8B Instruct, ministral-8B-Instruct-2512, Models Qwen Qwen3 0.6B, Ornith-1.5-9B, Qwen 0.6b Coder, Qwen2.5 Coder 1.5B Instruct GGUF, Qwen3 0.6B Instruct, Qwen3-0.6B, Qwen3-30B-A3B-Instruct-2507, Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen3-8B, Qwen3.5-4B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B, Tinystories Lay8 Hs512 Hd8 33M |
| Chip | Apple M1 Max, Apple M1 Pro, Apple M3 Ultra, Apple M4, Apple M4 Max, Apple M4 Pro, Apple M5, Apple M5 Max, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), AMD Ryzen 5 7600 6-Core Processor, Apple M1 Pro, Apple M4 Max, Apple M4 Pro, Apple M5 Max, Apple M5 Pro, NVIDIA GeForce RTX 4060 Laptop GPU, NVIDIA H100 80GB HBM3 × 8, Tesla T10/Tesla T10/Tesla T10/Tesla T10 |
| Backend | CUDA, Metal | BLAS + Metal, CPU, CUDA, ROCm, Vulkan |
| Conditioning | idle_reset_then_warm, warmup_only | idle_reset_before_workload_process, runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-8B Apple M5 Pro | 59.7 | 49.9 | 1.20× | 1,755 | 1,416 | 1.24× | 13 / 14 |
Qwen3-0.6B Apple M5 Pro | 415.8 | 357.2 | 1.16× | 20,451 | 14,530 | 1.41× | 6 / 4 |
Llama-3.1-8B-Instruct Apple M5 Pro | 48.8 | 55.5 | 0.88× | 1,784 | 1,424 | 1.25× | 4 / 5 |
Llama-3.2-3B-Instruct Apple M5 Pro | 94.6 | 115.2 | 0.82× | 4,297 | 3,318 | 1.30× | 5 / 3 |
Qwen3-30B-A3B-Instruct-2507 Apple M5 Pro | 98.1 | 95.1 | 1.03× | 3,493 | 1,769 | 1.97× | 2 / 6 |
gemma-4-E4B-it Apple M5 Pro | 66.6 | 69.4 | 0.96× | 4,563 | 2,263 | 2.02× | 4 / 3 |
Qwen3-4B-Instruct-2507 Apple M5 Pro | 76.9 | 92.6 | 0.83× | 3,379 | 2,568 | 1.32× | 4 / 3 |
gpt-oss-20b Apple M5 Pro | 106.9 | 81.5 | 1.31× | 1,110 | 1,851 | 0.60× | 4 / 2 |
Qwen3.6-27B Apple M5 Pro | 13.3 | 14.9 | 0.90× | 360 | 375 | 0.96× | 2 / 4 |
Qwen3-0.6B Apple M5 Max | 698.5 | 409.0 | 1.71× | 33,088 | 24,674 | 1.34× | 3 / 2 |
Qwen3-4B Apple M5 Pro | 103.9 | 93.6 | 1.11× | 3,384 | 2,705 | 1.25× | 3 / 1 |
Qwen3.6-35B-A3B Apple M5 Pro | 113.2 | 61.5 | 1.84× | 1,275 | 1,673 | 0.76× | 1 / 3 |
Qwen3.8-27B Apple M4 Max | 31.6 | 22.9 | 1.38× | 222 | 243 | 0.91× | 1 / 1 |
tinystories-lay8-hs512-hd8-33M Apple M5 Pro | 2,068.0 | 1,443.3 | 1.43× | 190,107 | 124,181 | 1.53× | 1 / 1 |
Top 25 of 375 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
ivnle/tinystories-lay8-hs512-hd8-33M BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Max Metal | 698.5 TG128 | 32,977 PP512 | lukas | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 583.3 TG128 @ 1 ctx | 9,998 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 501.7 TG128 @ 1 ctx | 12,830 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 458.8 TG128 @ 1 ctx | 20,970 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-cuda-q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 456.1 TG128 @ 1 ctx | 52,750 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 455.8 TG128 @ 1 ctx | 9,749 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 440.2 TG128 @ 1 ctx | 12,494 PP512 | basecompute | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 415.5 TG128 | 19,770 PP512 | lukas | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 414.2 TG128 | 23,019 PP512 | lukas | View Benchmark | |
Qwen3-0.6B-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 413.2 TG128 @ 1 ctx | 49,865 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 410.1 TG128 @ 1 ctx | 9,066 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 371.7 TG128 @ 1 ctx | 5,963 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 361.2 TG128 @ 1 ctx | 7,448 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
Llama-3.2-1B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 345.2 TG128 @ 1 ctx | 9,055 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 328.0 TG128 @ 1 ctx | 9,903 PP512 | basecompute | View Benchmark | |
Llama-3.2-1B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 320.5 TG128 @ 1 ctx | 6,058 PP512 | basecompute | View Benchmark | |
Qwen3-1.7B-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 318.4 TG128 @ 1 ctx | 4,146 PP512 | basecompute | View Benchmark |
Top 25 of 149 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Tinystories Lay8 Hs512 Hd8 33M llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 508.8 TG128 | 26,137 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 491.9 TG128 | 25,964 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 479.4 TG128 | 26,346 PP512 | arki05 | View Benchmark | |
Models Qwen Qwen3 0.6B llama.cpp b10360 (48d22e295)Q4_K_M | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | View Benchmark | |
Qwen3 0.6B Instruct llama.cpp b10360 (48d22e295)Q8_0 | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 368.6 TG128 | 23,972 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 367.7 TG128 | 23,718 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 313.7 TG128 | 24,691 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 229.7 TG128 | 5,737 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 223.4 TG128 | 5,721 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 198.1 TG128 | 2,837 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 195.6 TG128 | 3,292 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 186.7 TG128 | 2,868 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 185.1 TG128 | 4,783 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 184.5 TG128 | 2,474 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 181.3 TG128 | 2,827 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 180.3 TG128 | 4,793 PP512 | arki05 | View Benchmark | |
Qwen2.5 Coder 1.5B Instruct GGUF llama.cpp b10902 (df03399b8)Q4_K_M | NVIDIA GeForce RTX 4060 Laptop GPU CUDA | 178.8 TG128 | 9,465 PP512 | sarthak247 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 172.0 TG128 | 6,980 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 171.9 TG128 | 6,832 PP512 | arki05 | View Benchmark |