Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | Llama-3.2-3B-Instruct, Llama-3.2-3B-Instruct-cuda-q4mix, Llama-3.2-3B-Instruct-cuda-q8, Llama-3.2-3B-Instruct-Q4, Llama-3.2-3B-Instruct-Q8 | Llama 3.2 3B Instruct, Llama-3.2-3B-Instruct |
| Chip | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M4 Pro, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro |
| Backend | CUDA, Metal | BLAS + Metal, ROCm, Vulkan |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Llama-3.2-3B-Instruct Apple M5 Pro | 94.6 | 115.2 | 0.82× | 4,297 | 3,318 | 1.30× | 5 / 3 |
Top 17 of 17 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 195.4 TG128 @ 1 ctx | 3,241 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 178.8 TG128 @ 1 ctx | 2,202 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 154.4 TG128 @ 1 ctx | 3,316 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 133.1 TG128 @ 1 ctx | 2,218 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 110.7 TG128 @ 1 ctx | 1,229 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 110.5 TG128 @ 1 ctx | 14,651 PP512 | basecompute | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 106.8 TG128 | 4,297 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 106.1 TG128 | 4,291 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 95.0 TG128 @ 1 ctx | 943 PP512 | basecompute | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 94.6 TG128 | 3,845 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 89.7 TG128 @ 1 ctx | 14,073 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 88.4 TG128 @ 1 ctx | 1,240 PP512 | basecompute | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 79.5 TG128 | 4,343 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 79.5 TG128 | 4,344 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 71.2 TG128 @ 1 ctx | 943 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 69.4 TG128 @ 1 ctx | 14,988 PP512 | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 68.7 TG128 @ 1 ctx | 14,871 PP512 | basecompute | View Benchmark |
Top 9 of 9 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 229.7 TG128 | 5,737 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 223.4 TG128 | 5,721 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 172.0 TG128 | 6,980 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 171.9 TG128 | 6,832 PP512 | arki05 | View Benchmark | |
Llama 3.2 3B Instruct llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 167.5 TG128 | 5,846 PP512 | arki05 | View Benchmark | |
Llama 3.2 3B Instruct llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 142.3 TG128 | 7,210 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 117.7 TG128 | 3,318 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 115.2 TG128 | 3,315 PP512 | arki05 | View Benchmark | |
Llama 3.2 3B Instruct llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 76.6 TG128 | 3,516 PP512 | arki05 | View Benchmark |