Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | llama.cpp | BaseRT |
|---|---|---|
| Model | Llama-3.1-8B-Instruct, Meta Llama 3.1 8B Instruct | Llama-3.1-8B-Instruct, Llama-3.1-8B-Instruct-Q4, Llama-3.1-8B-Instruct-Q8 |
| Chip | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro, Tesla T10/Tesla T10/Tesla T10/Tesla T10 | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M4 Pro, Apple M5 Pro, NVIDIA GB10 |
| Backend | BLAS + Metal, CUDA, ROCm, Vulkan | CUDA, Metal |
| Conditioning | runtime_native_warmup | warmup_only |
| Decode workload | TG128 | TG128, TG128 @ 1 ctx |
| Harness schema | computearena-measurements/1 | basert-benchmark-harness/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read llama.cpp relative to BaseRT.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Llama-3.1-8B-Instruct Apple M5 Pro | 55.5 | 48.8 | 1.14× | 1,424 | 1,784 | 0.80× | 5 / 4 |
Top 16 of 16 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Meta Llama 3.1 8B Instruct llama.cpp b1 (e64c0ea)Q4_K_M | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 136.1 TG128 | 3,390 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 126.7 TG128 | 2,959 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 122.7 TG128 | 2,743 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 119.8 TG128 | 2,756 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_0 | AMD Radeon RX 7900 XT ROCm | 117.4 TG128 | 3,233 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q6_K | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 97.7 TG128 | 2,512 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 97.5 TG128 | 3,150 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 97.1 TG128 | 3,172 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q6_K | AMD Radeon RX 7900 XT ROCm | 87.4 TG128 | 2,851 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 68.6 TG128 | 2,286 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 66.2 TG128 | 3,292 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 58.7 TG128 | 1,538 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 55.6 TG128 | 1,424 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 55.5 TG128 | 1,423 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q6_K | Apple M5 Pro BLAS + Metal | 43.4 TG128 | 1,421 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 29.6 TG128 | 1,465 PP512 | arki05 | View Benchmark |
Top 14 of 14 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Llama-3.1-8B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 124.4 TG128 @ 1 ctx | 1,438 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 105.5 TG128 @ 1 ctx | 946 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 75.3 TG128 @ 1 ctx | 1,392 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 63.7 TG128 @ 1 ctx | 509 PP512 | basecompute | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 63.2 TG128 | 1,795 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 62.0 TG128 | 1,795 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 60.8 TG128 @ 1 ctx | 947 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 56.1 TG128 @ 1 ctx | 394 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 53.1 TG128 @ 1 ctx | 7,567 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 38.6 TG128 @ 1 ctx | 509 PP512 | basecompute | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 35.6 TG128 | 1,774 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 35.5 TG128 | 1,773 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 32.1 TG128 @ 1 ctx | 392 PP512 | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 25.4 TG128 @ 1 ctx | 7,299 PP512 | basecompute | View Benchmark |