Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 4 like-for-like pairs.
| Facet | Apple M5 Pro | AMD Radeon RX 7900 XT |
|---|---|---|
| Runtime | BaseRT, llama.cpp | llama.cpp |
| Runtime version | BaseRT 0.2.4, llama.cpp b10809 (5266f24da) | llama.cpp b10809 (5266f24da) |
| Quantisation | Q4, Q4_0, Q4_K_M, Q6_K, Q8_0 | Q4_0, Q4_K_M, Q6_K, Q8_0 |
| Backend | BLAS + Metal, Metal | ROCm |
| Conditioning | runtime_native_warmup, warmup_only | runtime_native_warmup |
| Harness schema | basert-benchmark-harness/1, computearena-measurements/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Pro relative to AMD Radeon RX 7900 XT.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Llama-3.1-8B-Instruct llama.cppQ4_K_M | 55.5 | 97.3 | 0.57× | 1,423 | 3,161 | 0.45× | 2 / 2 |
Llama-3.1-8B-Instruct llama.cppQ4_0 | 58.7 | 117.4 | 0.50× | 1,538 | 3,233 | 0.48× | 1 / 1 |
Llama-3.1-8B-Instruct llama.cppQ6_K | 43.4 | 87.4 | 0.50× | 1,421 | 2,851 | 0.50× | 1 / 1 |
Llama-3.1-8B-Instruct llama.cppQ8_0 | 29.6 | 66.2 | 0.45× | 1,465 | 3,292 | 0.44× | 1 / 1 |
Top 9 of 9 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 63.2 TG128 | 1,795 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 62.0 TG128 | 1,795 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 58.7 TG128 | 1,538 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 55.6 TG128 | 1,424 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 55.5 TG128 | 1,423 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q6_K | Apple M5 Pro BLAS + Metal | 43.4 TG128 | 1,421 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 35.6 TG128 | 1,774 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 35.5 TG128 | 1,773 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 29.6 TG128 | 1,465 PP512 | arki05 | View Benchmark |
Top 5 of 5 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_0 | AMD Radeon RX 7900 XT ROCm | 117.4 TG128 | 3,233 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 97.5 TG128 | 3,150 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 97.1 TG128 | 3,172 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q6_K | AMD Radeon RX 7900 XT ROCm | 87.4 TG128 | 2,851 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 66.2 TG128 | 3,292 PP512 | arki05 | View Benchmark |