Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 28 like-for-like pairs.
| Facet | AMD Radeon RX 7900 XT | Apple M5 Pro |
|---|---|---|
| Model | Gemma-4-E4B-It, Gpt-Oss-20B, Llama 3.2 3B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Qwen3-0.6B, Qwen3-30B-A3B-Instruct-2507, Qwen3-4B-Instruct-2507, Qwen3-8B, Qwen3.6-27B, Qwen3.6-35B-A3B | gemma-3-1b-it, gemma-4-26B-A4B-it, gemma-4-E2B-it, gemma-4-E4B-it, Gemma-4-E4B-It, gpt-oss-20b, Gpt-Oss-20B, Llama 3.2 3B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Muse-Glimmer-30B, NVIDIA-Nemotron-3-Nano-30B-A3B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen3-4B-Thinking-2507, Qwen3-8B, Qwen3-8B-base-q2, Qwen3-8B-base-q3, Qwen3-8B-base-q5, Qwen3-8B-base-q6, Qwen3.5-2B, Qwen3.5-2B-Base, Qwen3.5-35B-A3B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B, Tinystories Lay8 Hs512 Hd8 33M, tinystories-lay8-hs512-hd8-33M |
| Runtime | llama.cpp | BaseRT, llama.cpp |
| Runtime version | llama.cpp b10809 (5266f24da) | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10809 (5266f24da), llama.cpp b9960 (a935fbffe) |
| Quantisation | BF16, F16, Q2_K, Q3_K_M, Q4_0, Q4_1, Q4_K_M, Q5_K_M, Q6_K, Q8_0 | BF16, F16, mxfp4, passthrough_gguf, Q2, Q2_K, Q3, Q3_K_M, Q4, Q4_0, Q4_1, Q4_K_M, Q5, Q5_K_M, Q6, Q6_K, Q8, Q8_0 |
| Backend | ROCm | BLAS + Metal, Metal |
| Cooldown | off | off, on |
| Conditioning | runtime_native_warmup | idle_reset_then_warm, runtime_native_warmup, warmup_only |
| Decode workload | TG128 | TG128, TG128 @ 1 ctx |
| Harness schema | computearena-measurements/1 | basert-benchmark-harness/1, computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read AMD Radeon RX 7900 XT relative to Apple M5 Pro.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-0.6B llama.cppQ4_K_M | 368.2 | 358.1 | 1.03× | 23,845 | 14,509 | 1.64× | 2 / 3 |
gemma-4-E4B-it llama.cppQ4_K_M | 110.2 | 70.7 | 1.56× | 4,406 | 2,249 | 1.96× | 2 / 2 |
Llama-3.1-8B-Instruct llama.cppQ4_K_M | 97.3 | 55.5 | 1.75× | 3,161 | 1,423 | 2.22× | 2 / 2 |
Llama-3.2-3B-Instruct llama.cppQ4_K_M | 172.0 | 116.5 | 1.48× | 6,906 | 3,316 | 2.08× | 2 / 2 |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M | 128.5 | 95.1 | 1.35× | 2,741 | 1,838 | 1.49× | 2 / 2 |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M | 147.4 | 93.5 | 1.58× | 5,350 | 2,567 | 2.08× | 2 / 2 |
Qwen3.6-27B llama.cppQ4_K_M | 29.6 | 14.9 | 1.99× | 841 | 375 | 2.24× | 2 / 2 |
Qwen3-8B llama.cppQ2_K | 108.9 | 71.2 | 1.53× | 2,816 | 1,471 | 1.91× | 2 / 2 |
Qwen3-8B llama.cppQ3_K_M | 91.6 | 61.2 | 1.50× | 2,887 | 1,416 | 2.04× | 2 / 2 |
Qwen3-8B llama.cppQ4_K_M | 93.5 | 54.3 | 1.72× | 3,093 | 1,399 | 2.21× | 2 / 2 |
Qwen3-8B llama.cppQ5_K_M | 87.2 | 47.7 | 1.83× | 2,853 | 1,349 | 2.11× | 2 / 2 |
Qwen3-8B llama.cppQ6_K | 82.0 | 41.0 | 2.00× | 2,863 | 1,414 | 2.02× | 2 / 2 |
Qwen3-8B llama.cppQ8_0 | 69.1 | 31.8 | 2.17× | 3,224 | 1,460 | 2.21× | 2 / 2 |
gemma-4-E4B-it llama.cppQ8_0 | 84.3 | 47.9 | 1.76× | 4,564 | 2,346 | 1.95× | 1 / 1 |
gpt-oss-20b llama.cppF16 | 128.8 | 66.6 | 1.93× | 3,604 | 1,852 | 1.95× | 1 / 1 |
gpt-oss-20b llama.cppQ4_K_M | 160.1 | 96.3 | 1.66× | 3,502 | 1,849 | 1.89× | 1 / 1 |
Llama-3.1-8B-Instruct llama.cppQ4_0 | 117.4 | 58.7 | 2.00× | 3,233 | 1,538 | 2.10× | 1 / 1 |
Llama-3.1-8B-Instruct llama.cppQ6_K | 87.4 | 43.4 | 2.02× | 2,851 | 1,421 | 2.01× | 1 / 1 |
Llama-3.1-8B-Instruct llama.cppQ8_0 | 66.2 | 29.6 | 2.24× | 3,292 | 1,465 | 2.25× | 1 / 1 |
Llama-3.2-3B-Instruct llama.cppQ8_0 | 142.3 | 76.6 | 1.86× | 7,210 | 3,516 | 2.05× | 1 / 1 |
Qwen3-0.6B llama.cppQ8_0 | 313.7 | 283.0 | 1.11× | 24,691 | 14,942 | 1.65× | 1 / 1 |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K | 137.4 | 102.7 | 1.34× | 2,386 | 1,893 | 1.26× | 1 / 1 |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M | 128.8 | 97.3 | 1.32× | 2,596 | 1,708 | 1.52× | 1 / 1 |
Qwen3-4B-Instruct-2507 llama.cppQ8_0 | 110.8 | 60.8 | 1.82× | 5,581 | 2,686 | 2.08× | 1 / 1 |
Qwen3.6-27B llama.cppQ3_K_M | 29.7 | 16.7 | 1.78× | 807 | 373 | 2.16× | 1 / 1 |
Qwen3.6-35B-A3B llama.cppQ3_K_M | 100.3 | 67.5 | 1.49× | 2,501 | 1,737 | 1.44× | 1 / 1 |
Qwen3-8B llama.cppBF16 | 43.5 | 18.9 | 2.30× | 3,133 | 1,376 | 2.28× | 1 / 1 |
Qwen3-8B llama.cppQ4_1 | 107.3 | 51.9 | 2.07× | 2,894 | 1,535 | 1.89× | 1 / 1 |
Top 25 of 41 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 368.6 TG128 | 23,972 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 367.7 TG128 | 23,718 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 313.7 TG128 | 24,691 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 172.0 TG128 | 6,980 PP512 | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 171.9 TG128 | 6,832 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 160.1 TG128 | 3,502 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 147.8 TG128 | 5,364 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 147.1 TG128 | 5,337 PP512 | arki05 | View Benchmark | |
Llama 3.2 3B Instruct llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 142.3 TG128 | 7,210 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT ROCm | 137.4 TG128 | 2,386 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)F16 | AMD Radeon RX 7900 XT ROCm | 128.8 TG128 | 3,604 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT ROCm | 128.8 TG128 | 2,596 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 128.6 TG128 | 2,661 PP512 | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 128.5 TG128 | 2,820 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_0 | AMD Radeon RX 7900 XT ROCm | 117.4 TG128 | 3,233 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT ROCm | 112.3 TG128 | 2,798 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 111.0 TG128 | 4,443 PP512 | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 110.8 TG128 | 5,581 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 109.3 TG128 | 4,369 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_1 | AMD Radeon RX 7900 XT ROCm | 107.3 TG128 | 2,894 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT ROCm | 105.5 TG128 | 2,834 PP512 | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT ROCm | 100.3 TG128 | 2,501 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 97.5 TG128 | 3,150 PP512 | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 97.1 TG128 | 3,172 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 93.7 TG128 | 3,110 PP512 | arki05 | View Benchmark |
Top 25 of 129 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
ivnle/tinystories-lay8-hs512-hd8-33M BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | View Benchmark | |
Tinystories Lay8 Hs512 Hd8 33M llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 311.0 TG128 | 18,474 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 270.4 TG128 | 13,717 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 229.1 TG128 | 11,673 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 218.2 TG128 | 2,645 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 217.7 TG128 | 2,641 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 210.3 TG128 | 7,268 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 210.3 TG128 | 15,061 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 206.3 TG128 | 10,528 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 205.3 TG128 | 12,067 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E2B-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 144.6 TG128 | 13,240 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 139.7 TG128 | 8,043 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 131.4 TG128 | 2,642 PP512 | arki05 | View Benchmark |