Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 8 like-for-like pairs.
| Facet | Apple M5 Pro | AMD Radeon RX 7900 XT |
|---|---|---|
| Model | Qwen3-8B, Qwen3-8B-base-q2, Qwen3-8B-base-q3, Qwen3-8B-base-q5, Qwen3-8B-base-q6 | Qwen3-8B |
| Runtime | BaseRT, llama.cpp | llama.cpp |
| Runtime version | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10809 (5266f24da) | llama.cpp b10809 (5266f24da) |
| Quantisation | BF16, Q2, Q2_K, Q3, Q3_K_M, Q4, Q4_1, Q4_K_M, Q5, Q5_K_M, Q6, Q6_K, Q8, Q8_0 | BF16, Q2_K, Q3_K_M, Q4_1, Q4_K_M, Q5_K_M, Q6_K, Q8_0 |
| Backend | BLAS + Metal, Metal | ROCm |
| Conditioning | runtime_native_warmup, warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1, computearena-measurements/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Pro relative to AMD Radeon RX 7900 XT.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-8B llama.cppQ2_K | 71.2 | 108.9 | 0.65× | 1,471 | 2,816 | 0.52× | 2 / 2 |
Qwen3-8B llama.cppQ3_K_M | 61.2 | 91.6 | 0.67× | 1,416 | 2,887 | 0.49× | 2 / 2 |
Qwen3-8B llama.cppQ4_K_M | 54.3 | 93.5 | 0.58× | 1,399 | 3,093 | 0.45× | 2 / 2 |
Qwen3-8B llama.cppQ5_K_M | 47.7 | 87.2 | 0.55× | 1,349 | 2,853 | 0.47× | 2 / 2 |
Qwen3-8B llama.cppQ6_K | 41.0 | 82.0 | 0.50× | 1,414 | 2,863 | 0.49× | 2 / 2 |
Qwen3-8B llama.cppQ8_0 | 31.8 | 69.1 | 0.46× | 1,460 | 3,224 | 0.45× | 2 / 2 |
Qwen3-8B llama.cppBF16 | 18.9 | 43.5 | 0.43× | 1,376 | 3,133 | 0.44× | 1 / 1 |
Qwen3-8B llama.cppQ4_1 | 51.9 | 107.3 | 0.48× | 1,535 | 2,894 | 0.53× | 1 / 1 |
Top 25 of 27 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-8B-base-q2 BaseRT 0.2.4Q2 | Apple M5 Pro Metal | 81.8 TG128 | 1,786 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q2_K | Apple M5 Pro BLAS + Metal | 72.8 TG128 | 1,477 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q2_K | Apple M5 Pro BLAS + Metal | 69.6 TG128 | 1,465 PP512 | arki05 | View Benchmark | |
Qwen3-8B-base-q3 BaseRT 0.2.4Q3 | Apple M5 Pro Metal | 62.0 TG128 | 1,748 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q3_K_M | Apple M5 Pro BLAS + Metal | 61.5 TG128 | 1,418 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q3_K_M | Apple M5 Pro BLAS + Metal | 60.9 TG128 | 1,414 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 60.7 TG128 | 1,755 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 60.6 TG128 | 1,763 PP512 | isu | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 60.5 TG128 | 1,763 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 59.7 TG128 @ 1 ctx | 1,762 PP512 | isu | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 59.7 TG128 | 1,761 PP512 | isu | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 58.2 TG128 | 1,758 PP512 | isu | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 56.9 TG128 | 1,707 PP512 | isu | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 55.0 TG128 | 1,401 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 53.6 TG128 | 1,396 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_1 | Apple M5 Pro BLAS + Metal | 51.9 TG128 | 1,535 PP512 | arki05 | View Benchmark | |
Qwen3-8B-base-q5 BaseRT 0.2.4Q5 | Apple M5 Pro Metal | 48.4 TG128 | 1,730 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q5_K_M | Apple M5 Pro BLAS + Metal | 47.9 TG128 | 1,345 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q5_K_M | Apple M5 Pro BLAS + Metal | 47.6 TG128 | 1,353 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q6_K | Apple M5 Pro BLAS + Metal | 42.6 TG128 | 1,409 PP512 | arki05 | View Benchmark | |
Qwen3-8B-base-q6 BaseRT 0.2.4Q6 | Apple M5 Pro Metal | 42.1 TG128 | 1,723 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q6_K | Apple M5 Pro BLAS + Metal | 39.4 TG128 | 1,419 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 35.2 TG128 | 1,748 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-8B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 34.5 TG128 | 1,744 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 34.4 TG128 | 1,473 PP512 | arki05 | View Benchmark |
Top 14 of 14 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-8B llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT ROCm | 112.3 TG128 | 2,798 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_1 | AMD Radeon RX 7900 XT ROCm | 107.3 TG128 | 2,894 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q2_K | AMD Radeon RX 7900 XT ROCm | 105.5 TG128 | 2,834 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 93.7 TG128 | 3,110 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 93.2 TG128 | 3,076 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT ROCm | 91.6 TG128 | 2,864 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q3_K_M | AMD Radeon RX 7900 XT ROCm | 91.6 TG128 | 2,911 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q5_K_M | AMD Radeon RX 7900 XT ROCm | 87.5 TG128 | 2,837 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q5_K_M | AMD Radeon RX 7900 XT ROCm | 86.8 TG128 | 2,869 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q6_K | AMD Radeon RX 7900 XT ROCm | 84.0 TG128 | 2,823 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q6_K | AMD Radeon RX 7900 XT ROCm | 80.1 TG128 | 2,903 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 74.2 TG128 | 3,241 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 64.0 TG128 | 3,207 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b10809 (5266f24da)BF16 | AMD Radeon RX 7900 XT ROCm | 43.5 TG128 | 3,133 PP512 | arki05 | View Benchmark |