Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 2 like-for-like pairs.
| Facet | Tesla T10/Tesla T10/Tesla T10/Tesla T10 | Apple M5 Pro |
|---|---|---|
| Model | Gpt-Oss-20B, Meta Llama 3.1 8B Instruct, Ornith-1.5-9B, Qwen3-8B, Qwen3.6-27B, Qwen3.8-27B | gemma-3-1b-it, gemma-4-26B-A4B-it, gemma-4-E2B-it, gemma-4-E4B-it, Gemma-4-E4B-It, gpt-oss-20b, Gpt-Oss-20B, Llama 3.2 3B Instruct, Llama-3.1-8B-Instruct, Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct, Mistral-7B-Instruct-v0.3, Muse-Glimmer-30B, NVIDIA-Nemotron-3-Nano-30B-A3B, Qwen3-0.6B, Qwen3-1.7B, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-4B, Qwen3-4B-Instruct-2507, Qwen3-4B-Thinking-2507, Qwen3-8B, Qwen3-8B-base-q2, Qwen3-8B-base-q3, Qwen3-8B-base-q5, Qwen3-8B-base-q6, Qwen3.5-2B, Qwen3.5-2B-Base, Qwen3.5-35B-A3B, Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3.8-27B, Tinystories Lay8 Hs512 Hd8 33M, tinystories-lay8-hs512-hd8-33M |
| Runtime | llama.cpp | BaseRT, llama.cpp |
| Runtime version | llama.cpp b1 (e64c0ea) | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10809 (5266f24da), llama.cpp b9960 (a935fbffe) |
| Quantisation | Q4_0, Q4_K_M | BF16, F16, mxfp4, passthrough_gguf, Q2, Q2_K, Q3, Q3_K_M, Q4, Q4_0, Q4_1, Q4_K_M, Q5, Q5_K_M, Q6, Q6_K, Q8, Q8_0 |
| Backend | CUDA | BLAS + Metal, Metal |
| Cooldown | off | off, on |
| Conditioning | runtime_native_warmup | idle_reset_then_warm, runtime_native_warmup, warmup_only |
| Decode workload | TG128 | TG128, TG128 @ 1 ctx |
| Harness schema | computearena-measurements/1 | basert-benchmark-harness/1, computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Tesla T10/Tesla T10/Tesla T10/Tesla T10 relative to Apple M5 Pro.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Llama-3.1-8B-Instruct llama.cppQ4_K_M | 136.1 | 55.5 | 2.45× | 3,390 | 1,423 | 2.38× | 1 / 2 |
Qwen3-8B llama.cppQ4_K_M | 124.7 | 54.3 | 2.30× | 3,087 | 1,399 | 2.21× | 1 / 2 |
Top 7 of 7 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Gpt-Oss-20B llama.cpp b1 (e64c0ea)Q4_0 | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 157.9 TG128 | 3,038 PP512 | arki05 | View Benchmark | |
Meta Llama 3.1 8B Instruct llama.cpp b1 (e64c0ea)Q4_K_M | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 136.1 TG128 | 3,390 PP512 | arki05 | View Benchmark | |
Qwen3-8B llama.cpp b1 (e64c0ea)Q4_K_M | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 124.7 TG128 | 3,087 PP512 | arki05 | View Benchmark | |
Ornith-1.5-9B llama.cpp b1 (e64c0ea)Q4_K_M | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 104.8 TG128 | 2,935 PP512 | arki05 | View Benchmark | |
Qwen3.8-27B llama.cpp b1 (e64c0ea)Q4_0 | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 48.1 TG128 | 910 PP512 | arki05 | View Benchmark | |
Qwen3.6-27B llama.cpp b1 (e64c0ea)Q4_0 | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 48.0 TG128 | 1,174 PP512 | arki05 | View Benchmark | |
Qwen3.8-27B llama.cpp b1 (e64c0ea)Q4_K_M | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 46.0 TG128 | 1,132 PP512 | arki05 | View Benchmark |
Top 25 of 129 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
ivnle/tinystories-lay8-hs512-hd8-33M BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | View Benchmark | |
Tinystories Lay8 Hs512 Hd8 33M llama.cpp b10809 (5266f24da)Q4_0 | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 311.0 TG128 | 18,474 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 270.4 TG128 | 13,717 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 229.1 TG128 | 11,673 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 218.2 TG128 | 2,645 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 217.7 TG128 | 2,641 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 210.3 TG128 | 7,268 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 210.3 TG128 | 15,061 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 206.3 TG128 | 10,528 PP512 | arki05 | View Benchmark | |
basecompute/Llama-3.2-1B-Instruct BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 205.3 TG128 | 12,067 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E2B-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 144.6 TG128 | 13,240 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-1.7B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 139.7 TG128 | 8,043 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3.5-2B-Base BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 131.4 TG128 | 2,642 PP512 | arki05 | View Benchmark |