Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gpt-oss-20b, gpt-oss-20b-MXFP4, gpt-oss-20b-Q4, gpt-oss-20b-Q8 | Gpt-Oss-20B |
| Chip | Apple M1 Max, Apple M3 Ultra, Apple M4 Max, Apple M4 Pro, Apple M5 Max, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro, Tesla T10/Tesla T10/Tesla T10/Tesla T10 |
| Backend | CUDA, Metal | BLAS + Metal, CUDA, ROCm, Vulkan |
| Conditioning | warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
gpt-oss-20b Apple M5 Pro | 106.9 | 81.5 | 1.31× | 1,110 | 1,851 | 0.60× | 4 / 2 |
Top 20 of 20 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 185.1 TG128 @ 1 ctx | 2,554 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 175.8 TG128 @ 1 ctx | 2,627 PP512 | basecompute | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Max Metal | 169.0 TG128 | 2,179 PP512 | lukas | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M3 Ultra Metal | 163.9 TG128 @ 1 ctx | 2,590 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 160.5 TG128 @ 1 ctx | 1,709 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 146.9 TG128 @ 1 ctx | 1,700 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M4 Max Metal | 146.7 TG128 @ 1 ctx | 1,715 PP512 | basecompute | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 110.9 TG128 | 1,155 PP512 | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 110.1 TG128 | 1,156 PP512 | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4mxfp4 | Apple M5 Pro Metal | 103.6 TG128 | 1,013 PP512 | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 100.6 TG128 | 1,065 PP512 | arki05 | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 82.3 TG128 @ 1 ctx | 881 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 79.3 TG128 @ 1 ctx | 732 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 76.7 TG128 @ 1 ctx | 880 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M1 Max Metal | 76.6 TG128 @ 1 ctx | 890 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 73.7 TG128 @ 1 ctx | 731 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M4 Pro Metal | 72.7 TG128 @ 1 ctx | 734 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | NVIDIA GB10 CUDA | 60.2 TG128 @ 1 ctx | 4,424 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 57.3 TG128 @ 1 ctx | 4,292 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | NVIDIA GB10 CUDA | 51.4 TG128 @ 1 ctx | 4,421 PP512 | basecompute | View Benchmark |
Top 7 of 7 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Gpt-Oss-20B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 195.6 TG128 | 3,292 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 160.1 TG128 | 3,502 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b1 (e64c0ea)Q4_0 | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 157.9 TG128 | 3,038 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)F16 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 146.8 TG128 | 3,262 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)F16 | AMD Radeon RX 7900 XT ROCm | 128.8 TG128 | 3,604 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 96.3 TG128 | 1,849 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)F16 | Apple M5 Pro BLAS + Metal | 66.6 TG128 | 1,852 PP512 | arki05 | View Benchmark |