Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 3 like-for-like pairs.
| Facet | Apple M5 Pro | Apple M1 Max |
|---|---|---|
| Model | gpt-oss-20b, Gpt-Oss-20B | gpt-oss-20b-MXFP4, gpt-oss-20b-Q4, gpt-oss-20b-Q8 |
| Runtime | BaseRT, llama.cpp | BaseRT |
| Runtime version | BaseRT 0.2.4, llama.cpp b10809 (5266f24da) | BaseRT 0.2.6 |
| Quantisation | F16, mxfp4, Q4, Q4_K_M, Q8 | mxfp4, Q4, Q8 |
| Backend | BLAS + Metal, Metal | Metal |
| Conditioning | runtime_native_warmup, warmup_only | warmup_only |
| Decode workload | TG128 | TG128 @ 1 ctx |
| Harness schema | basert-benchmark-harness/1, computearena-measurements/1 | basert-benchmark-harness/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Pro relative to Apple M1 Max.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
gpt-oss-20b BaseRTQ4 | 110.5 | 82.3 | 1.34× | 1,156 | 881 | 1.31× | 2 / 1 |
gpt-oss-20b BaseRTQ8 | 100.6 | 76.7 | 1.31× | 1,065 | 880 | 1.21× | 1 / 1 |
gpt-oss-20b BaseRTmxfp4 | 103.6 | 76.6 | 1.35× | 1,013 | 890 | 1.14× | 1 / 1 |
Top 6 of 6 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 110.9 TG128 | 1,155 PP512 | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 110.1 TG128 | 1,156 PP512 | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4mxfp4 | Apple M5 Pro Metal | 103.6 TG128 | 1,013 PP512 | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 100.6 TG128 | 1,065 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 96.3 TG128 | 1,849 PP512 | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cpp b10809 (5266f24da)F16 | Apple M5 Pro BLAS + Metal | 66.6 TG128 | 1,852 PP512 | arki05 | View Benchmark |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
gpt-oss-20b-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 82.3 TG128 @ 1 ctx | 881 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 76.7 TG128 @ 1 ctx | 880 PP512 | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRT 0.2.6mxfp4 | Apple M1 Max Metal | 76.6 TG128 @ 1 ctx | 890 PP512 | basecompute | View Benchmark |