Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 2 like-for-like pairs.
Everything else matches: model, quantisation, backend, protocol, and workload are the same on both sides.
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M1 Max relative to Apple M3 Ultra.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-30B-A3B-Thinking-2507 BaseRTQ8 | 64.1 | 95.9 | 0.67× | 894 | 2,340 | 0.38× | 2 / 2 |
Qwen3-30B-A3B-Thinking-2507 BaseRTQ4 | 87.2 | 135.3 | 0.64× | 901 | 2,416 | 0.37× | 1 / 1 |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 87.2 TG128 @ 1 ctx | 901 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 64.6 TG128 @ 1 ctx | 893 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 63.7 TG128 @ 1 ctx | 895 PP512 | basecompute | View Benchmark |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 135.3 TG128 @ 1 ctx | 2,416 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 95.9 TG128 @ 1 ctx | 2,289 PP512 | basecompute | View Benchmark | |
Qwen3-30B-A3B-Thinking-2507-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 95.8 TG128 @ 1 ctx | 2,391 PP512 | basecompute | View Benchmark |