Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | Apple M5 Max | Apple M1 Max |
|---|---|---|
| Model | NVIDIA-Nemotron-3-Nano-30B-A3B | NVIDIA-Nemotron-3-Nano-30B-A3B-Q4, NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 |
| Runtime version | BaseRT 0.2.4 | BaseRT 0.2.6 |
| Quantisation | Q4 | Q4, Q8 |
| Cooldown | off, on | off |
| Conditioning | idle_reset_then_warm, warmup_only | warmup_only |
| Decode workload | TG128 | TG128 @ 1 ctx |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Max relative to Apple M1 Max.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 BaseRTQ4 | 179.2 | 110.3 | 1.62× | 4,942 | 870 | 5.68× | 6 / 1 |
Top 6 of 6 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 184.4 TG128 | 4,984 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 183.7 TG128 | 4,941 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 180.8 TG128 | 4,951 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 177.6 TG128 | 4,527 PP512 | basecompute | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 177.0 TG128 | 4,821 PP512 | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 109.2 TG128 | 4,943 PP512 | lukas | View Benchmark |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 110.3 TG128 @ 1 ctx | 870 PP512 | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 61.1 TG128 @ 1 ctx | 857 PP512 | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 60.6 TG128 @ 1 ctx | 858 PP512 | basecompute | View Benchmark |