Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | Apple M5 Pro | NVIDIA GB10 |
|---|---|---|
| Model | gemma-3-1b-it | gemma-3-1b-it-cuda-q4mix, gemma-3-1b-it-cuda-q8, gemma-3-1b-it-Q4 |
| Runtime version | BaseRT 0.2.4 | BaseRT 0.2.6 |
| Quantisation | Q4, Q8 | Q4 |
| Backend | Metal | CUDA |
| Decode workload | TG128 | TG128 @ 1 ctx |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Pro relative to NVIDIA GB10.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
gemma-3-1b-it-cuda-q8 BaseRTQ4 | 281.7 | 268.0 | 1.05× | 14,392 | 33,641 | 0.43× | 2 / 3 |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 270.4 TG128 | 13,717 PP512 | arki05 | View Benchmark | |
basecompute/gemma-3-1b-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 210.3 TG128 | 15,061 PP512 | arki05 | View Benchmark |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
gemma-3-1b-it-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 295.9 TG128 @ 1 ctx | 14,350 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 268.0 TG128 @ 1 ctx | 33,841 PP512 | basecompute | View Benchmark | |
gemma-3-1b-it-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 193.2 TG128 @ 1 ctx | 33,641 PP512 | basecompute | View Benchmark |