Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | NVIDIA GB10 | Apple M1 Max |
|---|---|---|
| Model | gemma-4-26B-A4B-it-cuda-q4mix, gemma-4-26B-A4B-it-cuda-q8, gemma-4-26B-A4B-it-Q4 | gemma-4-26B-A4B-it-Q4, gemma-4-26B-A4B-it-Q8 |
| Quantisation | Q4 | Q4, Q8 |
| Backend | CUDA | Metal |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read NVIDIA GB10 relative to Apple M1 Max.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
gemma-4-26B-A4B-it BaseRTQ4 | 44.0 | 51.5 | 0.85× | 6,179 | 871 | 7.09× | 3 / 1 |
Top 3 of 3 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
gemma-4-26B-A4B-it-cuda-q4mix BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 45.3 TG128 @ 1 ctx | 6,190 PP512 | basecompute | View Benchmark | |
gemma-4-26B-A4B-it-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 44.0 TG128 @ 1 ctx | 6,179 PP512 | basecompute | View Benchmark | |
gemma-4-26B-A4B-it-cuda-q8 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 40.4 TG128 @ 1 ctx | 6,045 PP512 | basecompute | View Benchmark |
Top 2 of 2 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
gemma-4-26B-A4B-it-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 51.5 TG128 @ 1 ctx | 871 PP512 | basecompute | View Benchmark | |
gemma-4-26B-A4B-it-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 45.5 TG128 @ 1 ctx | 877 PP512 | basecompute | View Benchmark |