Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 1 like-for-like pair.
| Facet | BaseRT | llama.cpp |
|---|---|---|
| Model | gemma-4-E4B-it, gemma-4-E4B-it-Q4, gemma-4-E4B-it-Q8 | Gemma-4-E4B-It |
| Chip | Apple M1 Max, Apple M1 Pro, Apple M3 Ultra, Apple M4, Apple M4 Max, Apple M4 Pro, Apple M5 Max, Apple M5 Pro, NVIDIA GB10 | AMD Radeon RX 7900 XT, AMD Radeon RX 7900 XT (RADV NAVI31), Apple M5 Pro |
| Backend | CUDA, Metal | BLAS + Metal, ROCm, Vulkan |
| Cooldown | off, on | off |
| Conditioning | idle_reset_then_warm, warmup_only | runtime_native_warmup |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
| Harness schema | basert-benchmark-harness/1 | computearena-measurements/1 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read BaseRT relative to llama.cpp.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
gemma-4-E4B-it Apple M5 Pro | 66.6 | 69.4 | 0.96× | 4,563 | 2,263 | 2.02× | 4 / 3 |
Top 17 of 17 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/gemma-4-E4B-it BaseRT 0.2.5Q4 | Apple M5 Max Metal | 121.1 TG128 | 8,156 PP512 | lukas | View Benchmark | |
gemma-4-E4B-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Max Metal | 109.3 TG128 @ 1 ctx | 2,552 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q4 BaseRT 0.2.6Q4 | Apple M3 Ultra Metal | 94.6 TG128 @ 1 ctx | 3,610 PP512 | basecompute | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 82.6 TG128 | 4,656 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 81.9 TG128 | 4,636 PP512 | arki05 | View Benchmark | |
gemma-4-E4B-it-Q8 BaseRT 0.2.6Q8 | Apple M4 Max Metal | 77.4 TG128 @ 1 ctx | 2,481 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q4 BaseRT 0.2.6Q4 | Apple M1 Max Metal | 75.9 TG128 @ 1 ctx | 1,394 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q8 BaseRT 0.2.6Q8 | Apple M3 Ultra Metal | 75.4 TG128 @ 1 ctx | 3,499 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q4 BaseRT 0.2.6Q4 | NVIDIA GB10 CUDA | 71.3 TG128 @ 1 ctx | 3,674 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q4 BaseRT 0.2.6Q4 | Apple M4 Pro Metal | 69.1 TG128 @ 1 ctx | 1,094 PP512 | basecompute | View Benchmark | |
gemma-4-E4B-it-Q8 BaseRT 0.2.6Q8 | Apple M1 Max Metal | 52.0 TG128 @ 1 ctx | 1,368 PP512 | basecompute | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 51.3 TG128 | 4,490 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 51.0 TG128 | 4,462 PP512 | arki05 | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q4 | Apple M1 Pro Metal | 45.9 TG128 | 565 PP512 | lukas | View Benchmark | |
gemma-4-E4B-it-Q8 BaseRT 0.2.6Q8 | Apple M4 Pro Metal | 44.5 TG128 @ 1 ctx | 1,081 PP512 | basecompute | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q4 | Apple M4 Metal | 15.8 TG128 | 259 PP512 | pyopyo | View Benchmark | |
basecompute/gemma-4-E4B-it BaseRT 0.2.4Q4 | Apple M4 Metal | 15.7 TG128 | 205 PP512 | pyopyo | View Benchmark |
Top 9 of 9 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 126.6 TG128 | 4,142 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 122.8 TG128 | 4,161 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 111.0 TG128 | 4,443 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | AMD Radeon RX 7900 XT ROCm | 109.3 TG128 | 4,369 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 95.6 TG128 | 4,214 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q8_0 | AMD Radeon RX 7900 XT ROCm | 84.3 TG128 | 4,564 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 72.1 TG128 | 2,263 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 69.4 TG128 | 2,235 PP512 | arki05 | View Benchmark | |
Gemma-4-E4B-It llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 47.9 TG128 | 2,346 PP512 | arki05 | View Benchmark |