Compare
Loading community benchmarks…
Loading community benchmarks…
Put two chips, runtimes, or runtime versions side by side. Everything else is held fixed or called out, so a ratio is worth exactly what the matched runs say.
Ratios read side A relative to side B. Geometric mean over 4 like-for-like pairs.
| Facet | Apple M5 Pro | Apple M5 Max |
|---|---|---|
| Model | Qwen3-0.6B | Models Qwen Qwen3 0.6B, Qwen3 0.6B Instruct, Qwen3-0.6B |
| Runtime version | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10809 (5266f24da) | BaseRT 0.2.4, BaseRT 0.2.5, llama.cpp b10360 (48d22e295) |
| Decode workload | TG128, TG128 @ 1 ctx | TG128 |
Each row is a configuration present on both sides. Values are per-cell medians; ratios read Apple M5 Pro relative to Apple M5 Max.
| Configuration | Decode A | Decode B | Ratio | Prefill A | Prefill B | Ratio | Runs A / B |
|---|---|---|---|---|---|---|---|
Qwen3-0.6B BaseRTQ4 | 512.0 | 703.4 | 0.73× | 20,603 | 33,557 | 0.61× | 3 / 2 |
Qwen3-0.6B BaseRTQ8 | 349.8 | 549.5 | 0.64× | 20,366 | 33,088 | 0.62× | 3 / 1 |
Qwen3-0.6B llama.cppQ4_K_M | 358.1 | 439.1 | 0.82× | 14,509 | 24,365 | 0.60× | 3 / 1 |
Qwen3-0.6B llama.cppQ8_0 | 283.0 | 379.0 | 0.75× | 14,942 | 24,983 | 0.60× | 1 / 1 |
Top 10 of 10 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Pro Metal | 512.0 TG128 | 20,603 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Pro Metal | 481.5 TG128 @ 1 ctx | 19,595 PP512 | isu | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,509 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 358.1 TG128 | 14,344 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q4_K_M | Apple M5 Pro BLAS + Metal | 356.2 TG128 | 14,552 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 349.8 TG128 | 20,366 PP512 | arki05 | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Pro Metal | 311.0 TG128 | 18,474 PP512 | arki05 | View Benchmark | |
Qwen3-0.6B llama.cpp b10809 (5266f24da)Q8_0 | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | View Benchmark |
Top 5 of 5 runs by decode throughput.
| Model / format | Chip / backend | Decode | Prefill | Contributor | Date | Report |
|---|---|---|---|---|---|---|
basecompute/Qwen3-0.6B BaseRT 0.2.4Q4 | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | basecompute | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.5Q4 | Apple M5 Max Metal | 698.5 TG128 | 32,977 PP512 | lukas | View Benchmark | |
basecompute/Qwen3-0.6B BaseRT 0.2.4Q8 | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | basecompute | View Benchmark | |
Models Qwen Qwen3 0.6B llama.cpp b10360 (48d22e295)Q4_K_M | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | basecompute | View Benchmark | |
Qwen3 0.6B Instruct llama.cpp b10360 (48d22e295)Q8_0 | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | basecompute | View Benchmark |