Chips
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Apple M3 Ultra. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
Apple M3 Ultra Metal |
72.5 tok/s TG128 @ 1 ctx |
1,338 tok/s PP512 |
| 10,360.9 MiB |
| basecompute |
| View Benchmark |
gemma-4-E4B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 75.4 tok/s TG128 @ 1 ctx | 3,499 tok/s PP512 | 5,523.3 MiB | basecompute | View Benchmark |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 75.3 tok/s TG128 @ 1 ctx | 1,392 tok/s PP512 | 10,318.1 MiB | basecompute | View Benchmark |
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 78.9 tok/s TG128 @ 1 ctx | 1,388 tok/s PP512 | 10,257.2 MiB | basecompute | View Benchmark |
gemma-4-E2B-it-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 116.6 tok/s TG128 @ 1 ctx | 8,610 tok/s PP512 | 2,608.4 MiB | basecompute | View Benchmark |
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 124.4 tok/s TG128 @ 1 ctx | 1,438 tok/s PP512 | 6,174.1 MiB | basecompute | View Benchmark |
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 113.6 tok/s TG128 @ 1 ctx | 1,357 tok/s PP512 | 6,254.9 MiB | basecompute | View Benchmark |
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 94.6 tok/s TG128 @ 1 ctx | 3,610 tok/s PP512 | 3,323.8 MiB | basecompute | View Benchmark |
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 129.3 tok/s TG128 @ 1 ctx | 1,440 tok/s PP512 | 6,113.4 MiB | basecompute | View Benchmark |
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 118.2 tok/s TG128 @ 1 ctx | 2,559 tok/s PP512 | 6,033.0 MiB | basecompute | View Benchmark |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 118.3 tok/s TG128 @ 1 ctx | 2,557 tok/s PP512 | 6,032.5 MiB | basecompute | View Benchmark |
Qwen3-4B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 116.8 tok/s TG128 @ 1 ctx | 2,441 tok/s PP512 | 6,032.7 MiB | basecompute | View Benchmark |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 154.4 tok/s TG128 @ 1 ctx | 3,316 tok/s PP512 | 4,768.0 MiB | basecompute | View Benchmark |
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 137.4 tok/s TG128 @ 1 ctx | 8,996 tok/s PP512 | 1,581.7 MiB | basecompute | View Benchmark |
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 155.6 tok/s TG128 @ 1 ctx | 2,582 tok/s PP512 | 3,894.0 MiB | basecompute | View Benchmark |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 153.7 tok/s TG128 @ 1 ctx | 2,529 tok/s PP512 | 3,894.0 MiB | basecompute | View Benchmark |
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 195.4 tok/s TG128 @ 1 ctx | 3,241 tok/s PP512 | 3,085.0 MiB | basecompute | View Benchmark |
Qwen3.5-2B-Base-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 222.4 tok/s TG128 @ 1 ctx | 1,696 tok/s PP512 | 2,949.2 MiB | basecompute | View Benchmark |
Qwen3.5-2B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 222.9 tok/s TG128 @ 1 ctx | 1,756 tok/s PP512 | 2,949.7 MiB | basecompute | View Benchmark |
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 165.7 tok/s TG128 @ 1 ctx | 2,461 tok/s PP512 | 3,894.0 MiB | basecompute | View Benchmark |
Qwen3-1.7B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 238.6 tok/s TG128 @ 1 ctx | 5,698 tok/s PP512 | 2,935.2 MiB | basecompute | View Benchmark |
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 345.2 tok/s TG128 @ 1 ctx | 9,055 tok/s PP512 | 1,670.5 MiB | basecompute | View Benchmark |
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 293.8 tok/s TG128 @ 1 ctx | 1,756 tok/s PP512 | 2,064.4 MiB | basecompute | View Benchmark |
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 307.4 tok/s TG128 @ 1 ctx | 5,420 tok/s PP512 | 2,080.1 MiB | basecompute | View Benchmark |
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 306.7 tok/s TG128 @ 1 ctx | 1,759 tok/s PP512 | 2,066.4 MiB | basecompute | View Benchmark |