Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Nemotron 3 Nano 30B A3B. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 171.6 tok/s TG128 @ 1 ctx | 2,295 tok/s PP512 | 1,105.0 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 169.1 tok/s TG128 @ 1 ctx | 1,632 tok/s PP512 | 1,109.7 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 102.1 tok/s TG128 @ 1 ctx | 2,311 tok/s PP512 | 1,106.3 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 101.9 tok/s TG128 @ 1 ctx | 2,315 tok/s PP512 | 1,105.2 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 110.3 tok/s TG128 @ 1 ctx | 870 tok/s PP512 | 1,112.5 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 61.1 tok/s TG128 @ 1 ctx | 857 tok/s PP512 | 1,110.5 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 60.6 tok/s TG128 @ 1 ctx | 858 tok/s PP512 | 1,110.8 MiB | basecompute | View Benchmark | |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 99.8 tok/s TG128 @ 1 ctx | 3,862 tok/s PP512 | 17,958.8 MiB | basecompute | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 66.0 tok/s TG128 | 1,748 tok/s PP512 | 1,118.6 MiB | arki05 | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 177.0 tok/s TG128 | 4,821 tok/s PP512 | 1,141.2 MiB | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 180.8 tok/s TG128 | 4,951 tok/s PP512 | 1,141.1 MiB | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 183.7 tok/s TG128 | 4,941 tok/s PP512 | 1,140.9 MiB | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 109.2 tok/s TG128 | 4,943 tok/s PP512 | 1,140.9 MiB | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 184.4 tok/s TG128 | 4,984 tok/s PP512 | 1,111.6 MiB | lukas | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 100.1 tok/s TG128 | 2,499 tok/s PP512 | 1,119.6 MiB | arki05 | View Benchmark | |
basecompute/NVIDIA-Nemotron-3-Nano-30B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 177.6 tok/s TG128 | 4,527 tok/s PP512 | 1,967.5 MiB | basecompute | View Benchmark |