Chips
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for NVIDIA GB10. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
NVIDIA GB10 CUDA |
13.1 tok/s TG128 @ 1 ctx |
1,136 tok/s PP512 |
| 17,368.9 MiB |
| basecompute |
| View Benchmark |
Qwen3.8-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 13.1 tok/s TG128 @ 1 ctx | 1,131 tok/s PP512 | 17,146.0 MiB | basecompute | View Benchmark |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 80.7 tok/s TG128 @ 1 ctx | 4,715 tok/s PP512 | 13,752.2 MiB | basecompute | View Benchmark |
Qwen3.6-27B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 13.0 tok/s TG128 @ 1 ctx | 344 tok/s PP512 | 15,228.4 MiB | basecompute | View Benchmark |
Qwen3-30B-A3B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 81.1 tok/s TG128 @ 1 ctx | 4,738 tok/s PP512 | 13,360.4 MiB | basecompute | View Benchmark |
gemma-4-26B-A4B-it-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 45.3 tok/s TG128 @ 1 ctx | 6,190 tok/s PP512 | 7,293.5 MiB | basecompute | View Benchmark |
gemma-4-26B-A4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 44.0 tok/s TG128 @ 1 ctx | 6,179 tok/s PP512 | 7,385.5 MiB | basecompute | View Benchmark |
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 60.2 tok/s TG128 @ 1 ctx | 4,424 tok/s PP512 | 16,377.1 MiB | basecompute | View Benchmark |
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 51.4 tok/s TG128 @ 1 ctx | 4,421 tok/s PP512 | 15,001.2 MiB | basecompute | View Benchmark |
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 57.3 tok/s TG128 @ 1 ctx | 4,292 tok/s PP512 | 13,579.2 MiB | basecompute | View Benchmark |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 25.4 tok/s TG128 @ 1 ctx | 7,299 tok/s PP512 | 8,425.1 MiB | basecompute | View Benchmark |
Mistral-7B-Instruct-v0.3-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 27.0 tok/s TG128 @ 1 ctx | 7,510 tok/s PP512 | 7,590.8 MiB | basecompute | View Benchmark |
gemma-4-E2B-it-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 87.7 tok/s TG128 @ 1 ctx | 19,246 tok/s PP512 | 5,786.0 MiB | basecompute | View Benchmark |
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 53.1 tok/s TG128 @ 1 ctx | 7,567 tok/s PP512 | 5,491.8 MiB | basecompute | View Benchmark |
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 50.3 tok/s TG128 @ 1 ctx | 1,452 tok/s PP512 | 5,106.7 MiB | basecompute | View Benchmark |
gemma-4-E4B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 71.3 tok/s TG128 @ 1 ctx | 3,674 tok/s PP512 | 4,738.5 MiB | basecompute | View Benchmark |
Mistral-7B-Instruct-v0.3-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 53.5 tok/s TG128 @ 1 ctx | 7,554 tok/s PP512 | 4,470.8 MiB | basecompute | View Benchmark |
gemma-4-E2B-it-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 79.0 tok/s TG128 @ 1 ctx | 18,372 tok/s PP512 | 4,988.4 MiB | basecompute | View Benchmark |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 54.8 tok/s TG128 @ 1 ctx | 10,970 tok/s PP512 | 4,520.0 MiB | basecompute | View Benchmark |
Qwen3-4B-Thinking-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 54.7 tok/s TG128 @ 1 ctx | 11,473 tok/s PP512 | 4,397.1 MiB | basecompute | View Benchmark |
Llama-3.2-3B-Instruct-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 69.4 tok/s TG128 @ 1 ctx | 14,988 tok/s PP512 | 3,611.8 MiB | basecompute | View Benchmark |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 68.7 tok/s TG128 @ 1 ctx | 14,871 tok/s PP512 | 3,611.6 MiB | basecompute | View Benchmark |
gemma-4-E2B-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 139.8 tok/s TG128 @ 1 ctx | 12,036 tok/s PP512 | 3,680.3 MiB | basecompute | View Benchmark |
Qwen3-4B-Thinking-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 73.2 tok/s TG128 @ 1 ctx | 11,020 tok/s PP512 | 3,105.8 MiB | basecompute | View Benchmark |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 73.7 tok/s TG128 @ 1 ctx | 11,246 tok/s PP512 | 3,379.4 MiB | basecompute | View Benchmark |