Chips
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for NVIDIA GB10. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
NVIDIA GB10 CUDA |
30.6 tok/s TG128 @ 1 ctx |
799 tok/s PP512 |
| 40,333.7 MiB |
| basecompute |
| View Benchmark |
gpt-oss-120b-MXFP4 BaseRTmxfp4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 41.4 tok/s TG128 @ 1 ctx | 2,622 tok/s PP512 | 42,371.5 MiB | basecompute | View Benchmark |
Qwen3.5-122B-A10B-Q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 31.2 tok/s TG128 @ 1 ctx | 795 tok/s PP512 | 41,236.1 MiB | basecompute | View Benchmark |
gpt-oss-120b-MXFP4 BaseRTmxfp4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 41.7 tok/s TG128 @ 1 ctx | 2,626 tok/s PP512 | 42,187.0 MiB | basecompute | View Benchmark |
gpt-oss-120b-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 35.8 tok/s TG128 @ 1 ctx | 2,638 tok/s PP512 | 37,583.4 MiB | basecompute | View Benchmark |
gpt-oss-120b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 39.7 tok/s TG128 @ 1 ctx | 2,619 tok/s PP512 | 38,800.7 MiB | basecompute | View Benchmark |
gpt-oss-120b-Q8 BaseRTQ8 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 35.5 tok/s TG128 @ 1 ctx | 2,278 tok/s PP512 | 40,877.5 MiB | basecompute | View Benchmark |
Qwen3.6-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 55.4 tok/s TG128 @ 1 ctx | 2,181 tok/s PP512 | 28,760.0 MiB | basecompute | View Benchmark |
Qwen3.5-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 54.7 tok/s TG128 @ 1 ctx | 2,447 tok/s PP512 | 22,036.2 MiB | basecompute | View Benchmark |
Qwen3.6-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 54.4 tok/s TG128 @ 1 ctx | 2,188 tok/s PP512 | 28,757.9 MiB | basecompute | View Benchmark |
Qwen3.5-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 54.2 tok/s TG128 @ 1 ctx | 2,428 tok/s PP512 | 22,015.2 MiB | basecompute | View Benchmark |
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 61.2 tok/s TG128 @ 1 ctx | 7,253 tok/s PP512 | 30,962.8 MiB | basecompute | View Benchmark |
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 7.5 tok/s TG128 @ 1 ctx | 1,127 tok/s PP512 | 27,414.0 MiB | basecompute | View Benchmark |
Qwen3.6-27B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 8.4 tok/s TG128 @ 1 ctx | 1,146 tok/s PP512 | 27,413.3 MiB | basecompute | View Benchmark |
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 49.4 tok/s TG128 @ 1 ctx | 7,315 tok/s PP512 | 21,317.5 MiB | basecompute | View Benchmark |
Qwen3.8-27B-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 7.6 tok/s TG128 @ 1 ctx | 1,152 tok/s PP512 | 27,413.4 MiB | basecompute | View Benchmark |
Qwen3.6-27B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 7.5 tok/s TG128 @ 1 ctx | 1,141 tok/s PP512 | 27,414.9 MiB | basecompute | View Benchmark |
gemma-4-26B-A4B-it-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 40.4 tok/s TG128 @ 1 ctx | 6,045 tok/s PP512 | 10,586.0 MiB | basecompute | View Benchmark |
Qwen3.6-35B-A3B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 72.8 tok/s TG128 @ 1 ctx | 2,279 tok/s PP512 | 14,955.8 MiB | basecompute | View Benchmark |
Qwen3.5-35B-A3B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 92.4 tok/s TG128 @ 1 ctx | 2,521 tok/s PP512 | 20,907.5 MiB | basecompute | View Benchmark |
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 84.3 tok/s TG128 @ 1 ctx | 2,028 tok/s PP512 | 16,676.7 MiB | basecompute | View Benchmark |
Qwen3.6-27B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 12.3 tok/s TG128 @ 1 ctx | 1,127 tok/s PP512 | 18,021.2 MiB | basecompute | View Benchmark |
Qwen3.5-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 84.4 tok/s TG128 @ 1 ctx | 2,093 tok/s PP512 | 19,310.7 MiB | basecompute | View Benchmark |
Qwen3-30B-A3B-Instruct-2507-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 99.0 tok/s TG128 @ 1 ctx | 7,439 tok/s PP512 | 17,897.6 MiB | basecompute | View Benchmark |
NVIDIA-Nemotron-3-Nano-30B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 99.8 tok/s TG128 @ 1 ctx | 3,862 tok/s PP512 | 17,958.8 MiB | basecompute | View Benchmark |