Chips
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for NVIDIA GB10. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
NVIDIA GB10 CUDA |
89.7 tok/s TG128 @ 1 ctx |
14,073 tok/s PP512 |
| 2,891.3 MiB |
| basecompute |
| View Benchmark |
Qwen3-4B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 91.3 tok/s TG128 @ 1 ctx | 2,822 tok/s PP512 | 2,862.5 MiB | basecompute | View Benchmark |
Qwen3.5-2B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 113.5 tok/s TG128 @ 1 ctx | 6,754 tok/s PP512 | 2,592.4 MiB | basecompute | View Benchmark |
Llama-3.2-3B-Instruct-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 110.5 tok/s TG128 @ 1 ctx | 14,651 tok/s PP512 | 2,444.5 MiB | basecompute | View Benchmark |
Qwen3.5-2B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 151.5 tok/s TG128 @ 1 ctx | 6,837 tok/s PP512 | 1,984.5 MiB | basecompute | View Benchmark |
Llama-3.2-1B-Instruct-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 173.7 tok/s TG128 @ 1 ctx | 34,296 tok/s PP512 | 1,742.4 MiB | basecompute | View Benchmark |
Llama-3.2-1B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 172.8 tok/s TG128 @ 1 ctx | 35,941 tok/s PP512 | 1,742.3 MiB | basecompute | View Benchmark |
Qwen3.5-2B-Base-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 181.9 tok/s TG128 @ 1 ctx | 4,191 tok/s PP512 | 1,769.2 MiB | basecompute | View Benchmark |
Qwen3.5-2B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 181.6 tok/s TG128 @ 1 ctx | 4,185 tok/s PP512 | 1,768.9 MiB | basecompute | View Benchmark |
Qwen3-1.7B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 202.5 tok/s TG128 @ 1 ctx | 7,394 tok/s PP512 | 1,727.8 MiB | basecompute | View Benchmark |
Llama-3.2-1B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 195.3 tok/s TG128 @ 1 ctx | 33,397 tok/s PP512 | 1,549.9 MiB | basecompute | View Benchmark |
gemma-3-1b-it-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 193.2 tok/s TG128 @ 1 ctx | 33,641 tok/s PP512 | 1,777.5 MiB | basecompute | View Benchmark |
Llama-3.2-1B-Instruct-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 261.3 tok/s TG128 @ 1 ctx | 36,647 tok/s PP512 | 1,252.3 MiB | basecompute | View Benchmark |
gemma-3-1b-it-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 295.9 tok/s TG128 @ 1 ctx | 14,350 tok/s PP512 | 1,317.4 MiB | basecompute | View Benchmark |
gemma-3-1b-it-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 268.0 tok/s TG128 @ 1 ctx | 33,841 tok/s PP512 | 1,402.8 MiB | basecompute | View Benchmark |
Qwen3-0.6B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 458.8 tok/s TG128 @ 1 ctx | 20,970 tok/s PP512 | 1,041.2 MiB | basecompute | View Benchmark |
Qwen3-0.6B-cuda-q4 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 456.1 tok/s TG128 @ 1 ctx | 52,750 tok/s PP512 | 1,236.9 MiB | basecompute | View Benchmark |
Qwen3-0.6B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 413.2 tok/s TG128 @ 1 ctx | 49,865 tok/s PP512 | 1,130.4 MiB | basecompute | View Benchmark |
Qwen3-0.6B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 315.8 tok/s TG128 @ 1 ctx | 51,373 tok/s PP512 | 1,359.7 MiB | basecompute | View Benchmark |
basecompute/gpt-oss-120b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 40.3 tok/s TG128 | 2,679 tok/s PP512 | 40,978.3 MiB | basecompute | View Benchmark |
Qwen/Qwen3.6-27B BaseRTQ4· dense-q4mix-cuda This report predates artifact identity, so ComputeArena mapped its model name to a model family by hand. The exact model bytes were not verified. | NVIDIA GB10 CUDA | 12.4 tok/s TG128 | 1,140 tok/s PP512 | 15,361.0 MiB | basecompute | View Benchmark |