Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for GPT OSS 20B. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
basecompute/gpt-oss-20b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 169.0 tok/s TG128 | 2,179 tok/s PP512 | 7,882.2 MiB | lukas | View Benchmark | |
basecompute/gpt-oss-20b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 110.9 tok/s TG128 | 1,155 tok/s PP512 | 7,888.4 MiB | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 100.6 tok/s TG128 | 1,065 tok/s PP512 | 8,240.3 MiB | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRTmxfp4· default-bf16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 103.6 tok/s TG128 | 1,013 tok/s PP512 | 10,492.4 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 66.6 tok/s TG128 | 1,852 tok/s PP512 | 13,776.5 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 96.3 tok/s TG128 | 1,849 tok/s PP512 | 11,709.7 MiB | arki05 | View Benchmark | |
basecompute/gpt-oss-20b BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 110.1 tok/s TG128 | 1,156 tok/s PP512 | 7,882.1 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 160.1 tok/s TG128 | 3,502 tok/s PP512 | 13,008.5 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 128.8 tok/s TG128 | 3,604 tok/s PP512 | 15,074.1 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 195.6 tok/s TG128 | 3,292 tok/s PP512 | 11,257.0 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 146.8 tok/s TG128 | 3,262 tok/s PP512 | 13,324.5 MiB | arki05 | View Benchmark | |
Gpt-Oss-20B llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 157.9 tok/s TG128 | 3,038 tok/s PP512 | 11.0 MiB | arki05 | View Benchmark | |
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 57.3 tok/s TG128 @ 1 ctx | 4,292 tok/s PP512 | 13,579.2 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 51.4 tok/s TG128 @ 1 ctx | 4,421 tok/s PP512 | 15,001.2 MiB | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 60.2 tok/s TG128 @ 1 ctx | 4,424 tok/s PP512 | 16,377.1 MiB | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 72.7 tok/s TG128 @ 1 ctx | 734 tok/s PP512 | 10,012.8 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 73.7 tok/s TG128 @ 1 ctx | 731 tok/s PP512 | 8,231.5 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 79.3 tok/s TG128 @ 1 ctx | 732 tok/s PP512 | 7,880.1 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 160.5 tok/s TG128 @ 1 ctx | 1,709 tok/s PP512 | 7,879.2 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 146.9 tok/s TG128 @ 1 ctx | 1,700 tok/s PP512 | 8,231.6 MiB | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 146.7 tok/s TG128 @ 1 ctx | 1,715 tok/s PP512 | 10,483.6 MiB | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 76.6 tok/s TG128 @ 1 ctx | 890 tok/s PP512 | 10,485.2 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 76.7 tok/s TG128 @ 1 ctx | 880 tok/s PP512 | 8,233.5 MiB | basecompute | View Benchmark | |
gpt-oss-20b-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 82.3 tok/s TG128 @ 1 ctx | 881 tok/s PP512 | 7,880.5 MiB | basecompute | View Benchmark | |
gpt-oss-20b-MXFP4 BaseRTmxfp4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 163.9 tok/s TG128 @ 1 ctx | 2,590 tok/s PP512 | 10,481.7 MiB | basecompute | View Benchmark |