Hardware2
NVIDIA GeForce RTX 4060 Laptop GPU
3 runsBest decode · TG128
178.8tok/s
Best prefill · PP512
9,465tok/s
NVIDIA H100 80GB HBM3 × 8
1 runBest decode · TG128
95.3tok/s
Best prefill · PP512
2,513tok/s
Badges4
Record1
Pioneer1
Coverage2
4 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
Model explorer
4 models covered
Coverage
Hardware explorer
2 devices covered
Coverage
Submissions4
Sort by
Order
Prefill size
4 submissions · PP512 selected
| Model / runtime / format | Device | Backend | Report | |||
|---|---|---|---|---|---|---|
Glm-4.6V-Flash llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GeForce RTX 4060 Laptop GPU | CUDA | 43.1 tok/s TG128 | 2,079 tok/s PP512 | ||
Qwen2.5 Coder 1.5B Instruct GGUF llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GeForce RTX 4060 Laptop GPU | CUDA | 178.8 tok/s TG128 | 9,465 tok/s PP512 | ||
Qwen3.8-27B llama.cppQ4_1 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA H100 80GB HBM3 × 8 | CUDA | 95.3 tok/s TG128 | 2,513 tok/s PP512 | ||
ministral-8B-Instruct-2512 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GeForce RTX 4060 Laptop GPU | CUDA | 45.1 tok/s TG128 | 2,034 tok/s PP512 |