Hardware2

NVIDIA GeForce RTX 4060 Laptop GPU

3 runs
Best decode · TG128
178.8tok/s
Best prefill · PP512
9,465tok/s

NVIDIA H100 80GB HBM3 × 8

1 run
Best decode · TG128
95.3tok/s
Best prefill · PP512
2,513tok/s
Badges4
Record1
Pioneer1
Coverage2
4 decode records
Fastest result in a tested configuration
Record
First run
Submitted a benchmark report
Pioneer
Model explorer
4 models covered
Coverage
Hardware explorer
2 devices covered
Coverage
Submissions4
Sort by
Order
Prefill size
4 submissions · PP512 selected
Model / runtime / formatDeviceBackendReport
Glm-4.6V-Flash
llama.cppQ4_0
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GeForce RTX 4060 Laptop GPUCUDA
43.1 tok/s
TG128
2,079 tok/s
PP512
Qwen2.5 Coder 1.5B Instruct GGUF
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GeForce RTX 4060 Laptop GPUCUDA
178.8 tok/s
TG128
9,465 tok/s
PP512
Qwen3.8-27B
llama.cppQ4_1
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA H100 80GB HBM3 × 8CUDA
95.3 tok/s
TG128
2,513 tok/s
PP512
ministral-8B-Instruct-2512
llama.cppQ4_K_M
The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution.
NVIDIA GeForce RTX 4060 Laptop GPUCUDA
45.1 tok/s
TG128
2,034 tok/s
PP512