Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Llama 3.1 8B Instruct. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 124.4 tok/s TG128 @ 1 ctx | 1,438 tok/s PP512 | 6,174.1 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 75.3 tok/s TG128 @ 1 ctx | 1,392 tok/s PP512 | 10,318.1 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 63.7 tok/s TG128 @ 1 ctx | 509 tok/s PP512 | 6,179.6 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 38.6 tok/s TG128 @ 1 ctx | 509 tok/s PP512 | 10,323.0 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 56.1 tok/s TG128 @ 1 ctx | 394 tok/s PP512 | 6,178.0 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 32.1 tok/s TG128 @ 1 ctx | 392 tok/s PP512 | 10,321.6 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 60.8 tok/s TG128 @ 1 ctx | 947 tok/s PP512 | 10,321.6 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 105.5 tok/s TG128 @ 1 ctx | 946 tok/s PP512 | 6,176.5 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 25.4 tok/s TG128 @ 1 ctx | 7,299 tok/s PP512 | 8,425.1 MiB | basecompute | View Benchmark | |
Llama-3.1-8B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 53.1 tok/s TG128 @ 1 ctx | 7,567 tok/s PP512 | 5,491.8 MiB | basecompute | View Benchmark | |
Meta Llama 3.1 8B Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 136.1 tok/s TG128 | 3,390 tok/s PP512 | 10.9 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 68.6 tok/s TG128 | 2,286 tok/s PP512 | 10,218.0 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 97.7 tok/s TG128 | 2,512 tok/s PP512 | 6,422.7 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 119.8 tok/s TG128 | 2,756 tok/s PP512 | 4,895.2 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 122.7 tok/s TG128 | 2,743 tok/s PP512 | 4,825.0 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 126.7 tok/s TG128 | 2,959 tok/s PP512 | 4,591.7 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 66.2 tok/s TG128 | 3,292 tok/s PP512 | 11,925.0 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 87.4 tok/s TG128 | 2,851 tok/s PP512 | 8,129.5 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 97.5 tok/s TG128 | 3,150 tok/s PP512 | 6,601.9 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 97.1 tok/s TG128 | 3,172 tok/s PP512 | 6,531.9 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ4_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 117.4 tok/s TG128 | 3,233 tok/s PP512 | 6,298.3 MiB | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 35.6 tok/s TG128 | 1,774 tok/s PP512 | 10,330.8 MiB | arki05 | View Benchmark | |
basecompute/Llama-3.1-8B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 63.2 tok/s TG128 | 1,795 tok/s PP512 | 6,186.8 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 29.6 tok/s TG128 | 1,465 tok/s PP512 | 12,319.5 MiB | arki05 | View Benchmark | |
Llama-3.1-8B-Instruct llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 43.4 tok/s TG128 | 1,421 tok/s PP512 | 8,524.1 MiB | arki05 | View Benchmark |