Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Llama 3.2 3B Instruct. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 95.0 tok/s TG128 @ 1 ctx | 943 tok/s PP512 | 3,089.2 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 71.2 tok/s TG128 @ 1 ctx | 943 tok/s PP512 | 4,772.4 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 195.4 tok/s TG128 @ 1 ctx | 3,241 tok/s PP512 | 3,085.0 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 154.4 tok/s TG128 @ 1 ctx | 3,316 tok/s PP512 | 4,768.0 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 110.7 tok/s TG128 @ 1 ctx | 1,229 tok/s PP512 | 3,091.7 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 88.4 tok/s TG128 @ 1 ctx | 1,240 tok/s PP512 | 4,774.1 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 133.1 tok/s TG128 @ 1 ctx | 2,218 tok/s PP512 | 4,770.8 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 178.8 tok/s TG128 @ 1 ctx | 2,202 tok/s PP512 | 3,087.5 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 69.4 tok/s TG128 @ 1 ctx | 14,988 tok/s PP512 | 3,611.8 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q8 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 68.7 tok/s TG128 @ 1 ctx | 14,871 tok/s PP512 | 3,611.6 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 89.7 tok/s TG128 @ 1 ctx | 14,073 tok/s PP512 | 2,891.3 MiB | basecompute | View Benchmark | |
Llama-3.2-3B-Instruct-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 110.5 tok/s TG128 @ 1 ctx | 14,651 tok/s PP512 | 2,444.5 MiB | basecompute | View Benchmark | |
Llama 3.2 3B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 167.5 tok/s TG128 | 5,846 tok/s PP512 | 3,394.5 MiB | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 223.4 tok/s TG128 | 5,721 tok/s PP512 | 437.0 MiB | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 229.7 tok/s TG128 | 5,737 tok/s PP512 | 442.4 MiB | arki05 | View Benchmark | |
Llama 3.2 3B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 142.3 tok/s TG128 | 7,210 tok/s PP512 | 5,102.3 MiB | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 171.9 tok/s TG128 | 6,832 tok/s PP512 | 3,804.4 MiB | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 172.0 tok/s TG128 | 6,980 tok/s PP512 | 3,764.9 MiB | arki05 | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 79.5 tok/s TG128 | 4,343 tok/s PP512 | 4,779.3 MiB | arki05 | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 106.8 tok/s TG128 | 4,297 tok/s PP512 | 3,096.1 MiB | arki05 | View Benchmark | |
Llama 3.2 3B Instruct llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 76.6 tok/s TG128 | 3,516 tok/s PP512 | 5,237.3 MiB | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 115.2 tok/s TG128 | 3,315 tok/s PP512 | 3,940.1 MiB | arki05 | View Benchmark | |
Llama-3.2-3B-Instruct llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 117.7 tok/s TG128 | 3,318 tok/s PP512 | 3,904.2 MiB | arki05 | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 94.6 tok/s TG128 | 3,845 tok/s PP512 | 3,096.0 MiB | arki05 | View Benchmark | |
basecompute/Llama-3.2-3B-Instruct BaseRTQ4· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 79.5 tok/s TG128 | 4,344 tok/s PP512 | 4,779.6 MiB | arki05 | View Benchmark |