Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Qwen3 8B. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 113.6 tok/s TG128 @ 1 ctx | 1,357 tok/s PP512 | 6,254.9 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 72.5 tok/s TG128 @ 1 ctx | 1,338 tok/s PP512 | 10,360.9 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 52.6 tok/s TG128 @ 1 ctx | 394 tok/s PP512 | 6,259.6 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 60.2 tok/s TG128 @ 1 ctx | 507 tok/s PP512 | 6,259.0 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 38.1 tok/s TG128 @ 1 ctx | 500 tok/s PP512 | 10,365.4 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 59.6 tok/s TG128 @ 1 ctx | 942 tok/s PP512 | 10,364.5 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 31.6 tok/s TG128 @ 1 ctx | 392 tok/s PP512 | 10,365.8 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 99.2 tok/s TG128 @ 1 ctx | 943 tok/s PP512 | 6,258.2 MiB | basecompute | View Benchmark | |
Qwen3-8B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 50.3 tok/s TG128 @ 1 ctx | 1,452 tok/s PP512 | 5,106.7 MiB | basecompute | View Benchmark | |
basecompute/Qwen3-8B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 59.7 tok/s TG128 @ 1 ctx | 1,762 tok/s PP512 | 6,181.0 MiB | isu | View Benchmark | |
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Tesla T10/Tesla T10/Tesla T10/Tesla T10 CUDA | 124.7 tok/s TG128 | 3,087 tok/s PP512 | 11.1 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppBF16 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 45.6 tok/s TG128 | 1,168 tok/s PP512 | 15,747.6 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 67.4 tok/s TG128 | 2,253 tok/s PP512 | 10,441.2 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 77.6 tok/s TG128 | 2,669 tok/s PP512 | 8,419.6 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 88.6 tok/s TG128 | 2,497 tok/s PP512 | 7,262.3 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 95.6 tok/s TG128 | 2,430 tok/s PP512 | 6,532.9 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 104.9 tok/s TG128 | 2,640 tok/s PP512 | 5,724.6 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 105.7 tok/s TG128 | 2,629 tok/s PP512 | 5,698.8 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ4_1 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 115.6 tok/s TG128 | 2,902 tok/s PP512 | 5,123.3 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 115.3 tok/s TG128 | 2,643 tok/s PP512 | 5,013.0 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 118.5 tok/s TG128 | 2,625 tok/s PP512 | 4,913.7 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 127.3 tok/s TG128 | 2,498 tok/s PP512 | 4,226.1 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 130.2 tok/s TG128 | 2,459 tok/s PP512 | 4,051.9 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 146.2 tok/s TG128 | 2,563 tok/s PP512 | 502.4 MiB | arki05 | View Benchmark | |
Qwen3-8B llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 150.7 tok/s TG128 | 2,528 tok/s PP512 | 332.2 MiB | arki05 | View Benchmark |