Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Qwen3 4B Instruct 2507. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 89.4 tok/s TG128 | 3,359 tok/s PP512 | 3,909.1 MiB | arki05 | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 64.3 tok/s TG128 | 3,398 tok/s PP512 | 6,047.5 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 94.4 tok/s TG128 | 2,568 tok/s PP512 | 4,860.1 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 92.6 tok/s TG128 | 2,566 tok/s PP512 | 4,902.8 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 60.8 tok/s TG128 | 2,686 tok/s PP512 | 6,558.2 MiB | arki05 | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 90.4 tok/s TG128 | 3,331 tok/s PP512 | 3,908.9 MiB | arki05 | View Benchmark | |
basecompute/Qwen3-4B-Instruct-2507 BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 62.8 tok/s TG128 | 3,405 tok/s PP512 | 6,047.2 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 147.8 tok/s TG128 | 5,364 tok/s PP512 | 4,211.3 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 147.1 tok/s TG128 | 5,337 tok/s PP512 | 4,254.3 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 110.8 tok/s TG128 | 5,581 tok/s PP512 | 5,908.0 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 185.1 tok/s TG128 | 4,783 tok/s PP512 | 424.9 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 180.3 tok/s TG128 | 4,793 tok/s PP512 | 425.5 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507 llama.cppQ8_0 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 134.3 tok/s TG128 | 4,824 tok/s PP512 | 4,200.9 MiB | arki05 | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 73.7 tok/s TG128 @ 1 ctx | 11,246 tok/s PP512 | 3,379.4 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 54.8 tok/s TG128 @ 1 ctx | 10,970 tok/s PP512 | 4,520.0 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 121.9 tok/s TG128 @ 1 ctx | 1,711 tok/s PP512 | 3,897.2 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 104.3 tok/s TG128 @ 1 ctx | 1,723 tok/s PP512 | 6,035.8 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 72.1 tok/s TG128 @ 1 ctx | 925 tok/s PP512 | 6,038.4 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 91.3 tok/s TG128 @ 1 ctx | 922 tok/s PP512 | 3,898.2 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 56.2 tok/s TG128 @ 1 ctx | 731 tok/s PP512 | 6,037.3 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 118.3 tok/s TG128 @ 1 ctx | 2,557 tok/s PP512 | 6,032.5 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 153.7 tok/s TG128 @ 1 ctx | 2,529 tok/s PP512 | 3,894.0 MiB | basecompute | View Benchmark | |
Qwen3-4B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Pro Metal | 77.9 tok/s TG128 @ 1 ctx | 727 tok/s PP512 | 3,898.9 MiB | basecompute | View Benchmark |