Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Qwen3 30B A3B Instruct 2507. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Contributor |
|---|
| Date |
|---|
| Report |
|---|
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 139.8 tok/s TG128 @ 1 ctx | 1,789 tok/s PP512 | 1,470.0 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 137.1 tok/s TG128 @ 1 ctx | 2,287 tok/s PP512 | 1,466.5 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 87.7 tok/s TG128 @ 1 ctx | 901 tok/s PP512 | 1,471.3 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 61.2 tok/s TG128 @ 1 ctx | 7,253 tok/s PP512 | 30,962.8 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 49.4 tok/s TG128 @ 1 ctx | 7,315 tok/s PP512 | 21,317.5 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 99.0 tok/s TG128 @ 1 ctx | 7,439 tok/s PP512 | 17,897.6 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 80.7 tok/s TG128 @ 1 ctx | 4,715 tok/s PP512 | 13,752.2 MiB | basecompute | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 181.3 tok/s TG128 | 2,827 tok/s PP512 | 17,816.1 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 186.7 tok/s TG128 | 2,868 tok/s PP512 | 16,990.0 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 184.5 tok/s TG128 | 2,474 tok/s PP512 | 13,311.2 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 198.1 tok/s TG128 | 2,837 tok/s PP512 | 11,357.9 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 128.6 tok/s TG128 | 2,661 tok/s PP512 | 19,571.4 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 128.5 tok/s TG128 | 2,820 tok/s PP512 | 18,741.8 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 128.8 tok/s TG128 | 2,596 tok/s PP512 | 15,066.6 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 137.4 tok/s TG128 | 2,386 tok/s PP512 | 13,115.8 MiB | arki05 | View Benchmark | |
basecompute/Qwen3-30B-A3B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 104.0 tok/s TG128 | 3,680 tok/s PP512 | 1,481.7 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ6_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 78.9 tok/s TG128 | 1,646 tok/s PP512 | 25,646.9 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 84.7 tok/s TG128 | 1,644 tok/s PP512 | 22,433.6 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 95.2 tok/s TG128 | 1,845 tok/s PP512 | 18,586.3 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 95.0 tok/s TG128 | 1,831 tok/s PP512 | 19,412.5 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 97.3 tok/s TG128 | 1,708 tok/s PP512 | 14,908.4 MiB | arki05 | View Benchmark | |
Qwen3-30B-A3B-Instruct-2507 llama.cppQ2_K The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 102.7 tok/s TG128 | 1,893 tok/s PP512 | 12,958.3 MiB | arki05 | View Benchmark | |
basecompute/Qwen3-30B-A3B-Instruct-2507 BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 92.2 tok/s TG128 | 3,305 tok/s PP512 | 1,482.0 MiB | arki05 | View Benchmark |