Models
Loading community benchmarks…
Loading community benchmarks…
Submitted benchmarks for Qwen3.6 35B A3B. Browse individual runs and their measurement settings, or open the report JSON for full benchmark details.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions. Model verification identifies published artifact bytes; community benchmark execution is not remotely attested.
| Date |
|---|
| Report |
|---|
basecompute/Qwen3.6-35B-A3B BaseRTQ8· default-q8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Max Metal | 119.6 tok/s TG128 | 1,540 tok/s PP512 | 3,727.4 MiB | lukas | View Benchmark | |
basecompute/Qwen3.6-35B-A3B BaseRTQ4· default-q4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro Metal | 113.2 tok/s TG128 | 1,275 tok/s PP512 | 2,765.7 MiB | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 67.5 tok/s TG128 | 1,737 tok/s PP512 | 16,676.6 MiB | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cppQ4_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 61.5 tok/s TG128 | 1,673 tok/s PP512 | 21,935.8 MiB | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cppQ5_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M5 Pro BLAS + Metal | 61.5 tok/s TG128 | 1,533 tok/s PP512 | 25,973.1 MiB | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT ROCm | 100.3 tok/s TG128 | 2,501 tok/s PP512 | 17,927.9 MiB | arki05 | View Benchmark | |
Qwen3.6-35B-A3B llama.cppQ3_K_M The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 124.1 tok/s TG128 | 2,597 tok/s PP512 | 16,221.2 MiB | arki05 | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | NVIDIA GB10 CUDA | 84.3 tok/s TG128 @ 1 ctx | 2,028 tok/s PP512 | 16,676.7 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 72.2 tok/s TG128 @ 1 ctx | 232 tok/s PP512 | 59,584.5 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-cuda-q4mix BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 72.8 tok/s TG128 @ 1 ctx | 2,279 tok/s PP512 | 14,955.8 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 70.8 tok/s TG128 @ 1 ctx | 247 tok/s PP512 | 59,566.5 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 54.4 tok/s TG128 @ 1 ctx | 2,188 tok/s PP512 | 28,757.9 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-cuda-q8 BaseRTQ4 Nothing in this report ties it to a published model file. Reports from CLI 0.1.0 record only a model name and a file hash, so the run is grouped by the name it reported until it is reconciled. | NVIDIA GB10 CUDA | 55.4 tok/s TG128 @ 1 ctx | 2,181 tok/s PP512 | 28,760.0 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M1 Max Metal | 94.2 tok/s TG128 @ 1 ctx | 278 tok/s PP512 | 2,749.9 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 109.5 tok/s TG128 @ 1 ctx | 913 tok/s PP512 | 3,677.9 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q8 BaseRTQ8 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 110.8 tok/s TG128 @ 1 ctx | 910 tok/s PP512 | 3,678.1 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M3 Ultra Metal | 133.0 tok/s TG128 @ 1 ctx | 922 tok/s PP512 | 2,747.0 MiB | basecompute | View Benchmark | |
Qwen3.6-35B-A3B-Q4 BaseRTQ4 The report's model SHA-256 matches a file published on Hugging Face, so the exact model bytes are known. This does not attest benchmark execution. | Apple M4 Max Metal | 139.7 tok/s TG128 @ 1 ctx | 886 tok/s PP512 | 2,747.8 MiB | basecompute | View Benchmark |