Community submitted benchmarks for Local AI.

Every result comes from a signed report produced on community hardware. Browse by model, quantisation and chip.

Install & start · macOS / Linux

curl -LsSf https://computearena.ai/install.sh | sh -s launch

Requires curl, tar and Python 3. Review and confirm the installation; the interactive menu opens in this terminal. No sudo required.

Latest runs

  1. Nemotron 3 Nano 30B A3BReported as nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 · nemotron_h_moebasecomputeAustralia
    BaseRTQ8Apple M4 Max
    94.9
    decode
    1,638
    prefill
  2. Qwen3.6 27BReported as Qwen/Qwen3.6-27B · qwen35basecomputeAustralia
    BaseRTQ8Apple M4 Max
    17.5
    decode
    222
    prefill
  3. Qwen3.8 27BReported as Qwen/Qwen3.8-27B · qwen35basecomputeAustralia
    BaseRTQ4Apple M4 Max
    17.7
    decode
    221
    prefill
  4. Qwen3.6 35B A3BReported as Qwen/Qwen3.6-35B-A3B · qwen35moebasecomputeAustralia
    BaseRTQ8Apple M4 Max
    104.6
    decode
    878
    prefill
  5. Qwen3.8 27BReported as Qwen/Qwen3.8-27B · qwen35basecomputeAustralia
    BaseRTQ4Apple M4 Max
    31.5
    decode
    221
    prefill
  6. Qwen3 30B A3B Thinking 2507Reported as Qwen/Qwen3-30B-A3B-Thinking-2507 · qwen3_moebasecomputeAustralia
    BaseRTQ8Apple M4 Max
    95.9
    decode
    1,795
    prefill
How it works
01

Install ComputeArena

Get the ComputeArena CLI for macOS or Linux, then choose BaseRT or llama.cpp as your runtime.

02

Run offline

No account is required. Signed reports remain on your machine until you submit them.

03

Submit when ready

Sign in, review the data in your saved reports, and submit the benchmarks you want to share publicly.

Leaderboard

Model × quantisation × chip

Headline rankings compare compatible PP512 prefill and TG128 decode workloads. The measured workload is shown beside every result.

Model
Runtime
Quantisation
Chip
Prefill size
Rank by
380 matching configurations
#Model / quantRuntimeChip / backendUserUpdatedShare
1Tinystories Lay8 HS512 HD8 33MReported as ivnle/tinystories-lay8-hs512-hd8-33M · llama · default-q4
Q4
BaseRTApple M5 Pro
Metal
2,068.0
TG128
190,107
PP512
arki05Germany
2Tinystories Lay8 HS512 HD8 33MReported as RichardErkhov/ivnle_-_tinystories-lay8-hs512-hd8-33M-gguf · llama
Q4_0
llama.cppApple M5 Pro
BLAS + Metal
1,443.3
TG128
124,181
PP512
arki05Germany
3Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4
Q4
BaseRTApple M5 Max
Metal
708.3
TG128
34,136
PP512
basecomputeAustralia
4Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q4
BaseRTApple M4 Max
Metal
583.3
TG128 @ 1 ctx
9,998
PP512
basecomputeAustralia
5Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8
Q8
BaseRTApple M5 Max
Metal
549.5
TG128
33,088
PP512
basecomputeAustralia
6Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4
Q4
BaseRTApple M5 Pro
Metal
516.8
TG128
20,778
PP512
arki05Germany
7Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppAMD Radeon RX 7900 XT (RADV NAVI31)
Vulkan
508.8
TG128
26,137
PP512
arki05Germany
8Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q4
BaseRTApple M3 Ultra
Metal
501.7
TG128 @ 1 ctx
12,830
PP512
basecomputeAustralia
9Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppAMD Radeon RX 7900 XT (RADV NAVI31)
Vulkan
479.4
TG128
26,346
PP512
arki05Germany
10Qwen3 0.6BReported as Qwen3-0.6B-cuda-q4, Qwen3-0.6B-cuda-q4mix, Qwen3-0.6B-cuda-q8, Qwen/Qwen3-0.6B · qwen
Q4
BaseRTNVIDIA GB10
CUDA
458.8
TG128 @ 1 ctx
52,750
PP512
basecomputeAustralia
11Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q8
BaseRTApple M4 Max
Metal
455.8
TG128 @ 1 ctx
9,749
PP512
basecomputeAustralia
12Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q8
BaseRTApple M3 Ultra
Metal
440.2
TG128 @ 1 ctx
12,494
PP512
basecomputeAustralia
13Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppApple M5 Max
BLAS + Metal
439.1
TG128
24,365
PP512
basecomputeAustralia
14Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4
Q4
BaseRTApple M5 Max
Metal
415.5
TG128
23,019
PP512
lukasAustralia
15Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTApple M3 Ultra
Metal
410.1
TG128 @ 1 ctx
9,066
PP512
basecomputeAustralia
16Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppApple M5 Max
BLAS + Metal
379.0
TG128
24,983
PP512
basecomputeAustralia
17Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTApple M4 Max
Metal
371.7
TG128 @ 1 ctx
6,058
PP512
basecomputeAustralia
18Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppAMD Radeon RX 7900 XT
ROCm
368.6
TG128
23,972
PP512
arki05Germany
19Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q4
BaseRTApple M4 Max
Metal
361.2
TG128 @ 1 ctx
7,448
PP512
basecomputeAustralia
20Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q4_K_M
llama.cppApple M5 Pro
BLAS + Metal
358.3
TG128
14,552
PP512
arki05Germany
21Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8
Q8
BaseRTApple M5 Pro
Metal
350.2
TG128
20,537
PP512
arki05Germany
22Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q4
BaseRTApple M3 Ultra
Metal
328.0
TG128 @ 1 ctx
9,903
PP512
basecomputeAustralia
23Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen
Q4
BaseRTApple M4 Max
Metal
318.4
TG128 @ 1 ctx
4,146
PP512
basecomputeAustralia
24Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppAMD Radeon RX 7900 XT
ROCm
313.7
TG128
24,691
PP512
arki05Germany
25Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen
Q4
BaseRTApple M3 Ultra
Metal
307.4
TG128 @ 1 ctx
5,420
PP512
basecomputeAustralia
26Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35
Q4
BaseRTApple M3 Ultra
Metal
306.7
TG128 @ 1 ctx
1,759
PP512
basecomputeAustralia
27Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35
Q4
BaseRTApple M4 Max
Metal
301.7
TG128 @ 1 ctx
1,718
PP512
basecomputeAustralia
28Gemma 3 1B ITReported as gemma-3-1b-it-cuda-q4mix, gemma-3-1b-it-cuda-q8, google/gemma-3-1b-it · gemma3
Q4
BaseRTNVIDIA GB10
CUDA
295.9
TG128 @ 1 ctx
33,841
PP512
basecomputeAustralia
29Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35
Q4
BaseRTApple M4 Max
Metal
295.1
TG128 @ 1 ctx
1,713
PP512
basecomputeAustralia
30Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35
Q4
BaseRTApple M3 Ultra
Metal
293.8
TG128 @ 1 ctx
1,756
PP512
basecomputeAustralia
31Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4
Q4
BaseRTApple M5 Pro
Metal
293.1
TG128
15,066
PP512
arki05Germany
32Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen
Q8
BaseRTApple M4 Pro
Metal
284.7
TG128 @ 1 ctx
4,787
PP512
basecomputeAustralia
33Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3
Q8_0
llama.cppApple M5 Pro
BLAS + Metal
283.0
TG128
14,942
PP512
arki05Germany
34Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q8
BaseRTApple M4 Max
Metal
281.0
TG128 @ 1 ctx
7,527
PP512
basecomputeAustralia
35Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4
Q4
BaseRTApple M1 Pro
Metal
268.8
TG128 @ 1 ctx
2,847
PP512
isuAustralia
36Llama 3.2 1B InstructReported as Llama-3.2-1B-Instruct-cuda-q4mix, Llama-3.2-1B-Instruct-cuda-q8, meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTNVIDIA GB10
CUDA
261.3
TG128 @ 1 ctx
36,647
PP512
basecomputeAustralia
37Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3
Q8
BaseRTApple M3 Ultra
Metal
257.8
TG128 @ 1 ctx
9,953
PP512
basecomputeAustralia
38Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama
Q4
BaseRTApple M1 Max
Metal
242.1
TG128 @ 1 ctx
3,372
PP512
basecomputeAustralia
39Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen
Q8
BaseRTApple M3 Ultra
Metal
238.6
TG128 @ 1 ctx
5,698
PP512
basecomputeAustralia
40Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen · default-q4
Q4
BaseRTApple M5 Pro
Metal
235.4
TG128
8,040
PP512
arki05Germany