16 community-reported setups from OWterminal · Reports 2026-09-16–2026-09-17. Saved 2026-09-21. Speeds depend on the exact build and test settings; DVIDIA has not reproduced these reports.
16 reported setups
Every result has a source.
How rankings & scores work
Rank orders the reported output speeds in these results. Speed index = reported tokens/sec ÷ fastest visible reported tokens/sec × 100. The fastest visible setup scores 100; equal speeds share a rank. Filtering recalculates both. Reports use different models, prompts, context lengths and builds, so compare the source conditions before deciding.
Qwen3.5 2B
Q4_K_M GGUF
#1
RTX 5090 32GBClaimed
Reported tokens/sec
351
Speed index
100 / 100
Runtime
Not reported
Reported peak memory
Not reported
Gemma-4 26B-A4B-it
Q4_K_M GGUFabliterated
#2
RTX 5090 32GBClaimed
Reported tokens/sec
173
Speed index
49.3 / 100
Runtime
Not reported
Reported peak memory
Not reported
Qwen3.8-Flash-Next
Build not specified
#3
M5 MaxClaimed
Reported tokens/sec
126.5
Speed index
36 / 100
peak · 61 @100k · 50 @200k
Runtime
MTPLX V2.11.3
Reported peak memory
Not reported
Qwen3.6-35B-A3B
NVFP4abliterated
#4
2× DGX SparkClaimed
Reported tokens/sec
94.4
Speed index
26.9 / 100
Runtime
Not reported
Reported peak memory
Not reported
Qwen3.6-35B-A3B
4-bit MLX
#5
M4 ProClaimed
Reported tokens/sec
85.5
Speed index
24.4 / 100
Runtime
rapid-mlx
Reported peak memory
21 GB RAM
Qwen3.5-4B
4-bit MLX
#6
M4 ProClaimed
Reported tokens/sec
82.8
Speed index
23.6 / 100
Runtime
rapid-mlx
Reported peak memory
8 GB RAM
E
Empero Qwen3.8-35B-A3B
Q4_K_M GGUF
#7
RTX 3060 12GBClaimed
Reported tokens/sec
≈ 50
Speed index
14.2 / 100
Runtime
llama.cpp
Reported peak memory
Not reported
Qwen3.5-9B
4-bit MLX
#8
M4 ProClaimed
Reported tokens/sec
49.3
Speed index
14 / 100
Runtime
rapid-mlx
Reported peak memory
Not reported
Qwen3-8B
4-bit MLX
#9
M4 ProClaimed
Reported tokens/sec
48.3
Speed index
13.8 / 100
Runtime
rapid-mlx
Reported peak memory
Not reported
Qwen3.8-Flash-Next
FP8
#10
2× DGX SparkClaimed
Reported tokens/sec
≈ 45
Speed index
12.8 / 100
Sustained with MTP, as reported
Runtime
Not reported
Reported peak memory
Not reported
Qwen3.8-27B
UD-Q4_K_XL GGUF
#11
RTX 4090 24GBClaimed
Reported tokens/sec
40.7
Speed index
11.6 / 100
60.1 MTP @130k
Runtime
llama.cpp
Reported peak memory
23.7 GB VRAM
O
Ornith
Build not specified
#12
GTX 1660 SUPER 6GBClaimed
Reported tokens/sec
≈ 40
Speed index
11.4 / 100
40–46 range
Runtime
Not reported
Reported peak memory
Not reported
Qwen3.8-27B
GSQ-RCO IQ3_XXS-mtp
#13
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
23
Speed index
6.6 / 100
Runtime
llama.cpp
Reported peak memory
Not reported
Qwen3.8-27B
UD-Q2_K_XL
#14
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
19
Speed index
5.4 / 100
Runtime
llama.cpp
Reported peak memory
Not reported
N
Nex-N2.5-mini
Q4_K_M GGUF
#15
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
14
Speed index
4 / 100
Runtime
llama.cpp
Reported peak memory
Not reported
Qwen3.8-27B TurboFCFusion
IQ2_M GGUFuncensored
#16
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
10
Speed index
2.8 / 100
Runtime
llama.cpp
Reported peak memory
Not reported
Hardware selection finds reports for that machine; it does not establish that a model fits your exact configuration. Missing memory or runtime details remain visible in each result.
02 / Put it to work
Choose your next step.
Run locally
Start on your own machine.
Choose a runtime, download a compatible build and try a small task. Keep the model license and its original source close.
Bring an open-weight model endpoint through the DVIDIA CLI. Local provider preparation and validation are implemented; public distribution and verified network enrollment are next.
The planned service links request usage to a verified provider address, then shares collected customer revenue. Network-token incentives will be tracked separately; rates, eligibility and rewards are not live.