DVIDIA × OWterminal

Open weights.
A place to start.

Find a model. Understand the hardware. Make it useful.

DVIDIA is where skills become useful work. Our sister site, OWterminal, helps you research the models and machines behind them.

01 / Explore the possibilities

What do you want to run?

Full comparison desk

16 community-reported setups from OWterminal · Reports 2026-09-162026-09-17. Saved 2026-09-21. Speeds depend on the exact build and test settings; DVIDIA has not reproduced these reports.

16 reported setups

Every result has a source.
How rankings & scores work

Rank orders the reported output speeds in these results. Speed index = reported tokens/sec ÷ fastest visible reported tokens/sec × 100. The fastest visible setup scores 100; equal speeds share a rank. Filtering recalculates both. Reports use different models, prompts, context lengths and builds, so compare the source conditions before deciding.

Qwen3.5 2B

Q4_K_M GGUF

#1
RTX 5090 32GBClaimed
Reported tokens/sec
351
Speed index
100 / 100
Runtime
Not reported
Reported peak memory
Not reported

Gemma-4 26B-A4B-it

Q4_K_M GGUFabliterated

#2
RTX 5090 32GBClaimed
Reported tokens/sec
173
Speed index
49.3 / 100
Runtime
Not reported
Reported peak memory
Not reported

Qwen3.8-Flash-Next

Build not specified

#3
M5 MaxClaimed
Reported tokens/sec
126.5
Speed index
36 / 100

peak · 61 @100k · 50 @200k

Runtime
MTPLX V2.11.3
Reported peak memory
Not reported

Qwen3.6-35B-A3B

NVFP4abliterated

#4
2× DGX SparkClaimed
Reported tokens/sec
94.4
Speed index
26.9 / 100
Runtime
Not reported
Reported peak memory
Not reported

Qwen3.6-35B-A3B

4-bit MLX

#5
M4 ProClaimed
Reported tokens/sec
85.5
Speed index
24.4 / 100
Runtime
rapid-mlx
Reported peak memory
21 GB RAM

Qwen3.5-4B

4-bit MLX

#6
M4 ProClaimed
Reported tokens/sec
82.8
Speed index
23.6 / 100
Runtime
rapid-mlx
Reported peak memory
8 GB RAM

Empero Qwen3.8-35B-A3B

Q4_K_M GGUF

#7
RTX 3060 12GBClaimed
Reported tokens/sec
≈ 50
Speed index
14.2 / 100
Runtime
llama.cpp
Reported peak memory
Not reported

Qwen3.5-9B

4-bit MLX

#8
M4 ProClaimed
Reported tokens/sec
49.3
Speed index
14 / 100
Runtime
rapid-mlx
Reported peak memory
Not reported

Qwen3-8B

4-bit MLX

#9
M4 ProClaimed
Reported tokens/sec
48.3
Speed index
13.8 / 100
Runtime
rapid-mlx
Reported peak memory
Not reported

Qwen3.8-Flash-Next

FP8

#10
2× DGX SparkClaimed
Reported tokens/sec
≈ 45
Speed index
12.8 / 100

Sustained with MTP, as reported

Runtime
Not reported
Reported peak memory
Not reported

Qwen3.8-27B

UD-Q4_K_XL GGUF

#11
RTX 4090 24GBClaimed
Reported tokens/sec
40.7
Speed index
11.6 / 100

60.1 MTP @130k

Runtime
llama.cpp
Reported peak memory
23.7 GB VRAM

Ornith

Build not specified

#12
GTX 1660 SUPER 6GBClaimed
Reported tokens/sec
≈ 40
Speed index
11.4 / 100

40–46 range

Runtime
Not reported
Reported peak memory
Not reported

Qwen3.8-27B

GSQ-RCO IQ3_XXS-mtp

#13
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
23
Speed index
6.6 / 100
Runtime
llama.cpp
Reported peak memory
Not reported

Qwen3.8-27B

UD-Q2_K_XL

#14
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
19
Speed index
5.4 / 100
Runtime
llama.cpp
Reported peak memory
Not reported

Nex-N2.5-mini

Q4_K_M GGUF

#15
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
14
Speed index
4 / 100
Runtime
llama.cpp
Reported peak memory
Not reported

Qwen3.8-27B TurboFCFusion

IQ2_M GGUFuncensored

#16
RTX 5060 Ti 16GBClaimed
Reported tokens/sec
10
Speed index
2.8 / 100
Runtime
llama.cpp
Reported peak memory
Not reported

Hardware selection finds reports for that machine; it does not establish that a model fits your exact configuration. Missing memory or runtime details remain visible in each result.

02 / Put it to work

Choose your next step.

Choose hardware

Research before you buy.

OWterminal brings hardware comparisons and sourced market listings together, so you can inspect the evidence and the seller’s current terms.

Compare hardware Browse the hardware market
Build with DVIDIA

Give the model a useful task.

Explore an open skill or record a demonstration. Check its artifact type and runtime requirements before pairing it with a model.

Explore skills Record a demonstration

Coming next / The provider network

Bring your own capacity.

Bring an open-weight model endpoint through the DVIDIA CLI. Local provider preparation and validation are implemented; public distribution and verified network enrollment are next.

The planned service links request usage to a verified provider address, then shares collected customer revenue. Network-token incentives will be tracked separately; rates, eligibility and rewards are not live.

  • Open model licenses stay visible
  • Pay for delivered service
  • Contributors share in collected revenue