Best hardware for Qwen3.6 35B-A3B
Fast all-round assistant.
35B parameters · MoE (3B active per token, so it runs faster than its size suggests) · qwen family
Our picks for this model
Cheapest
$1,299 on eBay ↗ · ~251 tok/s est.
one card, no extra setup
Runs with a small context window (about 4k): fine for chat, tight for long documents. Step down a quant for more room.
AMD card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.
Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.
Q4_K_M
Minimum VRAM: 25.41 GB (file 22.29 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-09-29.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× AMD Radeon RX 7900 XTX
our pick · Cheapest AMD card: more setup than NVIDIA for local AI
fits with a small context window (about 4k)
| $1,299 | 24 GB | 430 W | ~251 | 53 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 our pick · Easiest | $4,499 | 48 GB | 375 W (well below this table's average) | ~201 | 116 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 our pick · Fastest | $6,000 | 32 GB | 650 W | ~469 (well above this table's average) | 232 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090
fits with a small context window (about 4k)
| $1,530 | 24 GB | 425 W | ~245 | 44 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A5000
fits with a small context window (about 4k)
| $2,350 | 24 GB | 305 W (well below this table's average) | ~201 | 73 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | $3,070 | 48 GB | 775 W | ~245 | 112 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4090
fits with a small context window (about 4k)
| $3,200 | 24 GB | 525 W | ~264 | 112 mo | Renting wins at 40 h/mo |
| 4× NVIDIA GeForce RTX 3090 | $6,220 | 96 GB | 1475 W (well above this table's average) | ~245 | 513 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | $6,400 | 48 GB | 975 W (well above this table's average) | ~264 | 316 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A6000 | $8,999 | 96 GB | 675 W | ~201 | 281 mo | Renting wins at 40 h/mo |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards
fits with a small context window (about 4k)
| $374 | 24 GB | 325 W (well below this table's average) | ~90.8 (well below this table's average) | 9.8 mo | Buying wins by month 10 |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $758 | 48 GB | 575 W | ~90.8 (well below this table's average) | 24 mo | Buying wins by month 24 |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money
fits with a small context window (about 4k)
| $1,000 | 24 GB | 355 W (well below this table's average) | ~176 | 24 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,535 | 96 GB | 1075 W (well above this table's average) | ~90.8 (well below this table's average) | 75 mo | Renting wins at 40 h/mo |
Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Qwen3.6 35B-A3B (Q4_K_M) in the buy-vs-rent calculator →
Q8_0
Minimum VRAM: 41.71 GB (file 37.81 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-09-29.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 2× NVIDIA GeForce RTX 3090 | $3,070 | 48 GB | 775 W | ~144 | 112 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 | $4,499 | 48 GB | 375 W (well below this table's average) | ~118 | 116 mo | Renting wins at 40 h/mo |
| 4× NVIDIA GeForce RTX 3090 | $6,220 | 96 GB | 1475 W (well above this table's average) | ~144 | 513 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | $6,400 | 48 GB | 975 W | ~156 | 316 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A6000 | $8,999 | 96 GB | 675 W | ~118 | 281 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $758 | 48 GB | 575 W | ~53.6 (well below this table's average) | 24 mo | Buying wins by month 24 |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,535 | 96 GB | 1075 W | ~53.6 (well below this table's average) | 75 mo | Renting wins at 40 h/mo |
Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Qwen3.6 35B-A3B (Q8_0) in the buy-vs-rent calculator →