Best hardware for Gemma 4 12B
Light assistant for mid-range GPUs.
12B parameters · dense · gemma family
Our picks for this model
Cheapest
$320 on eBay ↗ · ~29.8 tok/s est.
one card, no extra setup
Intel card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.
Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.
Q4_K_M
Minimum VRAM: 10.05 GB (file 7.66 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: Vast.ai $0.056/hr (rtx-3060-12gb class, on-demand) · cloud prices checked 2026-09-23.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× Intel Arc B580
our pick · Cheapest Intel card: more setup than NVIDIA for local AI
| $320 | 12 GB | 265 W (well below this table's average) | ~29.8 | 213 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3060 12GB our pick · Easiest | $325 | 12 GB | 245 W (well below this table's average) | ~23.5 (well below this table's average) | 129 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 our pick · Fastest | $6,000 | 32 GB | 650 W | ~117 (well above this table's average) | never | Cloud always cheaper at this power price |
| 1× Intel Arc A770 16GB Intel card: more setup than NVIDIA for local AI
| $325 + shipping unknown | 16 GB | 300 W (well below this table's average) | ~36.5 | 348 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3080 10GB
fits with a small context window (about 4k)
| $398 + shipping unknown | 10 GB | 395 W | ~49.6 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 3080 12GB | $500 + shipping unknown | 12 GB | 425 W | ~59.6 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 5060 Ti 16GB | $700 | 16 GB | 255 W (well below this table's average) | ~29.2 (well below this table's average) | 393 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4070 Ti SUPER | $890 | 16 GB | 360 W | ~43.9 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A4000 | $980 | 16 GB | 215 W (well below this table's average) | ~29.2 (well below this table's average) | 390 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 | $1,080 | 16 GB | 395 W | ~46.8 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 4080 SUPER | $1,120 | 16 GB | 395 W | ~48.1 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 5070 Ti | $1,180 | 16 GB | 375 W | ~58.5 | never | Cloud always cheaper at this power price |
| 1× AMD Radeon RX 7900 XTX AMD card: more setup than NVIDIA for local AI
| $1,299 | 24 GB | 430 W | ~62.7 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 5080 | $1,500 | 16 GB | 435 W | ~62.7 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 3090 | $1,530 | 24 GB | 425 W | ~61.1 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A5000 | $2,350 | 24 GB | 305 W | ~50.1 | 3,044 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | $3,070 | 48 GB | 775 W (well above this table's average) | ~61.1 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 4090 | $3,200 | 24 GB | 525 W | ~65.8 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A6000 | $4,499 | 48 GB | 375 W | ~50.1 | never | Cloud always cheaper at this power price |
| 4× NVIDIA GeForce RTX 3090 | $6,220 | 96 GB | 1475 W (well above this table's average) | ~61.1 | never | Cloud always cheaper at this power price |
| 2× NVIDIA GeForce RTX 4090 | $6,400 | 48 GB | 975 W (well above this table's average) | ~65.8 | never | Cloud always cheaper at this power price |
| 2× NVIDIA RTX A6000 | $8,999 | 96 GB | 675 W | ~50.1 | never | Cloud always cheaper at this power price |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $374 | 24 GB | 325 W | ~22.7 (well below this table's average) | 765 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $758 | 48 GB | 575 W | ~22.7 (well below this table's average) | never | Cloud always cheaper at this power price |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money | $1,000 | 24 GB | 355 W | ~43.9 | never | Cloud always cheaper at this power price |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,535 | 96 GB | 1075 W (well above this table's average) | ~22.7 (well below this table's average) | never | Cloud always cheaper at this power price |
Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Gemma 4 12B (Q4_K_M) in the buy-vs-rent calculator →
Q8_0
Minimum VRAM: 15.31 GB (file 12.67 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: Vast.ai $0.102/hr (rtx-4060-ti-16gb class, on-demand) · cloud prices checked 2026-09-28.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× Intel Arc A770 16GB Intel card: more setup than NVIDIA for local AI
| $325 + shipping unknown | 16 GB | 300 W (well below this table's average) | ~22.1 | 49 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5060 Ti 16GB | $700 | 16 GB | 255 W (well below this table's average) | ~17.7 (well below this table's average) | 94 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4070 Ti SUPER | $890 | 16 GB | 360 W | ~26.5 | 116 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A4000 | $980 | 16 GB | 215 W (well below this table's average) | ~17.7 (well below this table's average) | 122 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 | $1,080 | 16 GB | 395 W | ~28.3 | 159 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 SUPER | $1,120 | 16 GB | 395 W | ~29.1 | 214 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5070 Ti | $1,180 | 16 GB | 375 W | ~35.4 | 233 mo | Renting wins at 40 h/mo |
| 1× AMD Radeon RX 7900 XTX AMD card: more setup than NVIDIA for local AI
| $1,299 | 24 GB | 430 W | ~37.9 | 413 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5080 | $1,500 | 16 GB | 435 W | ~37.9 | 293 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 | $1,530 | 24 GB | 425 W | ~36.9 | 334 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A5000 | $2,350 | 24 GB | 305 W (well below this table's average) | ~30.3 | 385 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | $3,070 | 48 GB | 775 W (well above this table's average) | ~36.9 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 4090 | $3,200 | 24 GB | 525 W | ~39.8 | 1,510 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 | $4,499 | 48 GB | 375 W | ~30.3 | 740 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 | $6,000 | 32 GB | 650 W | ~70.7 (well above this table's average) | never | Cloud always cheaper at this power price |
| 4× NVIDIA GeForce RTX 3090 | $6,220 | 96 GB | 1475 W (well above this table's average) | ~36.9 | never | Cloud always cheaper at this power price |
| 2× NVIDIA GeForce RTX 4090 | $6,400 | 48 GB | 975 W (well above this table's average) | ~39.8 | never | Cloud always cheaper at this power price |
| 2× NVIDIA RTX A6000 | $8,999 | 96 GB | 675 W | ~30.3 | never | Cloud always cheaper at this power price |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $374 | 24 GB | 325 W | ~13.7 (well below this table's average) | 54 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $758 | 48 GB | 575 W | ~13.7 (well below this table's average) | 556 mo | Renting wins at 40 h/mo |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money | $1,000 | 24 GB | 355 W | ~26.5 | 145 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,535 | 96 GB | 1075 W (well above this table's average) | ~13.7 (well below this table's average) | never | Cloud always cheaper at this power price |
Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Gemma 4 12B (Q8_0) in the buy-vs-rent calculator →