RigPrice.

Best hardware for Gemma 4 12B

Light assistant for mid-range GPUs.

12B parameters · dense · gemma family

Our picks for this model

Cheapest

1× Intel Arc B580

$320 on eBay ↗ · ~29.8 tok/s est.

one card, no extra setup

Intel card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.

Easiest

1× NVIDIA GeForce RTX 3060 12GB

$325 on eBay ↗ · ~23.5 tok/s est.

one card, no extra setup

Fastest

1× NVIDIA GeForce RTX 5090

$6,000 on eBay ↗ · ~117 tok/s est.

one card, no extra setup

Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.

Q4_K_M

Minimum VRAM: 10.05 GB (file 7.66 GB + context/runtime headroom at 8k ctx; see methodology).

Cheapest fitting cloud offer: Vast.ai $0.056/hr (rtx-3060-12gb class, on-demand) · cloud prices checked 2026-09-23.

ConfigScreened priceTotal VRAMEst. wattsEst. tokens/secBreakeven vs cloudVerdict
1× Intel Arc B580 our pick · Cheapest
Intel card: more setup than NVIDIA for local AI
$32012 GB265 W (well below this table's average)~29.8213 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 3060 12GB our pick · Easiest$32512 GB245 W (well below this table's average)~23.5 (well below this table's average)129 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5090 our pick · Fastest$6,00032 GB650 W~117 (well above this table's average)neverCloud always cheaper at this power price
1× Intel Arc A770 16GB
Intel card: more setup than NVIDIA for local AI
$325
+ shipping unknown
16 GB300 W (well below this table's average)~36.5348 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 3080 10GB
fits with a small context window (about 4k)
$398
+ shipping unknown
10 GB395 W~49.6neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 3080 12GB$500
+ shipping unknown
12 GB425 W~59.6neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 5060 Ti 16GB$70016 GB255 W (well below this table's average)~29.2 (well below this table's average)393 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 4070 Ti SUPER$89016 GB360 W~43.9neverCloud always cheaper at this power price
1× NVIDIA RTX A4000$98016 GB215 W (well below this table's average)~29.2 (well below this table's average)390 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 4080$1,08016 GB395 W~46.8neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 4080 SUPER$1,12016 GB395 W~48.1neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 5070 Ti$1,18016 GB375 W~58.5neverCloud always cheaper at this power price
1× AMD Radeon RX 7900 XTX
AMD card: more setup than NVIDIA for local AI
$1,29924 GB430 W~62.7neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 5080$1,50016 GB435 W~62.7neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 3090$1,53024 GB425 W~61.1neverCloud always cheaper at this power price
1× NVIDIA RTX A5000$2,35024 GB305 W~50.13,044 moRenting wins at 40 h/mo
2× NVIDIA GeForce RTX 3090$3,07048 GB775 W (well above this table's average)~61.1neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 4090$3,20024 GB525 W~65.8neverCloud always cheaper at this power price
1× NVIDIA RTX A6000$4,49948 GB375 W~50.1neverCloud always cheaper at this power price
4× NVIDIA GeForce RTX 3090$6,22096 GB1475 W (well above this table's average)~61.1neverCloud always cheaper at this power price
2× NVIDIA GeForce RTX 4090$6,40048 GB975 W (well above this table's average)~65.8neverCloud always cheaper at this power price
2× NVIDIA RTX A6000$8,99996 GB675 W~50.1neverCloud always cheaper at this power price
1× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$37424 GB325 W~22.7 (well below this table's average)765 moRenting wins at 40 h/mo
2× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$75848 GB575 W~22.7 (well below this table's average)neverCloud always cheaper at this power price
1× NVIDIA TITAN RTX
Turing-era card: much slower than an RTX 3090 for similar money
$1,00024 GB355 W~43.9neverCloud always cheaper at this power price
4× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$1,53596 GB1075 W (well above this table's average)~22.7 (well below this table's average)neverCloud always cheaper at this power price

Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.

  • Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
  • Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
  • Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
  • Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
  • Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
  • Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.

Run your own numbers for Gemma 4 12B (Q4_K_M) in the buy-vs-rent calculator →

Q8_0

Minimum VRAM: 15.31 GB (file 12.67 GB + context/runtime headroom at 8k ctx; see methodology).

Cheapest fitting cloud offer: Vast.ai $0.102/hr (rtx-4060-ti-16gb class, on-demand) · cloud prices checked 2026-09-28.

ConfigScreened priceTotal VRAMEst. wattsEst. tokens/secBreakeven vs cloudVerdict
1× Intel Arc A770 16GB
Intel card: more setup than NVIDIA for local AI
$325
+ shipping unknown
16 GB300 W (well below this table's average)~22.149 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5060 Ti 16GB$70016 GB255 W (well below this table's average)~17.7 (well below this table's average)94 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 4070 Ti SUPER$89016 GB360 W~26.5116 moRenting wins at 40 h/mo
1× NVIDIA RTX A4000$98016 GB215 W (well below this table's average)~17.7 (well below this table's average)122 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 4080$1,08016 GB395 W~28.3159 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 4080 SUPER$1,12016 GB395 W~29.1214 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5070 Ti$1,18016 GB375 W~35.4233 moRenting wins at 40 h/mo
1× AMD Radeon RX 7900 XTX
AMD card: more setup than NVIDIA for local AI
$1,29924 GB430 W~37.9413 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5080$1,50016 GB435 W~37.9293 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 3090$1,53024 GB425 W~36.9334 moRenting wins at 40 h/mo
1× NVIDIA RTX A5000$2,35024 GB305 W (well below this table's average)~30.3385 moRenting wins at 40 h/mo
2× NVIDIA GeForce RTX 3090$3,07048 GB775 W (well above this table's average)~36.9neverCloud always cheaper at this power price
1× NVIDIA GeForce RTX 4090$3,20024 GB525 W~39.81,510 moRenting wins at 40 h/mo
1× NVIDIA RTX A6000$4,49948 GB375 W~30.3740 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5090$6,00032 GB650 W~70.7 (well above this table's average)neverCloud always cheaper at this power price
4× NVIDIA GeForce RTX 3090$6,22096 GB1475 W (well above this table's average)~36.9neverCloud always cheaper at this power price
2× NVIDIA GeForce RTX 4090$6,40048 GB975 W (well above this table's average)~39.8neverCloud always cheaper at this power price
2× NVIDIA RTX A6000$8,99996 GB675 W~30.3neverCloud always cheaper at this power price
1× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$37424 GB325 W~13.7 (well below this table's average)54 moRenting wins at 40 h/mo
2× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$75848 GB575 W~13.7 (well below this table's average)556 moRenting wins at 40 h/mo
1× NVIDIA TITAN RTX
Turing-era card: much slower than an RTX 3090 for similar money
$1,00024 GB355 W~26.5145 moRenting wins at 40 h/mo
4× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$1,53596 GB1075 W (well above this table's average)~13.7 (well below this table's average)neverCloud always cheaper at this power price

Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.

  • Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
  • Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
  • Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
  • Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
  • Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
  • Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.

Run your own numbers for Gemma 4 12B (Q8_0) in the buy-vs-rent calculator →