Skip to content
selfhostr

run a model

Running Gemma 4 26B-A4B

A mixture of experts: it loads like a 25-billion model and answers faster than one.

25.2 billion parametersweights

memory needed

17.1 GB

the smallest hardware that runs it properly

Ne se vend plus neuve. Seules subsistent des offres de revente à 1 900–2 600 €, ni stables ni vérifiables : on n'affiche pas de prix.

price not checked

the cheapest one you can actually buy today

runs it, but only just

€1,169 · checked 2026-08-02product page

the alternatives

price not checked
€1,839 · checked 2026-08-02product page
€2,559 · checked 2026-08-02product page

on the processor alone: slow, and it works

price not checked
€4,099 · checked 2026-08-02product page
€7,659 · checked 2026-08-02product page

runs it, but only just

on the processor alone: slow, and it works

price not checked

Only just means it works, until the day you lengthen the context or open something else. Which is the day you conclude the tool lied to you.

not enough memory

if you change the format

4-bit weightswhat almost everyone runs, and what the hardware below is judged against
17.1 GB
8-bit weightsvisibly better on long or precise answers
30.7 GB
16-bit weightsthe model as it was trained, and rarely worth it at home
61 GB

Everything above assumes 4-bit weights and a 8k token context. Going up in precision costs memory; shortening the context is the other lever, and the cheaper one when something almost fits. Model details checked on 2026-08-02.

how to actually run it

what this page will not tell you

How many tokens per second you will get. That is a benchmark — it depends on memory bandwidth, on the engine, on the context and on how long the answers are. This site does not invent benchmarks, and everyone who quotes one for your exact setup is guessing.

builder

Price a whole machine around itA card is not a server. The builder adds the machine, the drives, the electricity and the backup.