run a model
Running gpt-oss 20B
OpenAI's open-weight model, in the size meant for one machine.
20 billion parameterslicence : Apache 2.0weights ↗
memory needed
13.6 GB
the smallest hardware that runs it properly
Le moins cher à ce palier : Mac mini M4, 24 Go.
the alternatives
on the processor alone: slow, and it works
on the processor alone: slow, and it works
runs it, but only just
runs it, but only just
Only just means it works, until the day you lengthen the context or open something else. Which is the day you conclude the tool lied to you.
not enough memory
- NVIDIA GeForce RTX 3060 8 GB · 8 GB usable
- NVIDIA GeForce RTX 3070 · 8 GB usable
- NVIDIA GeForce RTX 5060 8 GB (MSI Shadow 2X OC) · 8 GB usable
- No graphics card, 16 GB of RAM · 9.6 GB usable
- NVIDIA GeForce RTX 3080 · 10 GB usable
- NVIDIA GeForce RTX 3060 12 GB · 12 GB usable
- NVIDIA GeForce RTX 4070 · 12 GB usable
- NVIDIA GeForce RTX 5070 12 GB (MSI Shadow 2X OC) · 12 GB usable
- Apple Silicon, 16 GB unified · 12 GB usable
if you change the format
- 4-bit weightswhat almost everyone runs, and what the hardware below is judged against
- 13.6 GB
- 8-bit weightsvisibly better on long or precise answers
- 24.4 GB
- 16-bit weightsthe model as it was trained, and rarely worth it at home
- 48.4 GB
Everything above assumes 4-bit weights and a 8k token context. Going up in precision costs memory; shortening the context is the other lever, and the cheaper one when something almost fits. Model details checked on 2026-08-02.
how to actually run it
what this page will not tell you
How many tokens per second you will get. That is a benchmark — it depends on memory bandwidth, on the engine, on the context and on how long the answers are. This site does not invent benchmarks, and everyone who quotes one for your exact setup is guessing.
builder
Price a whole machine around itA card is not a server. The builder adds the machine, the drives, the electricity and the backup.