run a model
What you can run on an NVIDIA GeForce RTX 3080
10 GB usable out of 10.
fits comfortably
Qwen3.5 4BRuns on a laptop with no graphics card. Rewriting, summarising, sorting — not reasoning.2.7 GB Llama 3.1 8BThe one every tutorial assumes. Old enough that every problem you hit has already been answered somewhere.5.4 GB Qwen3.5 9BThe everyday assistant, and the smallest one that answers usefully about a document you paste in.6.1 GB fits, but only just
Gemma 4 12BGoogle's middle size: good at ordinary language, and it reads images.8.2 GB Phi-4 14BPunches above its size on reasoning and maths for how little memory it takes.9.5 GB does not fit
gpt-oss 20BOpenAI's open-weight model, in the size meant for one machine.13.6 GB Gemma 4 26B-A4BA mixture of experts: it loads like a 25-billion model and answers faster than one.17.1 GB Qwen3.6 27BThe size where a machine at home starts to replace a paid subscription for most everyday work.18.4 GB Qwen3-Coder 30B-A3BBuilt for code, and the one to run if that is the whole reason you are here.20.4 GB Gemma 4 31BThe largest Gemma that still fits on one consumer card.20.9 GB Qwen3.6 35B-A3BOne notch above 27B, and the point where 24 GB stops being enough.23.8 GB Llama 3.1 70BMeta's large one. Needs hardware bought for this and nothing else.47.6 GB gpt-oss 120BFits on a single 80 GB card, or a Mac with a lot of unified memory.81.6 GB Figured at 4-bit weights — what almost everyone runs — with a 8k token context. Both change the numbers: shorter context is the first lever when something almost fits. Model details checked on 2026-08-02.
what generation costs to run
- 1 h a day
- €29 a year · 117 kWh
- 2 h a day
- €58 a year · 234 kWh
- 8 h a day
- €234 a year · 934 kWh
At 320 W under load and 0.25 € per kWh — the European average, since this page does not ask where you live. Idle draw is not counted here: that is the builder's job, and counting it twice would bill the same socket twice.
what this page will not tell you
How many tokens per second you will get. That is a benchmark — it depends on memory bandwidth, on the engine, on the context and on how long the answers are. This site does not invent benchmarks, and everyone who quotes one for your exact setup is guessing.
A card is not a server. The builder works out the rest — the machine, the drives, the electricity and the backup.