Skip to content
selfhostr

run a model

What you can run on an NVIDIA GeForce RTX 5060 Ti 16 GB (MSI Shadow 2X OC Plus)

16 GB usable out of 16.

fits comfortably

Qwen3.5 4BRuns on a laptop with no graphics card. Rewriting, summarising, sorting — not reasoning.2.7 GB
Llama 3.1 8BThe one every tutorial assumes. Old enough that every problem you hit has already been answered somewhere.5.4 GB
Qwen3.5 9BThe everyday assistant, and the smallest one that answers usefully about a document you paste in.6.1 GB
Gemma 4 12BGoogle's middle size: good at ordinary language, and it reads images.8.2 GB
Phi-4 14BPunches above its size on reasoning and maths for how little memory it takes.9.5 GB

fits, but only just

gpt-oss 20BOpenAI's open-weight model, in the size meant for one machine.13.6 GB

does not fit

Gemma 4 26B-A4BA mixture of experts: it loads like a 25-billion model and answers faster than one.17.1 GB
Qwen3.6 27BThe size where a machine at home starts to replace a paid subscription for most everyday work.18.4 GB
Qwen3-Coder 30B-A3BBuilt for code, and the one to run if that is the whole reason you are here.20.4 GB
Gemma 4 31BThe largest Gemma that still fits on one consumer card.20.9 GB
DeepSeek-R1 Distill 32BShows its reasoning as it works. Slower to answer, and worth it on hard questions.21.8 GB
Qwen3.6 35B-A3BOne notch above 27B, and the point where 24 GB stops being enough.23.8 GB
Llama 3.1 70BMeta's large one. Needs hardware bought for this and nothing else.47.6 GB
gpt-oss 120BFits on a single 80 GB card, or a Mac with a lot of unified memory.81.6 GB

Figured at 4-bit weights — what almost everyone runs — with a 8k token context. Both change the numbers: shorter context is the first lever when something almost fits. Model details checked on 2026-08-02.

what generation costs to run

1 h a day
€16 a year · 66 kWh
2 h a day
€33 a year · 131 kWh
8 h a day
€131 a year · 526 kWh

At 180 W under load and 0.25 € per kWh — the European average, since this page does not ask where you live. Idle draw is not counted here: that is the builder's job, and counting it twice would bill the same socket twice.

how to actually run it

what this page will not tell you

How many tokens per second you will get. That is a benchmark — it depends on memory bandwidth, on the engine, on the context and on how long the answers are. This site does not invent benchmarks, and everyone who quotes one for your exact setup is guessing.

Price the machine around it

A card is not a server. The builder works out the rest — the machine, the drives, the electricity and the backup.