run a model
What you can run on an Apple Silicon, 64 GB unified
48 GB usable out of 64. macOS does not hand the whole of unified memory to the GPU — about three quarters of it is the honest figure.
fits comfortably
fits, but only just
does not fit
Figured at 4-bit weights — what almost everyone runs — with a 8k token context. Both change the numbers: shorter context is the first lever when something almost fits. Model details checked on 2026-08-02.
what generation costs to run
- 1 h a day
- €5 a year · 22 kWh
- 2 h a day
- €11 a year · 44 kWh
- 8 h a day
- €44 a year · 175 kWh
At 60 W under load and 0.25 € per kWh — the European average, since this page does not ask where you live. Idle draw is not counted here: that is the builder's job, and counting it twice would bill the same socket twice.
how to actually run it
what this page will not tell you
How many tokens per second you will get. That is a benchmark — it depends on memory bandwidth, on the engine, on the context and on how long the answers are. This site does not invent benchmarks, and everyone who quotes one for your exact setup is guessing.
A card is not a server. The builder works out the rest — the machine, the drives, the electricity and the backup.