Nico

What fits on one RTX 3080

I wanted a language model running on my own PC, on the graphics card I already had: an RTX 3080 with 10 GB of memory. These are the numbers I measured, in case they save someone an evening.

The setup

LM Studio running Qwythos 9B, an uncensored nine-billion-parameter model compressed down to a 6.2 GB file. The part that lets it read images is another 0.9 GB.

Context is the limit, not the model

The model itself fits easily. What fills the card is context, meaning how much text the model can hold in its head at once.

Out of the box it loaded with about 24,000 tokens of context, split across four parallel slots. That was too small. A single Gmail page handed to the model as a snapshot overflowed it.

  • 65,536 tokens, one slot. Everything stays on the card. It uses about 9.9 of the 10 GB and runs at around 84 tokens a second once warm.
  • 131,072 tokens. It loads, but spills into ordinary RAM and runs about three times slower.

So 64K is the practical ceiling on this card. More system RAM doesn't move it, because the slow part is leaving the graphics card at all.

The settings that gave me 12 GB back

With the defaults, the model server held 13.9 GB of system RAM on top of the graphics memory. Turning off "keep model in memory" and memory-mapping dropped that to 1.3 GB. The whole system went from 28.2 GB in use to 16.7 GB.

Generation ran at 40 to 60 tokens a second afterwards, slower than the 84 I'd seen before. I haven't worked out how much of that is down to those two settings.

Where it goes wrong

  • With its thinking step switched off it answered faster and was wrong more often.
  • It invented an import for a library I had installed. The code was confident, tidy, and not real.

A model this size is useful for sorting and summarising. I check anything it tells me about code.


← All notes