What hardware do I need for a tiny LLM?

Updated October 2026 · How we answer

Short answerFor a tiny LLM (roughly 0.5B to 3B parameters), a modern laptop with 8-16 GB of RAM can run quantized models on CPU. A GPU with 6-12 GB of VRAM makes it much faster, and a Raspberry Pi 5 with 8 GB can handle the smallest models slowly.

Memory is the real constraint

Model size in memory depends on parameter count and precision. A rough rule: at 4-bit quantization, you need about 0.5 to 0.7 GB of memory per billion parameters, plus overhead for the context window.

So a 1B model at 4-bit fits in roughly 1 GB, a 3B model in about 2-3 GB, and a 7B model in about 4-6 GB. Longer conversations and bigger context windows add more, sometimes a lot more.

That means a machine with 8 GB of RAM can comfortably run 1B-3B models on CPU, and 16 GB opens up 7B-8B models at 4-bit. Unified-memory machines, like recent Apple Silicon laptops, share one pool between CPU and GPU, which helps.

  • 8 GB RAM: smallest models (0.5B-3B), CPU only, modest speed
  • 16 GB RAM: 7B-8B at 4-bit, usable on CPU, faster with a GPU
  • 6-8 GB VRAM GPU: small models run fast, 7B at 4-bit often fits
  • 12-24 GB VRAM GPU: comfortable for 7B-14B quantized models
  • Raspberry Pi 5 (8 GB): sub-1B to 1B models, slow but workable

CPU vs. GPU vs. phone

CPU inference works and needs no special hardware, but expect a few tokens per second on a laptop for a 3B model. A discrete GPU can push that into the tens or hundreds of tokens per second, depending on the model and card.

You do not need a data-center card. Consumer GPUs from the last several years handle tiny models well, and integrated graphics are improving but still limited by shared memory bandwidth.

Phones and single-board computers can run very small models. They are fine for experiments, offline demos, and simple tasks, but they are not a substitute for a desktop if you want interactive speed.

Common mistakes

  • Buying a big GPU before checking whether a quantized model already runs fine on the machine you own.
  • Ignoring context length; a model that fits at 2K tokens may run out of memory at 32K.
  • Assuming VRAM and system RAM are interchangeable; a model must fit in the memory type the runtime actually uses.
From our shopsCaseMorph: Type an idea, see a custom phone case in seconds, then print a one-of-one.