Best Ollama Models by RAM: What to Run on 4GB, 8GB, 16GB, and 32GB in 2026

The most common local-AI question in 2026 is still the simplest one: "I have X GB of RAM — what should I run?" The answer changed a lot in the last year as newer model generations got dramatically better at small sizes. A 4B model from the current crop now beats 7B models from two years ago, which means old "best models" lists are actively misleading.
Here's the current pick per RAM tier, based on actually running these on real hardware — with the honest speed expectations for CPU versus GPU inference.
The one rule that decides everything
Model weights at Q4 quantization (the sweet spot for quality-per-memory) need roughly 0.6–0.7 GB per billion parameters, plus 1–3 GB of overhead for context and the KV cache. Quick mental math for a machine with 16 GB RAM:
- Leave 4+ GB for the OS and browser → ~12 GB usable
- That's a 14B model at Q4 comfortably, or a 24B at Q4 if you close everything else and keep context short
If you have a GPU, the constraint is VRAM instead — and GPU VRAM is worth roughly 3–5x CPU RAM for usefulness, because offloaded layers generate several times faster.
4 GB RAM: small, fast, surprisingly capable
The era of "4 GB can't run LLMs" is over. Current small models handle summarization, rewriting, and simple Q&A genuinely well.
| Model | Size on disk | Context | CPU speed feel |
|---|---|---|---|
| Qwen3 1.7B (Q4) | ~1.1 GB | 32k | Snappy, instant-feeling replies |
| Llama 3.2 3B (Q4) | ~2.0 GB | 128k | Usable, a paragraph in a few seconds |
| Phi-4-mini (Q4) | ~2.6 GB | 128k | Slightly slower, strong reasoning for its size |
| Gemma 3 4B (Q4) | ~2.6 GB | 128k | Comfortable if it's the only big app open |
Pick: Qwen3 1.7B for speed, Gemma 3 4B if you can spare the RAM. Avoid anything 7B+ — it'll thrash.
8 GB RAM: the sweet spot for capable small models
8 GB machines (base Mac mini territory, older laptops) are the first tier where local AI feels like a real assistant rather than a toy.
| Model | Size on disk | Context | Notes |
|---|---|---|---|
| Llama 3.1 8B / Llama 3.8B-class (Q4) | ~4.7 GB | 128k | The default pick; balanced and reliable |
| Qwen3 8B (Q4) | ~5.0 GB | 32k+ | Stronger coding + math, similar speed |
| Gemma 3 12B (Q4) | ~7.3 GB | 128k | Fits, but close everything else; strong multimodal option |
| DeepSeek-R1 distill 7B/8B (Q4) | ~4.5–5 GB | 64k | Visible chain-of-thought reasoning; fun to watch think |
Pick: Qwen3 8B or Llama 3.1 8B. Go Gemma 3 12B only if 6+ GB is genuinely free — at Q4 with a browser open, 12B on 8 GB is flirting with swap.
16 GB RAM: where local AI gets serious
This is the tier most mid-range laptops and desktops sit at in 2026, and it unlocks 14B-class models that rival cloud models from a couple of years ago.
| Model | Size on disk | Context | Notes |
|---|---|---|---|
| Qwen3 14B (Q4) | ~9 GB | 32k+ | The best all-rounder at this tier |
| Phi-4 14B (Q4) | ~9 GB | 16k | Excellent reasoning, shorter context |
| Gemma 3 27B (Q4) | ~16 GB | 128k | Only fits at Q3 or with very short context — usually skip on 16 GB |
| DeepSeek-R1 distill 14B (Q4) | ~9 GB | 64k | Reasoning workloads, math, logic puzzles |
Pick: Qwen3 14B at Q4 as the daily driver. On CPU expect roughly 6–10 tokens/sec — reading speed, fine for chat. With a GPU holding most layers, 25–40 tokens/sec.
32 GB RAM: flagship territory
32 GB comfortably runs the models people actually mean when they say "local AI is as good as the cloud now."
| Model | Size on disk | Context | Notes |
|---|---|---|---|
| Qwen3 32B (Q4) | ~20 GB | 32k+ | GPT-4-class quality for most everyday tasks |
| Gemma 3 27B (Q4) | ~16 GB | 128k | Now comfortable; multimodal; excellent writing |
| Qwen3 30B-A3B (MoE) (Q4) | ~18 GB | 32k+ | The speed hack — MoE runs like a 3B, thinks like a 30B |
| Llama 3.3 70B (Q2/Q3) | ~26 GB | 128k | Fits only at aggressive quants; quality loss is real but usable |
Pick: Qwen3 30B-A3B if you want speed, Gemma 3 27B or Qwen3 32B for maximum quality. On CPU, dense 32B is slow (~3–5 tok/s) — this tier really wants a GPU with 12+ GB VRAM, where it flies at 20–40 tok/s.
Speed expectations: set them before you install
Honest numbers for Q4 models, single-user chat:
- Modern CPU only: 7–8B → 10–20 tok/s; 14B → 6–10 tok/s; 32B → 3–5 tok/s
- GPU (8 GB VRAM, most layers offloaded): 7–8B → 40–70 tok/s; 14B → 25–40 tok/s
- GPU (16+ GB VRAM): 32B fully offloaded → 20–40 tok/s
Tokens-per-second over ~15 feels instant for chat. Under 5 is where people give up — so on CPU-only machines, cap yourself at 14B.
Getting started takes five minutes
Install Ollama, run ollama pull qwen3:14b, and you're off. The nicer upgrade is skipping the terminal entirely: a browser chat interface that talks to your local models keeps history on your device and feels like a normal AI product. The private AI chat on Practical Web Tools does exactly that — no account, no uploads, works with locally running models and cloud models alike.
FAQ
Does GPU VRAM count toward these tiers?
Yes — arguably more than system RAM. A model fully held in VRAM runs several times faster than one split across CPU RAM. An 8 GB GPU running an 8B model beats 32 GB of RAM running the same model.
Is Q4 quantization really good enough?
For 8B+ models, Q4_K_M-type quants lose roughly 1–3% versus full precision on most benchmarks — imperceptible in chat. Below 4B parameters, quantization hurts more; consider Q6/Q8 for tiny models since they're small anyway.
Can I run a model bigger than my RAM?
Technically yes via mmap streaming, practically no — it's orders of magnitude slower and will hammer your SSD. Stay within the tier table.
Why is my model so slow even though it fits?
Three usual suspects: context window set very high (KV cache eats bandwidth), the model is split between CPU and GPU awkwardly, or a browser with 40 tabs is competing for RAM. Check all three before blaming the model.
Are these model names stable on Ollama's registry?
Tags shift as versions update (e.g., qwen3:14b today may point to a newer revision later). Pin exact tags if reproducibility matters.
The bottom line
Match the model to the machine, not the hype: 4 GB → Qwen3 1.7B, 8 GB → Qwen3 8B, 16 GB → Qwen3 14B, 32 GB → Qwen3 30B-A3B or Gemma 3 27B. Every tier now delivers genuinely useful output — the difference between tiers is how much you can ask of it, not whether it works.


