Ollama — run LLMs locally
Ollama downloads a language model onto your own computer and runs it there, so nothing you type leaves your machine and nothing costs money per question.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Ollama downloads a language model onto your own computer and runs it there.
The analogy you have already lived
You have downloaded songs before a long train journey. The tunnel kills your signal, but the music keeps playing, because the file is already on the phone.
Ollama does that for AI. It pulls the model onto your disk once. After that it answers with the Wi-Fi switched off.
Streaming needs signal and eats data. Downloading needs storage and works anywhere. The same trade-off applies here, exactly.
Why it exists
Using ChatGPT or a similar service means your words travel to somebody else's computer. That is fine for a poem. It is not fine for a friend's medical report, an unreleased assignment, or a company's customer list.
There were three problems with the online-only route.
Privacy. Whatever you paste is now on a server you do not control.
Cost. Every question costs a small amount of money. A loop that runs ten thousand times stops being small.
The internet. Hostel Wi-Fi dies. A train goes through a tunnel. Your demo is tomorrow morning.
Running a model on your own machine was possible before Ollama, and it was miserable. You compiled C++ code, hunted for the right model file, and guessed at settings. Ollama turned all of that into one command.
How it works
ollama pull llama3.2:1b
|
v
a 1.3 GB file lands on your disk
|
ollama run llama3.2:1b
|
v
a small server starts on your own machine
|
v
your chat, or your Python code, talks to it
(no internet needed from here on)The download happens once. The server is a program on your own computer. It listens on a private address that nobody outside your machine can reach.
The word that explains everything: quantized
A large model stores millions of learned numbers, and each number normally takes two bytes of memory. Quantizing means storing those numbers roughly instead of exactly, using about half a byte each.
Think of writing down a price as "about 500 rupees" instead of "497.83 rupees". You lose a little accuracy. You save a lot of space.
Almost every model in Ollama is quantized. That is why a model needing 15 GB of memory fits into 5 GB. It is also why it runs on a laptop at all.
The loss is real. A quantized model is slightly worse than the original. For most learning and building, the trade is worth it.
Where you have already seen this idea
- Offline maps in Google Maps, downloaded before a trip.
- Downloaded episodes on Netflix or YouTube for a flight.
- Offline dictionaries and offline translation packs.
Same pattern: pay in storage once, stop depending on the network.
The honest part
This is the section people skip and then get angry about.
Small models are noticeably weaker. A model you can run on a laptop is not the model behind ChatGPT. It will confuse facts, lose track of long instructions, and make up citations. It is genuinely useful for summarising, rewriting, classifying and learning. It is not a research assistant.
Your laptop probably cannot run the big ones. A seven-billion-parameter model needs roughly 8 GB of memory free. That is 8 GB after your operating system and browser have taken theirs. On an 8 GB laptop it means closing everything, and it will still be slow.
Slow means slow. Without a graphics card, expect a few words per second. Reading this paragraph aloud is faster than the model will write it.
Start with a 1-billion-parameter model. It downloads in minutes, runs on anything, and teaches you the whole workflow. Move up only when you have hit its limits and know why.
Remember this
- Ollama puts the model on your machine, so your text never leaves it.
- Models are quantized — stored roughly to save memory, at a small cost in quality.
- Start small. A 1B model on a laptop teaches you more than a 7B model that never loads.
What to learn next
- vLLM — the same idea built for many users at once, on a GPU.
- Function calling and tools — let a local model run your Python.
- Local LLM assistant — build something with what you now have installed.
Developer — Code and libraries.
Install
macOS and Windows: download the installer from ollama.com/download and run it. It installs a background service that starts with your machine.
Linux:
curl -fsSL https://ollama.com/install.sh | shPiping a script from the internet into a shell deserves a pause. Open the URL in a browser and read it first. It is short.
Check the install:
ollama --versionollama version is 0.32.14
Your version number will be higher. The commands below have been stable for a long time.
Your first model, end to end
ollama pull llama3.2:1bpulling manifest pulling 74701a8c35f6: 100% ▕██████████████████▏ 1.3 GB pulling fcc5a6bec9da: 100% ▕██████████████████▏ 7.7 KB verifying sha256 digest writing manifest success
The tag after the colon is the size. llama3.2:1b is the 1-billion-parameter build. Leaving the tag off gives you :latest, which for many models is a much larger download than you expected.
Now talk to it:
ollama run llama3.2:1b "Reply with one word. What colour is the sky on a clear day?"Blue.
Run it without a prompt and you get an interactive chat. Type /bye to leave, /? for the in-chat commands.
The commands you will actually use
| Command | What it does |
|---|---|
ollama pull <model> | Download without running |
ollama run <model> | Chat, downloading first if needed |
ollama list | Every model on your disk, with sizes |
ollama ps | Models currently loaded in memory |
ollama stop <model> | Unload it and free the memory now |
ollama show <model> | Architecture, context length, quantization |
ollama rm <model> | Delete it from disk |
ollama serve | Start the server manually, in the foreground |
ollama show is the one people ignore and should not:
ollama show llama3.2:1b Model
architecture llama
parameters 1.2B
context length 131072
embedding length 2048
quantization Q8_0
Capabilities
completion
toolsquantization Q8_0 means eight bits per stored number. Q4_K_M — the more common one — means roughly four. Capabilities: tools tells you this model was trained to request function calls; many small models were not.
Sizes, honestly
Download sizes below were read from the Ollama registry, so they are real. Memory is what the model needs while loaded, and depends on your settings.
| Model tag | Download | Runs comfortably on |
|---|---|---|
qwen3:0.6b | 0.52 GB | any laptop, 4 GB RAM |
gemma3:1b | 0.82 GB | 4 GB RAM |
llama3.2:1b | 1.32 GB | 8 GB RAM |
llama3.2:3b | 2.02 GB | 8 GB RAM, closing other apps |
gemma3:4b | 3.34 GB | 16 GB RAM |
mistral:7b | 4.37 GB | 16 GB RAM, or 8 GB GPU |
llama3.1:8b | 4.92 GB | 16 GB RAM, or 8 GB GPU |
"Runs comfortably" assumes a browser and an editor are also open, because they always are.
The trap that costs beginners a whole evening
The file size is not the memory size. Here is ollama ps for the same 1.3 GB model, loaded twice with different settings:
NAME ID SIZE PROCESSOR CONTEXT llama3.2:1b baf6a787fdff 6.4 GB 100% GPU 131072
NAME ID SIZE PROCESSOR CONTEXT llama3.2:1b baf6a787fdff 1.5 GB 100% GPU 4096
Same model file. 6.4 GB versus 1.5 GB. The difference is the context length — how many tokens of conversation the model can hold at once.
Holding conversation needs a scratchpad in memory called the KV cache, and that scratchpad grows in direct proportion to the context length. A 131,072-token window reserves space for a whole book you are never going to send.
So when a model refuses to load and your maths said it should fit, this is almost always why. Cut the context:
OLLAMA_CONTEXT_LENGTH=4096 ollama serveOr set it per request, as shown next. The vLLM lesson works through the arithmetic behind this.
Talking to it from Python
Ollama exposes an HTTP server on localhost:11434. Nothing needs to be installed to use it.
import json
import urllib.request
payload = {
"model": "llama3.2:1b",
"prompt": "Name the capital of France in one word.",
"stream": False, # False = wait and return the whole answer
"options": {"temperature": 0, "num_ctx": 4096},
}
request = urllib.request.Request(
"http://localhost:11434/api/generate",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(request) as response:
result = json.loads(response.read())
print("answer :", result["response"])
print("stop reason :", result["done_reason"])
print("tokens in :", result["prompt_eval_count"])
print("tokens out :", result["eval_count"])answer : Paris. stop reason : stop tokens in : 34 tokens out : 3
The answer text will not always match. temperature: 0 makes the model pick its most likely next word every time, which removes most randomness on one machine. It does not guarantee the same words on different hardware or a different Ollama build. Token counts shift with the model's prompt template too. Treat the shape of this output as fixed and the exact words as a sample.
Prefer a real library once you are past the first script:
pip install ollamaIt wraps the same endpoints, handles streaming, and matches the OpenAI message format.
Making your own model with a Modelfile
A Modelfile is a small recipe: start from an existing model, bake in a system prompt and settings, save it under a new name.
FROM llama3.2:1b
PARAMETER temperature 0.2
PARAMETER num_ctx 4096
SYSTEM """
You are a study helper for Indian college students.
Answer in at most three short sentences. Never invent exam dates.
"""ollama create study-buddy -f Modelfile
ollama run study-buddy "What is gradient descent?"using existing layer sha256:74701a8c35f6... creating new layer sha256:65e8cfd66c08... writing manifest success
ollama list will now show study-buddy at 1.3 GB. That is a display quirk: the weights are shared with llama3.2:1b, so the extra disk cost is a few kilobytes.
Common mistakes
Forgetting the size tag. ollama pull llama3.1 fetches the 8B build — nearly 5 GB. On a metered connection that hurts. Always write the tag.
Leaving the default context length. Covered above, and it is the top cause of "it worked yesterday" memory failures.
Expecting temperature: 0 to be a guarantee. It removes the sampling randomness. It does not make output reproducible across machines, versions or GPU drivers.
Exposing the server to the network. Setting OLLAMA_HOST=0.0.0.0 puts an unauthenticated model server on your Wi-Fi. Anyone on that network can use it, and on a public network that is a real problem. Keep the default 127.0.0.1 unless you have added authentication in front.
Judging a model by one bad answer. Change the system prompt and lower the temperature before concluding a model is useless. Small models are far more sensitive to prompt wording than large ones.
Letting models sit in memory. Ollama keeps a model loaded for five minutes after the last request. If RAM is tight, ollama stop <model>, or set OLLAMA_KEEP_ALIVE=30s.
Try it yourself
Pull qwen3:0.6b and llama3.2:3b. Ask both the same five questions from your own coursework. Time them with ollama run --verbose. You will get a concrete feel for the size-versus-quality curve that no benchmark table can give you, on your actual hardware.
Then write a Modelfile that turns one of them into a tool you would use weekly — a commit-message writer, a Marathi-to-English helper, a flashcard maker.
What to learn next
- vLLM — the same idea built for many users at once, on a GPU.
- Function calling and tools — let a local model run your Python.
- Local LLM assistant — build something with what you now have installed.
Researcher — Mathematics and papers.
What Ollama is, architecturally
Ollama is a Go service wrapping llama.cpp (Gerganov, 2023) as its inference engine, plus a content-addressed model store using an OCI-style manifest and blob layout, plus an HTTP API.
The abstraction it adds over llama.cpp is packaging: a model reference resolves to a manifest listing blobs for weights, template, parameters and licence. Blobs are shared across tags, so derived models created with ollama create cost only the changed layers. This is why a Modelfile that only alters the system prompt adds kilobytes rather than gigabytes.
Weights are stored in GGUF, a single-file format holding tensors plus key-value metadata (architecture, tokenizer, chat template, RoPE parameters). It is memory-mappable, so load time is dominated by page faults rather than parsing.
Quantization, concretely
Q4_K_M is a k-quant block format. Weights are split into blocks; each block stores 4-bit quantized values plus per-block scale and minimum in higher precision. The _M suffix marks a mixed assignment: attention and feed-forward tensors that are more sensitive get 5 or 6 bits, the rest get 4.
Effective bits per weight land near 4.8 rather than 4.0, because of scale overhead. Memory for weights is therefore:
weight_bytes ≈ n_params × bits_per_weight / 8For an 8B model at ~4.8 bits: 8e9 × 4.8 / 8 ≈ 4.8 GB, which matches the observed 4.92 GB for llama3.1:8b.
Perplexity degradation for Q4_K_M on Llama-class models is typically under 1% relative to fp16, and Q8_0 is close to lossless. Below 4 bits the curve steepens sharply, and IQ2-class quants are usable only with importance-matrix calibration. Dettmers and Zettlemoyer (2023) argue 4-bit is near the accuracy-per-bit optimum for a fixed memory budget — a larger model at 4 bits generally beats a smaller model at 8.
The KV cache is the real budget
Weights are static. The KV cache is what scales with load, and it is what the Developer tab's 6.4 GB versus 1.5 GB observation is measuring:
kv_bytes = 2 × n_layers × n_kv_heads × head_dim × bytes_per_elem × n_ctx × n_parallel2— one key tensor and one value tensor per layer.n_layers— transformer blocks.n_kv_heads— key/value heads. Under grouped-query attention (Ainslie et al., 2023) this is far smaller than the query-head count, which is the single largest reduction available.head_dim— dimension per head.bytes_per_elem— 2 for f16, 1 for q8_0 cache, 0.5 for q4_0.n_ctx— context length.n_parallel— concurrent sequences (OLLAMA_NUM_PARALLEL).
For Llama-3.1-8B: 2 × 32 × 8 × 128 × 2 = 131,072 bytes = 128 KiB per token. At 128k context that is 16 GiB of cache for one sequence — larger than the quantized weights by a factor of three.
OLLAMA_KV_CACHE_TYPE=q8_0 halves it, at a small quality cost that is usually invisible. OLLAMA_FLASH_ATTENTION=1 reduces attention working memory, not cache size.
Why CPU inference is slow: it is memory bandwidth, not FLOPs
Autoregressive decoding at batch size 1 reads every weight from memory to produce one token. Arithmetic intensity is roughly one multiply-accumulate per weight loaded, so the kernel is bandwidth-bound, not compute-bound.
tokens_per_second ≲ memory_bandwidth / weight_bytesA laptop with dual-channel DDR4-3200 has about 51 GB/s theoretical, and perhaps 35 GB/s achievable. Against 4.8 GB of weights that ceiling is roughly 7 tokens/second, before any overhead. An RTX 4090 has about 1000 GB/s, and an H100 with HBM3 about 3.35 TB/s. The two-orders-of-magnitude gap in observed speed is almost entirely this ratio.
Prompt processing behaves differently. It is a batched matrix multiply over all prompt tokens at once, so it is compute-bound and far more efficient per token. This asymmetry is why a long prompt costs less wall-clock time than generating an equally long answer.
Speculative decoding (Leviathan et al., 2022; Chen et al., 2023) attacks the bandwidth bound directly: a small draft model proposes k tokens, the target model verifies them in one batched forward pass, and rejection sampling preserves the target distribution exactly. Typical speedups are 2–3× with no quality loss.
Where Ollama is the wrong tool
Ollama serves requests through llama.cpp's scheduler. It is built for one user, or a handful. It does not implement continuous batching with paged attention, so throughput under concurrency is far below a purpose-built server.
For multi-user serving, vLLM, SGLang or TensorRT-LLM are the correct choices, and the difference at high concurrency is measured in multiples, not percentages. Ollama's design centre is a developer laptop, and it is excellent there.
References
- Gerganov et al., llama.cpp — github.com/ggml-org/llama.cpp
- Dettmers and Zettlemoyer, The case for 4-bit precision: k-bit inference scaling laws, 2023 — arxiv.org/abs/2212.09720
- Ainslie et al., GQA: Training Generalized Multi-Query Transformer Models, 2023 — arxiv.org/abs/2305.13245
- Leviathan et al., Fast Inference from Transformers via Speculative Decoding, 2022 — arxiv.org/abs/2211.17192
- Frantar et al., GPTQ: Accurate Post-Training Quantization, 2022 — arxiv.org/abs/2210.17323
- Lin et al., AWQ: Activation-aware Weight Quantization, 2023 — arxiv.org/abs/2306.00978
What to learn next
- vLLM — the same idea built for many users at once, on a GPU.
- Function calling and tools — let a local model run your Python.
- Local LLM assistant — build something with what you now have installed.