LLM Development

Ollama — run LLMs locally

Ollama downloads a language model onto your own computer and runs it there, so nothing you type leaves your machine and nothing costs money per question.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The word that explains everything: quantized
  6. Where you have already seen this idea
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Ollama downloads a language model onto your own computer and runs it there.

The analogy you have already lived

You have downloaded songs before a long train journey. The tunnel kills your signal, but the music keeps playing, because the file is already on the phone.

Ollama does that for AI. It pulls the model onto your disk once. After that it answers with the Wi-Fi switched off.

Streaming needs signal and eats data. Downloading needs storage and works anywhere. The same trade-off applies here, exactly.

Why it exists

Using ChatGPT or a similar service means your words travel to somebody else's computer. That is fine for a poem. It is not fine for a friend's medical report, an unreleased assignment, or a company's customer list.

There were three problems with the online-only route.

Privacy. Whatever you paste is now on a server you do not control.

Cost. Every question costs a small amount of money. A loop that runs ten thousand times stops being small.

The internet. Hostel Wi-Fi dies. A train goes through a tunnel. Your demo is tomorrow morning.

Running a model on your own machine was possible before Ollama, and it was miserable. You compiled C++ code, hunted for the right model file, and guessed at settings. Ollama turned all of that into one command.

How it works

   ollama pull llama3.2:1b
          |
          v
   a 1.3 GB file lands on your disk
          |
   ollama run llama3.2:1b
          |
          v
   a small server starts on your own machine
          |
          v
   your chat, or your Python code, talks to it
          (no internet needed from here on)

The download happens once. The server is a program on your own computer. It listens on a private address that nobody outside your machine can reach.

The word that explains everything: quantized

A large model stores millions of learned numbers, and each number normally takes two bytes of memory. Quantizing means storing those numbers roughly instead of exactly, using about half a byte each.

Think of writing down a price as "about 500 rupees" instead of "497.83 rupees". You lose a little accuracy. You save a lot of space.

Almost every model in Ollama is quantized. That is why a model needing 15 GB of memory fits into 5 GB. It is also why it runs on a laptop at all.

The loss is real. A quantized model is slightly worse than the original. For most learning and building, the trade is worth it.

Where you have already seen this idea

  • Offline maps in Google Maps, downloaded before a trip.
  • Downloaded episodes on Netflix or YouTube for a flight.
  • Offline dictionaries and offline translation packs.

Same pattern: pay in storage once, stop depending on the network.

The honest part

This is the section people skip and then get angry about.

Small models are noticeably weaker. A model you can run on a laptop is not the model behind ChatGPT. It will confuse facts, lose track of long instructions, and make up citations. It is genuinely useful for summarising, rewriting, classifying and learning. It is not a research assistant.

Your laptop probably cannot run the big ones. A seven-billion-parameter model needs roughly 8 GB of memory free. That is 8 GB after your operating system and browser have taken theirs. On an 8 GB laptop it means closing everything, and it will still be slow.

Slow means slow. Without a graphics card, expect a few words per second. Reading this paragraph aloud is faster than the model will write it.

Start with a 1-billion-parameter model. It downloads in minutes, runs on anything, and teaches you the whole workflow. Move up only when you have hit its limits and know why.

Remember this

  • Ollama puts the model on your machine, so your text never leaves it.
  • Models are quantized — stored roughly to save memory, at a small cost in quality.
  • Start small. A 1B model on a laptop teaches you more than a 7B model that never loads.

What to learn next

Developer — Code and libraries.

Install

macOS and Windows: download the installer from ollama.com/download and run it. It installs a background service that starts with your machine.

Linux:

bash
curl -fsSL https://ollama.com/install.sh | sh

Piping a script from the internet into a shell deserves a pause. Open the URL in a browser and read it first. It is short.

Check the install:

bash
ollama --version
Output
ollama version is 0.32.14

Your version number will be higher. The commands below have been stable for a long time.

Your first model, end to end

bash
ollama pull llama3.2:1b
Output
pulling manifest
pulling 74701a8c35f6: 100% ▕██████████████████▏ 1.3 GB
pulling fcc5a6bec9da: 100% ▕██████████████████▏ 7.7 KB
verifying sha256 digest
writing manifest
success

The tag after the colon is the size. llama3.2:1b is the 1-billion-parameter build. Leaving the tag off gives you :latest, which for many models is a much larger download than you expected.

Now talk to it:

bash
ollama run llama3.2:1b "Reply with one word. What colour is the sky on a clear day?"
Output
Blue.

Run it without a prompt and you get an interactive chat. Type /bye to leave, /? for the in-chat commands.

The commands you will actually use

CommandWhat it does
ollama pull <model>Download without running
ollama run <model>Chat, downloading first if needed
ollama listEvery model on your disk, with sizes
ollama psModels currently loaded in memory
ollama stop <model>Unload it and free the memory now
ollama show <model>Architecture, context length, quantization
ollama rm <model>Delete it from disk
ollama serveStart the server manually, in the foreground

ollama show is the one people ignore and should not:

bash
ollama show llama3.2:1b
Output
  Model
    architecture        llama
    parameters          1.2B
    context length      131072
    embedding length    2048
    quantization        Q8_0

  Capabilities
    completion
    tools

quantization Q8_0 means eight bits per stored number. Q4_K_M — the more common one — means roughly four. Capabilities: tools tells you this model was trained to request function calls; many small models were not.

Sizes, honestly

Download sizes below were read from the Ollama registry, so they are real. Memory is what the model needs while loaded, and depends on your settings.

Model tagDownloadRuns comfortably on
qwen3:0.6b0.52 GBany laptop, 4 GB RAM
gemma3:1b0.82 GB4 GB RAM
llama3.2:1b1.32 GB8 GB RAM
llama3.2:3b2.02 GB8 GB RAM, closing other apps
gemma3:4b3.34 GB16 GB RAM
mistral:7b4.37 GB16 GB RAM, or 8 GB GPU
llama3.1:8b4.92 GB16 GB RAM, or 8 GB GPU

"Runs comfortably" assumes a browser and an editor are also open, because they always are.

The trap that costs beginners a whole evening

The file size is not the memory size. Here is ollama ps for the same 1.3 GB model, loaded twice with different settings:

Output
NAME              ID              SIZE      PROCESSOR    CONTEXT
llama3.2:1b       baf6a787fdff    6.4 GB    100% GPU     131072
Output
NAME              ID              SIZE      PROCESSOR    CONTEXT
llama3.2:1b       baf6a787fdff    1.5 GB    100% GPU     4096

Same model file. 6.4 GB versus 1.5 GB. The difference is the context length — how many tokens of conversation the model can hold at once.

Holding conversation needs a scratchpad in memory called the KV cache, and that scratchpad grows in direct proportion to the context length. A 131,072-token window reserves space for a whole book you are never going to send.

So when a model refuses to load and your maths said it should fit, this is almost always why. Cut the context:

bash
OLLAMA_CONTEXT_LENGTH=4096 ollama serve

Or set it per request, as shown next. The vLLM lesson works through the arithmetic behind this.

Talking to it from Python

Ollama exposes an HTTP server on localhost:11434. Nothing needs to be installed to use it.

ask.py
import json
import urllib.request

payload = {
    "model": "llama3.2:1b",
    "prompt": "Name the capital of France in one word.",
    "stream": False,                       # False = wait and return the whole answer
    "options": {"temperature": 0, "num_ctx": 4096},
}

request = urllib.request.Request(
    "http://localhost:11434/api/generate",
    data=json.dumps(payload).encode(),
    headers={"Content-Type": "application/json"},
)

with urllib.request.urlopen(request) as response:
    result = json.loads(response.read())

print("answer      :", result["response"])
print("stop reason :", result["done_reason"])
print("tokens in   :", result["prompt_eval_count"])
print("tokens out  :", result["eval_count"])
Output
answer      : Paris.
stop reason : stop
tokens in   : 34
tokens out  : 3

The answer text will not always match. temperature: 0 makes the model pick its most likely next word every time, which removes most randomness on one machine. It does not guarantee the same words on different hardware or a different Ollama build. Token counts shift with the model's prompt template too. Treat the shape of this output as fixed and the exact words as a sample.

Prefer a real library once you are past the first script:

bash
pip install ollama

It wraps the same endpoints, handles streaming, and matches the OpenAI message format.

Making your own model with a Modelfile

A Modelfile is a small recipe: start from an existing model, bake in a system prompt and settings, save it under a new name.

Modelfile
FROM llama3.2:1b

PARAMETER temperature 0.2
PARAMETER num_ctx 4096

SYSTEM """
You are a study helper for Indian college students.
Answer in at most three short sentences. Never invent exam dates.
"""
bash
ollama create study-buddy -f Modelfile
ollama run study-buddy "What is gradient descent?"
Output
using existing layer sha256:74701a8c35f6...
creating new layer sha256:65e8cfd66c08...
writing manifest
success

ollama list will now show study-buddy at 1.3 GB. That is a display quirk: the weights are shared with llama3.2:1b, so the extra disk cost is a few kilobytes.

Common mistakes

Forgetting the size tag. ollama pull llama3.1 fetches the 8B build — nearly 5 GB. On a metered connection that hurts. Always write the tag.

Leaving the default context length. Covered above, and it is the top cause of "it worked yesterday" memory failures.

Expecting temperature: 0 to be a guarantee. It removes the sampling randomness. It does not make output reproducible across machines, versions or GPU drivers.

Exposing the server to the network. Setting OLLAMA_HOST=0.0.0.0 puts an unauthenticated model server on your Wi-Fi. Anyone on that network can use it, and on a public network that is a real problem. Keep the default 127.0.0.1 unless you have added authentication in front.

Judging a model by one bad answer. Change the system prompt and lower the temperature before concluding a model is useless. Small models are far more sensitive to prompt wording than large ones.

Letting models sit in memory. Ollama keeps a model loaded for five minutes after the last request. If RAM is tight, ollama stop <model>, or set OLLAMA_KEEP_ALIVE=30s.

Try it yourself

Pull qwen3:0.6b and llama3.2:3b. Ask both the same five questions from your own coursework. Time them with ollama run --verbose. You will get a concrete feel for the size-versus-quality curve that no benchmark table can give you, on your actual hardware.

Then write a Modelfile that turns one of them into a tool you would use weekly — a commit-message writer, a Marathi-to-English helper, a flashcard maker.

What to learn next

Researcher — Mathematics and papers.

What Ollama is, architecturally

Ollama is a Go service wrapping llama.cpp (Gerganov, 2023) as its inference engine, plus a content-addressed model store using an OCI-style manifest and blob layout, plus an HTTP API.

The abstraction it adds over llama.cpp is packaging: a model reference resolves to a manifest listing blobs for weights, template, parameters and licence. Blobs are shared across tags, so derived models created with ollama create cost only the changed layers. This is why a Modelfile that only alters the system prompt adds kilobytes rather than gigabytes.

Weights are stored in GGUF, a single-file format holding tensors plus key-value metadata (architecture, tokenizer, chat template, RoPE parameters). It is memory-mappable, so load time is dominated by page faults rather than parsing.

Quantization, concretely

Q4_K_M is a k-quant block format. Weights are split into blocks; each block stores 4-bit quantized values plus per-block scale and minimum in higher precision. The _M suffix marks a mixed assignment: attention and feed-forward tensors that are more sensitive get 5 or 6 bits, the rest get 4.

Effective bits per weight land near 4.8 rather than 4.0, because of scale overhead. Memory for weights is therefore:

weight_bytes  ≈  n_params × bits_per_weight / 8

For an 8B model at ~4.8 bits: 8e9 × 4.8 / 8 ≈ 4.8 GB, which matches the observed 4.92 GB for llama3.1:8b.

Perplexity degradation for Q4_K_M on Llama-class models is typically under 1% relative to fp16, and Q8_0 is close to lossless. Below 4 bits the curve steepens sharply, and IQ2-class quants are usable only with importance-matrix calibration. Dettmers and Zettlemoyer (2023) argue 4-bit is near the accuracy-per-bit optimum for a fixed memory budget — a larger model at 4 bits generally beats a smaller model at 8.

The KV cache is the real budget

Weights are static. The KV cache is what scales with load, and it is what the Developer tab's 6.4 GB versus 1.5 GB observation is measuring:

kv_bytes  =  2 × n_layers × n_kv_heads × head_dim × bytes_per_elem × n_ctx × n_parallel
  • 2 — one key tensor and one value tensor per layer.
  • n_layers — transformer blocks.
  • n_kv_heads — key/value heads. Under grouped-query attention (Ainslie et al., 2023) this is far smaller than the query-head count, which is the single largest reduction available.
  • head_dim — dimension per head.
  • bytes_per_elem — 2 for f16, 1 for q8_0 cache, 0.5 for q4_0.
  • n_ctx — context length.
  • n_parallel — concurrent sequences (OLLAMA_NUM_PARALLEL).

For Llama-3.1-8B: 2 × 32 × 8 × 128 × 2 = 131,072 bytes = 128 KiB per token. At 128k context that is 16 GiB of cache for one sequence — larger than the quantized weights by a factor of three.

OLLAMA_KV_CACHE_TYPE=q8_0 halves it, at a small quality cost that is usually invisible. OLLAMA_FLASH_ATTENTION=1 reduces attention working memory, not cache size.

Why CPU inference is slow: it is memory bandwidth, not FLOPs

Autoregressive decoding at batch size 1 reads every weight from memory to produce one token. Arithmetic intensity is roughly one multiply-accumulate per weight loaded, so the kernel is bandwidth-bound, not compute-bound.

tokens_per_second  ≲  memory_bandwidth / weight_bytes

A laptop with dual-channel DDR4-3200 has about 51 GB/s theoretical, and perhaps 35 GB/s achievable. Against 4.8 GB of weights that ceiling is roughly 7 tokens/second, before any overhead. An RTX 4090 has about 1000 GB/s, and an H100 with HBM3 about 3.35 TB/s. The two-orders-of-magnitude gap in observed speed is almost entirely this ratio.

Prompt processing behaves differently. It is a batched matrix multiply over all prompt tokens at once, so it is compute-bound and far more efficient per token. This asymmetry is why a long prompt costs less wall-clock time than generating an equally long answer.

Speculative decoding (Leviathan et al., 2022; Chen et al., 2023) attacks the bandwidth bound directly: a small draft model proposes k tokens, the target model verifies them in one batched forward pass, and rejection sampling preserves the target distribution exactly. Typical speedups are 2–3× with no quality loss.

Where Ollama is the wrong tool

Ollama serves requests through llama.cpp's scheduler. It is built for one user, or a handful. It does not implement continuous batching with paged attention, so throughput under concurrency is far below a purpose-built server.

For multi-user serving, vLLM, SGLang or TensorRT-LLM are the correct choices, and the difference at high concurrency is measured in multiples, not percentages. Ollama's design centre is a developer laptop, and it is excellent there.

References

What to learn next

What to learn next

These follow on from what you just read.

  • LLM Development

    vLLM

    vLLM is a server that runs one language model on a GPU for many users at once, by never letting the graphics card sit idle waiting.

  • LLM Development

    LangChain

    LangChain is a toolkit that joins prompts, models and parsers into one pipeline, so you can swap any part without rewriting the rest.

  • LLM Development

    LlamaIndex

    LlamaIndex loads your documents, cuts them into searchable pieces and finds the right piece when a question arrives, so a model can answer from your files.