AI glossary

GGUF

In one sentence GGUF is the file format for running quantized LLMs on ordinary computers — one self-contained file that llama.cpp, Ollama and LM Studio all read.

By Updated

GGUF is a single-file format for packaging LLMs — weights, tokenizer and metadata together, usually quantized — designed for running models on everyday hardware.

Think of it as the MP3 of language models. Studio recordings are enormous and need studio equipment; the MP3 compressed them into a single file any phone could play, and that is what made music portable. GGUF plays the same role: a 16-bit research checkpoint needs a serious GPU, while its GGUF version — quantized to 4 or 5 bits — runs on a student's laptop, CPU included.

The format grew out of the llama.cpp project (successor to its GGML format, hence the name) and its design choices explain its spread. Everything lives in one file — weights, tokenizer, chat template, architecture metadata — so "download one file, point the runner at it" is the whole install. It is memory-mappable, meaning the OS pages weights in as needed rather than loading everything up front. And it carries many quantization levels, named like Q4_K_M:

Q8_0    ~8 bits/weight   near-lossless, biggest
Q5_K_M  ~5.5 bits        strong quality, popular
Q4_K_M  ~4.8 bits        the usual sweet spot
Q2_K    ~2.6 bits        heavy damage — last resort

Rule of thumb: pick the largest quant that fits your RAM/VRAM with room for the KV cache. Ollama, LM Studio, llamafile and llama.cpp all consume GGUF directly, and popular models appear on Hugging Face in GGUF within days of release — it is the de facto standard for local LLMs. Note it is an inference format: for training or fine-tuning you go back to the original checkpoint.

Where to go next