Inside a Transformer Section 084

Quantised LLM Inference

Shrinking a language model so it fits and runs fast, and knowing exactly what that costs you in quality.

3 of 8 lessons published Three reading levels on every lesson

Start with “Compressing the KV cache”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. Compressing the KV cache
  2. fp32, bf16, fp8 and int4
  3. GPTQ
  4. AWQ soon
  5. GGUF and llama.cpp quantisation soon
  6. Measuring what quantisation costs you soon
  7. Separating prefill from decode soon
  8. CUDA graphs and kernel fusion soon