Inside a Transformer Section 084
Quantised LLM Inference
Shrinking a language model so it fits and runs fast, and knowing exactly what that costs you in quality.
3 of 8 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Compressing the KV cache
- fp32, bf16, fp8 and int4
- GPTQ
- AWQ soon
- GGUF and llama.cpp quantisation soon
- Measuring what quantisation costs you soon
- Separating prefill from decode soon
- CUDA graphs and kernel fusion soon