GPTQ
GPTQ quantises a layer one column at a time and pushes each rounding error into the weights it has not reached yet, so the layer's output stays close to the original.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
GPTQ rounds the weights one column at a time. After each column it adjusts the columns not yet reached, to make up for the error.
The analogy you have already lived
Six of you split a restaurant bill of 2,347 rupees. Nobody wants coins, so each share gets rounded to the nearest ten.
Round each share on its own and the six rounded shares do not add up to 2,347. Somebody is short.
The sensible way is different. Round the first share. Notice you are eight rupees over. Take eight rupees off the next share before rounding that. The errors cancel instead of piling up.
That is exactly GPTQ.
Why it exists
The plain way to shrink a model is to round every weight to the nearest allowed value, independently.
At 8 bits that is fine. At 4 bits it is not. Each weight moves a little. All those little moves add up. The layer's output drifts away from where it should be.
Rounding is unavoidable. Rounding errors accumulating is not.
How it works
Two ingredients.
A little real data. You run a few hundred real examples through the model and record what actually arrives at each layer. This is called calibration — measuring what the layer really sees, instead of guessing.
Error carried forward. Round the first column of weights. Measure how far off you are. Nudge the remaining columns to compensate. Move to the next column. Repeat.
column 1 round it -> error e1
adjust cols 2..n to cancel e1
column 2 round it -> error e2
adjust cols 3..n to cancel e2
...By the last column, the errors that would have piled up have been absorbed along the way.
The condition it depends on
This only helps if the layer's inputs are related to each other.
Say input channel 5 and channel 6 tend to move together. Then over-shooting on channel 5 can be cancelled by under-shooting on channel 6. The calibration data is what tells you that.
Suppose every input were unrelated to every other. There would be nothing to cancel with, and the method would do nothing.
Real layer inputs are strongly related. The measurement below shows GPTQ cutting the error more than fourfold in that case. In the imaginary case of unrelated inputs, it does nothing at all.
What is honestly hard here
The calibration data matters, and it is easy to be careless about.
Calibrate on English Wikipedia and serve Hindi customer queries, and you have tuned the compensation for the wrong inputs. Calibrate on ten examples and you have measured noise.
A few hundred samples from the distribution you will actually serve is the practical rule.
Where you have already seen this
- Carrying a remainder forward when splitting a bill.
- A carpenter cutting one plank slightly long because the last cut ran short.
- Rounding a grocery bill and adjusting the last item so the total matches.
Remember this
- Round column by column, and push each error into the columns not yet done.
- It needs calibration data, and it needs the inputs to be related to each other.
- Get the calibration data wrong and you have optimised for the wrong workload.
What to learn next
- AWQ — the other main 4-bit method, with a different idea.
- Measuring what quantisation costs you — how to compare the two honestly.
- GGUF and llama.cpp quantisation — the CPU-first family of formats.
Developer — Code and libraries.
Setup
pip install torchRuns on a CPU in a couple of seconds.
GPTQ implemented, and its condition demonstrated
import torch
torch.manual_seed(0)
D_IN, D_OUT, NCAL, BITS = 256, 128, 1024, 4
QMAX = 2 ** (BITS - 1) - 1
W0 = torch.randn(D_OUT, D_IN) * 0.05 # the layer to be quantised
def rtn(W, group):
"""Round to nearest: every weight quantised on its own."""
out = W.clone()
for g in range(0, W.shape[1], group):
b = W[:, g:g + group]
s = b.abs().amax(1, keepdim=True) / QMAX
out[:, g:g + group] = (b / s).round().clamp(-QMAX - 1, QMAX) * s
return out
def gptq(W0, X, group, damp=0.01):
"""Quantise column by column, pushing each rounding error into the columns
not yet done, weighted by the inverse Hessian of the calibration data."""
W = W0.clone().double()
H = X.double().T @ X.double() # the Hessian
H += damp * torch.diag(H).mean() * torch.eye(H.shape[0], dtype=torch.double)
Hinv = torch.linalg.cholesky(
torch.cholesky_inverse(torch.linalg.cholesky(H)), upper=True)
Q = torch.zeros_like(W)
for g in range(0, W.shape[1], group):
end = min(g + group, W.shape[1])
s = W[:, g:end].abs().amax(1, keepdim=True) / QMAX # one scale per group
for j in range(g, end):
q = (W[:, j:j + 1] / s).round().clamp(-QMAX - 1, QMAX) * s
Q[:, j:j + 1] = q
err = (W[:, j:j + 1] - q) / Hinv[j, j]
W[:, j:] -= err @ Hinv[j:j + 1, j:] # spread the error forward
return Q.float()
def experiment(name, A):
"""A mixes independent noise into the layer's input channels."""
X = torch.randn(NCAL, D_IN) @ A
Xt = torch.randn(4096, D_IN) @ A
ref = Xt @ W0.T
corr = torch.corrcoef(X.T).abs()
off = corr[~torch.eye(D_IN, dtype=bool)].mean()
print(f"\n=== {name} (mean |correlation| between input channels: {off:.3f}) ===")
print(f"{'group size':>11} {'round-to-nearest':>18} {'GPTQ':>10} {'improvement':>13}")
for group in (256, 128, 32):
a = ((Xt @ rtn(W0, group).T - ref).norm() / ref.norm()).item()
b = ((Xt @ gptq(W0, X, group).T - ref).norm() / ref.norm()).item()
print(f"{group:>11} {a:>18.5f} {b:>10.5f} {a / b:>12.1f}x")
# Real transformer activations are strongly correlated across channels.
mix = torch.randn(D_IN, D_IN) / D_IN ** 0.5
mix = mix + 2.0 * torch.randn(D_IN, 16) @ torch.randn(16, D_IN) / D_IN ** 0.5
experiment("correlated inputs, like a real layer", mix)
experiment("independent inputs (a fantasy)", torch.eye(D_IN))=== correlated inputs, like a real layer (mean |correlation| between input channels: 0.206) ===
group size round-to-nearest GPTQ improvement
256 0.12150 0.02753 4.4x
128 0.11281 0.02572 4.4x
32 0.09426 0.02191 4.3x
=== independent inputs (a fantasy) (mean |correlation| between input channels: 0.025) ===
group size round-to-nearest GPTQ improvement
256 0.12586 0.13488 0.9x
128 0.11736 0.12553 0.9x
32 0.09647 0.10309 0.9xWritten against PyTorch 2.5.1, CPU, seeded — reproducible on this build.
The two tables together are the whole lesson
With correlated inputs, GPTQ cut the output error by a factor of 4.4. Same 4 bits, same group size, same weights. All that changed is the order of operations and a Hessian computed from 1,024 calibration samples.
With independent inputs, GPTQ was slightly worse than plain rounding: 0.9×. This is not a bug in the code — it is the method's precondition failing. Error compensation works by cancelling one column's error against later columns. If the columns are uncorrelated, there is nothing to cancel against, and perturbing the later weights only adds noise.
Real activations sit firmly in the first regime. Mean absolute off-diagonal correlation of 0.206 in the first experiment against 0.025 in the second. Transformer layer inputs are heavily correlated, which is why GPTQ works on real models.
Notice how little the group size mattered for GPTQ. 0.0275, 0.0257, 0.0219 across an 8× change in group size, against 0.1215 → 0.0943 for plain rounding. Error compensation does most of the work that small groups would otherwise have to do — so GPTQ can use larger groups and spend fewer bits on scales.
damp is not decoration. The Hessian from a finite sample is often singular or near-singular, and torch.linalg.cholesky fails on it. Adding 1% of the mean diagonal is the standard ridge. Raise it if you see a Cholesky error; the published implementations do exactly that in a retry loop.
Using it for real
GPTQ's active backend is GPTQModel; auto-gptq is no longer supported in Transformers.
from transformers import AutoModelForCausalLM, GPTQConfig
cfg = GPTQConfig(bits=4, group_size=128, dataset="c4", tokenizer=tok)
model = AutoModelForCausalLM.from_pretrained("your/model", quantization_config=cfg)No output block — this needs a GPU, a real model and a calibration download, and it takes minutes to hours. Two things about that snippet are worth internalising:
dataset="c4" is a default, not a recommendation. Pass your own calibration texts if your workload is not generic English web text.
group_size=128 is the usual choice, and combined with GPTQ's error compensation it is generally enough. Group 32 costs 0.5 bits per weight for a small further gain.
Common mistakes
Calibrating on the wrong distribution. The single largest source of disappointing GPTQ results. Use samples that look like your traffic.
Too few calibration samples. 128 sequences of 2048 tokens is the common setting. With 8 samples the Hessian is rank-deficient and the compensation is fitted to noise.
Quantising the embeddings and the output head. Skip them. Nearly every recipe does.
Using act_order=False when quality matters. Ordering columns by activation magnitude (also called desc_act) quantises the important channels first, while the compensation budget is largest. It costs inference speed on some kernels.
Assuming a 4-bit model is 4× faster. It is roughly 4× smaller. Speed depends on whether a kernel exists for your format on your hardware, and on whether you are memory-bound.
Comparing GPTQ and AWQ on a single perplexity number. See measuring what quantisation costs you.
Try it yourself
Change damp from 0.01 to 1.0 and re-run. Heavy damping pushes the Hessian toward a diagonal, which removes the correlation information — and GPTQ's advantage should shrink toward the second table's numbers. That is the mechanism, isolated with one number.
What to learn next
- AWQ — the other main 4-bit method, with a different idea.
- Measuring what quantisation costs you — how to compare the two honestly.
- GGUF and llama.cpp quantisation — the CPU-first family of formats.
Researcher — Mathematics and papers.
The layer-wise objective
GPTQ (Frantar et al., 2023) does not minimise the weight error. It minimises the output error of each layer, independently:
$$ \hat W = \arg\min_{W'} \; \lVert W X - W' X \rVert_2^2 \quad \text{s.t. } W' \text{ quantised} $$
$X \in \mathbb{R}^{d_{\text{in}} \times N}$ holds $N$ calibration activations. The objective is quadratic in $W'$, with Hessian
$$ H = 2 X X^{\top} $$
which is why calibration data is required at all: $H$ is the calibration data, in second-moment form. Note $H$ is shared by every output row, so one Cholesky serves the whole matrix.
From OBS to GPTQ
The lineage is Optimal Brain Surgeon (Hassibi and Stork, 1992) → OBQ (Frantar and Alistarh, 2022) → GPTQ.
OBQ greedily selects the weight whose quantisation costs least, quantises it, and applies the closed-form optimal update to the remaining unquantised weights:
$$ \delta_F = -\frac{w_q - \mathrm{quant}(w_q)}{[H_F^{-1}]_{qq}} \cdot \left(H_F^{-1}\right)_{:,q} $$
with $F$ the set of not-yet-quantised indices. This is exact, and it costs $O(d_{\text{row}} \cdot d_{\text{col}}^3)$ — days for a billion-parameter model.
GPTQ makes three changes that turn it into an hour:
- Fixed order. Quantise columns left to right rather than greedily. Empirically the ordering barely matters for large layers, and it means every row uses the same $H^{-1}$, so it is computed once.
- Lazy batch updates. Apply updates to a block of 128 columns at a time and only propagate to the rest of the matrix once per block. This turns a memory-bandwidth-bound inner loop into compute-bound matrix multiplies.
- Cholesky reformulation. Repeated row/column removal from $H^{-1}$ is numerically unstable at scale. Precomputing the upper Cholesky factor of $H^{-1}$ gives exactly the required sequence of rows, stably.
Reported cost: OPT-175B and BLOOM-176B quantised to 3–4 bits in about four GPU-hours, with near-negligible perplexity increase — the first demonstration that 175B-class models could be quantised to 4 bits without retraining.
act_order / desc_act
Processing columns in descending order of $\mathrm{diag}(H)$ — the most active input channels first — measurably improves quality, especially at 3 bits and on smaller models. The reason is that the compensation budget is largest early, so the channels with the most influence on the output get the best treatment.
Its cost is structural: with group-wise scales, reordering means group members are no longer contiguous, requiring an index permutation at inference. Some fast kernels (older Marlin paths) do not support that combination, so desc_act is a real quality-versus-throughput decision rather than a free win.
Where GPTQ is weak
It is layer-wise and greedy. Each layer is optimised assuming the previous layers are unquantised, so errors compound across depth. Nagel et al., 2020 (AdaRound) and Li et al., 2021 (BRECQ) address this with block-wise reconstruction; Shao et al., 2024 (OmniQuant) learn clipping thresholds and equivalent transformations by gradient descent on a block objective, and beat GPTQ especially at 3 bits and below.
It overfits the calibration set at low bit widths. At 2–3 bits, results become noticeably sensitive to which calibration texts were used. Reporting a 3-bit result without stating the calibration set is not a reproducible claim.
It does not touch activations. GPTQ is a weight-only method. For W8A8 you need SmoothQuant or LLM.int8() as well.
GPTQ against AWQ
They optimise different things and the comparison is genuinely close.
| GPTQ | AWQ | |
|---|---|---|
| Uses | second-order error compensation | per-channel scaling before rounding |
| Needs | Hessian from calibration | activation magnitudes from calibration |
| Sensitivity to calibration | higher | lower (magnitudes are more robust than a Hessian) |
| Quantisation time | minutes to hours | usually faster |
| Reordering complications | desc_act breaks some kernels | none |
Published head-to-heads disagree by model and bit width. The practical advice is to run both on your own evaluation and pick, rather than to trust a table.
Kernels
A GPTQ checkpoint is only fast if a matching kernel exists. The main families:
- ExLlamaV2 — fast, widely available, batch-1 oriented.
- Marlin — mixed-precision int4×fp16 kernels reaching near-ideal 4× speed-up at moderate batch sizes; needs Ampere or later and constrains group size and ordering.
- Triton / torch fallback — always works, often slower than fp16 at large batch, because dequantisation cost dominates when you are compute-bound rather than memory-bound.
That last point is the one people are surprised by: weight-only 4-bit quantisation speeds up small-batch decoding and can slow down large-batch prefill.
Papers
- Hassibi and Stork, Second Order Derivatives for Network Pruning: Optimal Brain Surgeon, NeurIPS 1992.
- Nagel et al., Up or Down? Adaptive Rounding for Post-Training Quantization (AdaRound), ICML 2020 — arxiv.org/abs/2004.10568
- Li et al., BRECQ, ICLR 2021 — arxiv.org/abs/2102.05426
- Frantar and Alistarh, Optimal Brain Compression, NeurIPS 2022 — arxiv.org/abs/2208.11580
- Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, ICLR 2023 — arxiv.org/abs/2210.17323
- Shao et al., OmniQuant, ICLR 2024 — arxiv.org/abs/2308.13137
- Frantar et al., MARLIN: Mixed-precision Auto-Regressive Parallel Inference on Large Language Models, 2024 — arxiv.org/abs/2408.11743
What to learn next
- AWQ — the other main 4-bit method, with a different idea.
- Measuring what quantisation costs you — how to compare the two honestly.
- GGUF and llama.cpp quantisation — the CPU-first family of formats.