Autograd in Depth

Checking your gradients with gradcheck

gradcheck compares your backward rule against slow-but-honest numerical differentiation, catching wrong gradients that would otherwise fail silently for months.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

gradcheck tests a gradient rule by comparing it against a second, slower way of getting the same answer.

Think of checking a shop bill. The cashier's machine printed a total — probably right, but you re-add the items on your phone, slowly, item by item. Two independent methods, one answer. If they agree, you trust the bill. If they disagree, something is wrong, and you know before paying.

A hand-written backward rule is the cashier's total. gradcheck is the slow re-adding.

Why it exists

A custom backward rule with a mistake in it does not crash. It trains — badly, quietly, for months. Wrong gradients are among the most expensive silent bugs in all of machine learning, because the symptom is "results are mediocre", which has a hundred other explanations.

Fortunately, gradients can be measured without any rule. Nudge one input slightly, watch the output move, divide. Painfully slow — one nudge per input — but rule-free, so it cannot share your rule's mistake. Two independent answers, compared: that is the entire idea.

How it works

your rule:      backward()  ──────>  fast answer
honest nudges:  f(x + tiny) vs f(x - tiny), divide  ──>  slow answer

          agree (to many decimals)  ->  PASS
          disagree                  ->  FAIL, with both answers shown

A real example you have seen

Shops tally the cash drawer against the register's receipts at closing time. Two independent records of the same day — and disagreement means investigate now, not at year-end audit. Cheap independent verification, run routinely, catches expensive quiet errors early. gradcheck is closing-time tally for your maths.

Remember this

  • Gradients can be measured by nudging inputs — no rule involved.
  • gradcheck compares your rule against that measurement.
  • Every hand-written backward gets gradchecked. No exceptions.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs verified with torch 2.5.1, CPU.

A correct rule passes, a wrong one is caught

check.py
import torch
from torch.autograd import gradcheck

class Cube(torch.autograd.Function):
    @staticmethod
    def forward(ctx, x):
        ctx.save_for_backward(x)
        return x ** 3
    @staticmethod
    def backward(ctx, grad_out):
        (x,) = ctx.saved_tensors
        return grad_out * 3 * x ** 2

class BadCube(torch.autograd.Function):
    @staticmethod
    def forward(ctx, x):
        ctx.save_for_backward(x)
        return x ** 3
    @staticmethod
    def backward(ctx, grad_out):
        (x,) = ctx.saved_tensors
        return grad_out * 2 * x ** 2      # wrong on purpose: 2 instead of 3

torch.manual_seed(0)
probe = torch.randn(3, dtype=torch.double, requires_grad=True)
print(gradcheck(Cube.apply, (probe,)))

try:
    gradcheck(BadCube.apply, (probe,))
except Exception as err:
    print(type(err).__name__)
    print(str(err)[:420])
Output
True
GradcheckError
Jacobian mismatch for output 0 with respect to input 0,
numerical:tensor([[ 7.1240,  0.0000,  0.0000],
        [ 0.0000,  0.2583,  0.0000],
        [ 0.0000,  0.0000, 14.2414]], dtype=torch.float64)
analytical:tensor([[4.7493, 0.0000, 0.0000],
        [0.0000, 0.1722, 0.0000],
        [0.0000, 0.0000, 9.4942]], dtype=torch.float64)

The walkthrough

Read the failure like a detective. "Numerical" is the honest nudging; "analytical" is your rule. Divide any matching pair: 7.1240 / 4.7493 = 1.5 — exactly 3/2, the planted bug. Failure reports frequently hand you the ratio of your mistake, which is often enough to spot the missing factor by inspection.

The diagonal shape is information too. Element-wise operations have gradients only on the diagonal (each output touches one input). A mismatch off the diagonal in an operation you believed element-wise reveals accidental cross-talk — a broadcast or reshape bug in the forward, not the backward.

dtype=torch.double is not optional. The honest method computes (f(x+h) − f(x−h)) / 2h with h around 1e-6. In float32, subtracting near-equal numbers at that scale destroys most significant digits, and gradcheck drowns in noise — reporting failures for correct rules. Always probe in float64.

Probe small. gradcheck nudges every input element separately: an n-element input means n forward evaluations, and it builds full Jacobians. Three elements — fine. A production-sized batch — minutes of pointless work. Tiny random doubles are the correct probe; correctness at size 3 is correctness.

The habit that makes it worthwhile

Put the check in your test suite, not your training script:

test_cube.py
import torch
from torch.autograd import gradcheck

def test_cube_gradients():
    torch.manual_seed(0)
    probe = torch.randn(4, dtype=torch.double, requires_grad=True)
    assert gradcheck(Cube.apply, (probe,))

(Requires Cube imported from your code.) Now the backward rule is guarded: anyone editing the forward next year and forgetting the backward gets a red test in seconds, instead of a mediocre model in March.

Common mistakes

Testing in float32 and "fixing" the failure. The failure was the dtype's, not the rule's. Double precision first; only then believe the verdict.

Testing at unrepresentative points. Operations with kinks — like your Clamp01 from the previous lesson — are non-differentiable exactly at the kink (0 and 1). A probe landing there fails honestly but unhelpfully. Test away from kinks; test the kink's two sides separately if the behaviour there matters.

Skipping the check for straight-through estimators. Correct: STE backward rules are deliberately wrong, and gradcheck will fail them — the tool verifies mathematical truth, not usefulness. Do not gradcheck the lie; gradcheck everything you claim to be a true derivative.

Believing training success means gradients are right. Networks are forgiving; they train around moderate gradient errors and settle a few points below their potential. That forgiveness is precisely why the silent bug survives — and why measurement, not vibes, is the standard.

Forgetting requires_grad=True on the probe. gradcheck raises a clear complaint — the analytical side needs a graph to walk.

Try it yourself

Gradcheck the Clamp01 you wrote last lesson, probing at [-0.5, 0.3, 0.7, 1.5] (away from the kinks) — it should pass. Then probe at exactly 0.0 and 1.0, watch what happens, and write one sentence on why that failure is the mathematics being honest rather than the code being wrong.

What to learn next

Researcher — Mathematics and papers.

The method: central differences

gradcheck estimates each Jacobian column by symmetric perturbation:

J_{:,i} ≈ [f(x + h·e_i) − f(x − h·e_i)] / 2h

Where e_i is the i-th standard basis vector and h the perturbation (eps, default 1e-6). Taylor expansion cancels the even terms: error is O(h²) in truncation, versus O(h) for one-sided differences — the reason for the two-sided form. Against truncation error stands round-off error O(u/h) with unit round-off u; the total error is minimised near h* ≈ (u)^{1/3}. For float64 (u ≈ 1.1e-16), h* ≈ 5e-6 — the default sits at the sweet spot. For float32 (u ≈ 6e-8), even optimal h leaves ~1e-3 relative error: numerically incapable of certifying a gradient to the default tolerances, which is the precise version of "use double".

The analytical side is assembled column-by-column from VJPs: seeding backward with each e_j recovers row j (see non-scalar backward). Comparison uses allclose semantics with atol (default 1e-5) and rtol (1e-3), all configurable.

Cost and scope

Numerical side: 2n forward passes for n input elements; analytical side: m backward passes for m output elements; memory O(mn) for both Jacobians. Hence the doctrine of tiny probes — the test certifies the rule, which is size-independent, not the workload.

gradgradcheck extends the scheme to second derivatives, verifying your Function's double-backward (needed the moment anyone puts your op inside a gradient-penalty or higher-order training objective). Non-obvious options that matter in practice: check_undefined_grad (None-gradient handling), nondet_tol for legitimately nondeterministic kernels (atomics — see reproducibility), and fast_mode=True, which checks a random projection uᵀJv instead of the full Jacobian — O(1) passes rather than O(n), catching most bugs at a fraction of the cost, at the price of projection-direction blind spots.

Non-differentiable points, honestly

Central differences straddle kinks: at x = 0, |x| yields numerical estimate 0 while any valid subgradient choice in [−1, 1] may appear analytically — legitimate disagreement without error. ReLU-family ops adopt a convention at 0 (PyTorch: derivative 0); gradcheck at measure-zero kink sets tests the convention, not correctness. The measure-theoretic footing — a.e.-differentiability of the piecewise-smooth functions in practice — is treated rigorously in Griewank and Walther (2008), Evaluating Derivatives, ch. 14, and for the Clarke subdifferential view, Bolte and Pauwels (2020), A mathematical model for automatic differentiation in machine learning.

Lineage

Gradient checking predates autograd as the debugging ritual of hand-derived backprop — standard curriculum in Ng's CS231n-era materials ("always gradient check") — and survives today precisely at the boundary where derivation becomes manual again: custom Functions, custom kernels, and framework development itself (PyTorch's own operator test suite runs gradcheck across the operator catalogue via OpInfo entries).

References

  • Griewank and Walther (2008), Evaluating Derivatives, SIAM — finite-difference error analysis and AD verification.
  • Bolte and Pauwels (2020), NeurIPS — nonsmooth AD semantics.
  • PyTorch docs, torch.autograd.gradcheck — options, tolerances, fast mode.

What to learn next