TensorFlow and Keras

GradientTape

GradientTape is TensorFlow's recorder — it watches your calculations as they run, then walks the recording backwards to tell you every gradient.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

GradientTape is a recorder: while your calculation runs, it writes down every step, so TensorFlow can replay the recording backwards and work out what to blame for the result.

Think of a shopkeeper's carbon-copy receipt book. Every sale gets written down as it happens, without slowing the shop. At closing time, the shopkeeper flips back through the copies to answer: which item moved today's total the most?

The tape does that for maths. Run your calculation inside the recorder, and afterwards ask: if I nudge this input, how much does the final answer move? That "how much" is called a gradient — the direction and strength of influence each number has on the result.

Why it exists

Training a network means answering one question millions of times: which knob, turned which way, reduces the error? With thousands of knobs, testing each one by hand would take forever.

The recording trick answers it all at once. Because every operation was written down, TensorFlow can walk the record backwards from the error, splitting blame among all knobs in a single backward pass. This is the engine under every fit call — fit hides the tape; here you hold it directly.

How it works

  with recorder on:
      guess  =  knobs  combined with  inputs
      error  =  distance(guess, truth)

  recorder, played backwards:
      error  →  blame for each knob  →  turn each knob slightly  →  less error

The forward direction computes the answer. The backward replay computes the blame. Then you nudge every knob against its blame and repeat.

A real example you have seen

Phone cameras that sharpen night photos were trained this way. The model guessed a cleaner image, and the error against a real sharp photo was measured. Then a tape-like recorder assigned blame across millions of knobs — nudge, repeat, for weeks. The photo mode you tap is the settled result of that blame game.

Remember this

  • The tape records while the calculation runs, then replays it backwards for gradients.
  • A gradient = how much (and in which direction) one value influences the result.
  • fit uses this machinery invisibly; custom training loops hold the tape directly.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tensorflow

Outputs verified with TensorFlow 2.21, CPU. Everything below is deterministic — no seeds needed.

From one derivative to a full training loop

tape_basics.py
import tensorflow as tf

# 1. The smallest possible example
x = tf.Variable(3.0)
with tf.GradientTape() as tape:
    y = x ** 2
print("dy/dx at x=3:", tape.gradient(y, x).numpy())

# 2. Constants are not recorded unless you ask
c = tf.constant(3.0)
with tf.GradientTape() as tape:
    y = c ** 2
print("gradient for an unwatched constant:", tape.gradient(y, c))

with tf.GradientTape() as tape:
    tape.watch(c)
    y = c ** 2
print("after tape.watch:", tape.gradient(y, c).numpy())

# 3. A real training step, no Keras anywhere
w = tf.Variable(0.0)
b = tf.Variable(0.0)
xs = tf.constant([1.0, 2.0, 3.0, 4.0])
ys = tf.constant([3.0, 5.0, 7.0, 9.0])       # the truth is y = 2x + 1

for step in range(200):
    with tf.GradientTape() as tape:
        loss = tf.reduce_mean((w * xs + b - ys) ** 2)
    dw, db = tape.gradient(loss, [w, b])
    w.assign_sub(0.01 * dw)                   # walk downhill
    b.assign_sub(0.01 * db)

print(f"learned w={w.numpy():.2f}, b={b.numpy():.2f}, loss={loss.numpy():.4f}")
Output
dy/dx at x=3: 6.0
gradient for an unwatched constant: None
after tape.watch: 6.0
learned w=2.05, b=0.84, loss=0.0042

The walkthrough

6.0 is the calculus checking out: the derivative of x² is 2x, and 2 × 3 = 6. When learning autodiff, always start with a function whose derivative you can verify on paper.

tf.Variable vs tf.constant is the recording rule. Variables are watched automatically — they are the "knobs" training exists to turn. Constants are treated as fixed scenery and ignored, which is why the unwatched gradient came back None rather than zero. None means never recorded; zero would mean recorded, no influence. The distinction diagnoses many bugs.

The training loop is fit with the covers off. Loss inside the tape, tape.gradient(loss, [w, b]) for the blame, assign_sub to walk downhill with step size 0.01. Two hundred steps land near the truth (w=2, b=1); the remaining gap closes with more steps or a schedule — the walk itself is gradient descent.

One tape, one replay. After tape.gradient returns, the tape frees its recording. Ask twice and you get: RuntimeError: A non-persistent GradientTape can only be used to compute one set of gradients. When you genuinely need two replays — two different outputs, one recording — create it as tf.GradientTape(persistent=True) and del tape when done.

Common mistakes

Computing outside the with block. Any operation above the with line is off the record, and its gradient path is dead. The forward pass — all of it — goes inside.

Breaking the chain with .numpy(). Converting a tensor to a NumPy value mid-calculation leaves the tape's world; everything downstream is invisible to it. Stay in tensors until the loss is computed. The same rule exists in PyTorch as detach.

Expecting gradients through integers. Gradients flow through float tensors only. An int32 input silently yields None — check dtypes when a gradient is mysteriously missing.

Reassigning instead of assign_sub. Writing w = w - 0.01 * dw rebinds the Python name to a new tensor, and the original Variable never changes. In-place updates on Variables go through assign, assign_add, assign_sub.

Try it yourself

Replace the loss with mean absolute error: tf.reduce_mean(tf.abs(w * xs + b - ys)). Rerun and watch the loop still work — the tape differentiates the new expression without any other change. That indifference to what you wrote is the whole point of automatic differentiation.

What to learn next

Researcher — Mathematics and papers.

Reverse-mode autodiff

GradientTape implements reverse-mode automatic differentiation. The forward pass records a directed acyclic graph of primitive operations $v_1, \dots, v_n$; the backward pass propagates adjoints $\bar{v}_i = \partial L / \partial v_i$ from the output to every recorded node via the chain rule:

$$ \bar{v}i = \sum{j \in \text{consumers}(i)} \bar{v}_j \, \frac{\partial v_j}{\partial v_i} $$

Symbols: $L$ — the scalar output being differentiated; $v_i$ — an intermediate value; $\bar{v}_i$ — its adjoint (gradient of $L$ with respect to it); the sum runs over operations that consumed $v_i$.

Cost: one backward sweep computes gradients with respect to all $n$ inputs in $O(1)$ times the forward cost — against $O(n)$ forward passes for numerical or forward-mode differentiation. With millions of parameters and one scalar loss, this asymmetry is the reason deep learning is computable at all. The price is memory: intermediate activations must be retained until the sweep, $O(\text{ops})$ space, which is why long unrolled computations exhaust RAM and why gradient checkpointing (recompute instead of store; Chen et al. 2016, Training deep nets with sublinear memory cost) exists.

Tape versus graph-based autodiff

TF1 built a static graph and derived gradients symbolically ahead of time. The tape (TF2, 2019) records dynamically — per execution, following whatever Python control flow actually ran. This matches PyTorch's design (autograd graph mechanics) and makes data-dependent branching differentiable for free. The historical arc: Wengert (1964) introduced the recording-list formulation ("Wengert tape" — the name survives); Speelpenning (1980) the efficient reverse sweep; Rumelhart, Hinton, Williams (1986) its rediscovery for neural networks as backpropagation — the fuller story is in backpropagation.

Higher-order gradients nest tapes: an outer tape watching an inner tape's replay differentiates the differentiation, needed for Hessian-vector products and meta-learning (Finn et al. 2017, Model-agnostic meta-learning). Each nesting level multiplies memory retention accordingly.

References

  • Baydin et al. (2018), Automatic differentiation in machine learning: a survey — the standard modern reference.
  • Griewank and Walther (2008), Evaluating Derivatives — the textbook treatment, including tape memory trade-offs.
  • Agrawal et al. (2019), TensorFlow Eager: a multi-stage, Python-embedded DSL for machine learning.

What to learn next