TensorFlow and Keras

compile, fit and evaluate

Three calls take a Keras model from an untrained shell to a scored result — compile sets the rules, fit does the learning, evaluate checks it honestly.

On this page 5
  1. Why the split exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

compile tells the model how to learn, fit makes it practise on your data, and evaluate gives it a test it has not seen the answers to.

Think of coaching a student for an exam. Before practice starts, you agree the rules: what counts as a mistake, how to correct after each mistake, and what score to report to the parents. That agreement is compile — a one-time setup, before any learning happens.

Then come weeks of solving papers and correcting errors. That repetition is fit. Finally you hand over a mock exam the student has never seen. That honest check is evaluate.

Why the split exists

Learning needs three separate decisions, and Keras refuses to guess them:

  • Loss — the definition of "how wrong was that answer", as a single number.
  • Optimizer — the correction rule: how to nudge the model after each mistake.
  • Metrics — the scores a human wants to watch, like accuracy.

These choices change with the problem. Predicting a price and predicting spam-or-not need different definitions of wrong. So Keras makes you state them once, up front, in compile.

How it works

compile:   model + (loss, optimizer, metrics)  →  ready to learn

fit:       data  →  guess  →  measure wrongness  →  correct  →  repeat
           (one full pass over all data = 1 epoch)

evaluate:  unseen data  →  guess  →  score        (no correcting!)

An epoch is one complete pass through the training data. A batch is the small group of examples processed between corrections — the model does not wait to see everything before improving.

A real example you have seen

Every spam filter went through exactly this loop. Someone compiled a model with "wrongness = how confidently you misjudged this email". They fitted it on millions of labelled emails, then evaluated it on emails held back in a drawer. Only after the held-back score looked good did it reach your inbox.

Remember this

  • compile = choose loss, optimizer and metrics. No learning happens yet.
  • fit = the actual practice loop, measured in epochs and batches.
  • evaluate = scoring on data the model never trained on. The only score you should trust.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tensorflow numpy

Outputs verified with TensorFlow 2.21 (Keras 3.12), CPU. The seed makes runs repeatable on one machine; exact loss digits still drift across versions and hardware, so treat them as the shape of the story, not gospel.

The whole loop in one file

train_pass_fail.py
import numpy as np
import keras
from keras import layers

keras.utils.set_random_seed(0)          # same starting weights every run

# Hours studied, hours slept -> passed the exam (1) or not (0)
x = np.array([[8.0, 7.0], [7.5, 8.0], [6.0, 7.5], [5.5, 6.0],
              [2.0, 4.0], [1.5, 8.0], [3.0, 5.0], [0.5, 6.5]], dtype="float32")
y = np.array([1, 1, 1, 1, 0, 0, 0, 0], dtype="float32")

model = keras.Sequential([
    keras.Input(shape=(2,)),
    layers.Dense(4, activation="relu"),
    layers.Dense(1, activation="sigmoid"),
])

model.compile(optimizer=keras.optimizers.Adam(learning_rate=0.05),
              loss="binary_crossentropy",
              metrics=["accuracy"])

history = model.fit(x, y, epochs=200, batch_size=4, verbose=0)

for epoch in [0, 49, 99, 199]:
    print(f"epoch {epoch + 1:3d}  loss {history.history['loss'][epoch]:.4f}"
          f"  accuracy {history.history['accuracy'][epoch]:.2f}")

loss, acc = model.evaluate(x, y, verbose=0)
print(f"final: loss {loss:.4f}, accuracy {acc:.2f}")
Output
epoch   1  loss 1.5627  accuracy 0.00
epoch  50  loss 0.0616  accuracy 1.00
epoch 100  loss 0.0120  accuracy 1.00
epoch 200  loss 0.0029  accuracy 1.00
final: loss 0.0029, accuracy 1.00

The walkthrough

loss="binary_crossentropy" — the right wrongness measure for yes/no questions with a sigmoid output. The string is shorthand for keras.losses.BinaryCrossentropy(). Match the trio deliberately: sigmoid + binary crossentropy for two classes, softmax + categorical crossentropy for many, no activation + MSE for regression. Mixing the trios is the most common silent failure in Keras — see loss functions for why each pairing exists.

Adam(learning_rate=0.05) — Adam is the default-choice optimizer; the learning rate is high here because the dataset is eight rows. Real datasets want the default 0.001 as a starting point.

batch_size=4 — eight samples, batches of four, so two corrections per epoch. Smaller batches mean noisier but more frequent updates.

history — fit returns a log object. history.history is a plain dict of lists, one entry per epoch. Plot it before believing any model; reading training curves starts there.

evaluate here reuses the training data, which is cheating — done only to keep the file tiny. Real code holds data back with validation_split=0.2 inside fit, or a separate test set. The honest version of this topic is train/test split.

Common mistakes

Trusting training accuracy. The 1.00 above says the model memorised eight rows — nothing more. Only a held-out score means anything. Pass validation_split=0.2 to fit and watch val_loss instead.

Forgetting to compile. Calling fit on an uncompiled model raises ValueError: You must call compile() before using the model. The fix is to call compile first — but the deeper point is that loss and optimizer are not defaults Keras will pick for you.

Wrong loss for integer labels. Labels 0, 1, 2 with categorical_crossentropy crash with a shape error, because that loss expects one-hot rows. Integer labels want sparse_categorical_crossentropy. This one error message confuses everyone the first time.

Calling fit twice and thinking it restarts. A second fit continues from the current weights. That is a feature (it is how fine-tuning works), but if you wanted a fresh start, rebuild the model.

Try it yourself

Add validation_split=0.25 to the fit call and print history.history["val_loss"] at the same four epochs. With eight samples the validation set is two students — watch how jumpy the numbers are, and connect that to why tiny validation sets mislead.

What to learn next

Researcher — Mathematics and papers.

What fit runs per step

For parameters $\theta$, one optimizer step over a mini-batch $B$ minimises the empirical risk estimate:

$$ \theta_{t+1} = \theta_t - \eta \cdot A!\left(\nabla_\theta \frac{1}{|B|} \sum_{i \in B} \ell(f_\theta(x_i), y_i)\right) $$

Symbols: $\theta$ — all trainable weights; $\eta$ — the learning rate; $\ell$ — the per-sample loss; $f_\theta$ — the model; $|B|$ — batch size; $A$ — the optimizer's transformation of the raw gradient (identity for SGD; bias-corrected first/second moment scaling for Adam, Kingma and Ba 2015).

Binary crossentropy for label $y \in {0,1}$ and prediction $p$: $\ell = -y \log p - (1-y)\log(1-p)$. Keras computes it from logits internally when the last layer is sigmoid, folding sigmoid and log together for numerical stability (the log-sum-exp trick), which is why from_logits=True with no activation is exactly equivalent and slightly better conditioned.

Cost accounting

Per epoch: $\lceil N/|B| \rceil$ steps, each costing one forward and one backward pass — roughly 3× the FLOPs of inference-only. Memory is dominated by stored activations, $O(|B| \sum_l d_l)$, which is why the practical fix for out-of-memory is lowering batch size, not model size. Epoch loop overhead in Keras is amortised by compiling the train step into a graph — the mechanism is covered in tf.function and retracing.

The three-way contract

compile fixes the objective, fit performs stochastic optimisation, evaluate estimates generalisation. The statistical claim behind evaluate is that held-out loss is an unbiased estimate of expected loss under the data distribution — a claim that survives only if the test set never influenced any choice: not the weights, not the architecture, not the stopping epoch. Repeated evaluation during development erodes it (adaptive overfitting; Dwork et al. 2015, The reusable holdout).

References

  • Kingma and Ba (2015), Adam: a method for stochastic optimization.
  • Bottou (2010), Large-scale machine learning with stochastic gradient descent — why mini-batches beat full-batch.
  • Dwork et al. (2015), The reusable holdout: preserving validity in adaptive data analysis.

What to learn next