compile, fit and evaluate
Three calls take a Keras model from an untrained shell to a scored result — compile sets the rules, fit does the learning, evaluate checks it honestly.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
compile tells the model how to learn, fit makes it practise on your data, and evaluate gives it a test it has not seen the answers to.
Think of coaching a student for an exam. Before practice starts, you agree the rules: what counts as a mistake, how to correct after each mistake, and what score to report to the parents. That agreement is compile — a one-time setup, before any learning happens.
Then come weeks of solving papers and correcting errors. That repetition is fit. Finally you hand over a mock exam the student has never seen. That honest check is evaluate.
Why the split exists
Learning needs three separate decisions, and Keras refuses to guess them:
- Loss — the definition of "how wrong was that answer", as a single number.
- Optimizer — the correction rule: how to nudge the model after each mistake.
- Metrics — the scores a human wants to watch, like accuracy.
These choices change with the problem. Predicting a price and predicting spam-or-not need different definitions of wrong. So Keras makes you state them once, up front, in compile.
How it works
compile: model + (loss, optimizer, metrics) → ready to learn
fit: data → guess → measure wrongness → correct → repeat
(one full pass over all data = 1 epoch)
evaluate: unseen data → guess → score (no correcting!)An epoch is one complete pass through the training data. A batch is the small group of examples processed between corrections — the model does not wait to see everything before improving.
A real example you have seen
Every spam filter went through exactly this loop. Someone compiled a model with "wrongness = how confidently you misjudged this email". They fitted it on millions of labelled emails, then evaluated it on emails held back in a drawer. Only after the held-back score looked good did it reach your inbox.
Remember this
- compile = choose loss, optimizer and metrics. No learning happens yet.
- fit = the actual practice loop, measured in epochs and batches.
- evaluate = scoring on data the model never trained on. The only score you should trust.
What to learn next
- Callbacks — reacting to training while it runs, including stopping at the right time.
- Loss functions — choosing the definition of wrong.
- Overfitting and underfitting — why training accuracy lies.
Developer — Code and libraries.
Setup
pip install tensorflow numpyOutputs verified with TensorFlow 2.21 (Keras 3.12), CPU. The seed makes runs repeatable on one machine; exact loss digits still drift across versions and hardware, so treat them as the shape of the story, not gospel.
The whole loop in one file
import numpy as np
import keras
from keras import layers
keras.utils.set_random_seed(0) # same starting weights every run
# Hours studied, hours slept -> passed the exam (1) or not (0)
x = np.array([[8.0, 7.0], [7.5, 8.0], [6.0, 7.5], [5.5, 6.0],
[2.0, 4.0], [1.5, 8.0], [3.0, 5.0], [0.5, 6.5]], dtype="float32")
y = np.array([1, 1, 1, 1, 0, 0, 0, 0], dtype="float32")
model = keras.Sequential([
keras.Input(shape=(2,)),
layers.Dense(4, activation="relu"),
layers.Dense(1, activation="sigmoid"),
])
model.compile(optimizer=keras.optimizers.Adam(learning_rate=0.05),
loss="binary_crossentropy",
metrics=["accuracy"])
history = model.fit(x, y, epochs=200, batch_size=4, verbose=0)
for epoch in [0, 49, 99, 199]:
print(f"epoch {epoch + 1:3d} loss {history.history['loss'][epoch]:.4f}"
f" accuracy {history.history['accuracy'][epoch]:.2f}")
loss, acc = model.evaluate(x, y, verbose=0)
print(f"final: loss {loss:.4f}, accuracy {acc:.2f}")epoch 1 loss 1.5627 accuracy 0.00 epoch 50 loss 0.0616 accuracy 1.00 epoch 100 loss 0.0120 accuracy 1.00 epoch 200 loss 0.0029 accuracy 1.00 final: loss 0.0029, accuracy 1.00
The walkthrough
loss="binary_crossentropy" — the right wrongness measure for yes/no questions with a sigmoid output. The string is shorthand for keras.losses.BinaryCrossentropy(). Match the trio deliberately: sigmoid + binary crossentropy for two classes, softmax + categorical crossentropy for many, no activation + MSE for regression. Mixing the trios is the most common silent failure in Keras — see loss functions for why each pairing exists.
Adam(learning_rate=0.05) — Adam is the default-choice optimizer; the learning rate is high here because the dataset is eight rows. Real datasets want the default 0.001 as a starting point.
batch_size=4 — eight samples, batches of four, so two corrections per epoch. Smaller batches mean noisier but more frequent updates.
history — fit returns a log object. history.history is a plain dict of lists, one entry per epoch. Plot it before believing any model; reading training curves starts there.
evaluate here reuses the training data, which is cheating — done only to keep the file tiny. Real code holds data back with validation_split=0.2 inside fit, or a separate test set. The honest version of this topic is train/test split.
Common mistakes
Trusting training accuracy. The 1.00 above says the model memorised eight rows — nothing more. Only a held-out score means anything. Pass validation_split=0.2 to fit and watch val_loss instead.
Forgetting to compile. Calling fit on an uncompiled model raises ValueError: You must call compile() before using the model. The fix is to call compile first — but the deeper point is that loss and optimizer are not defaults Keras will pick for you.
Wrong loss for integer labels. Labels 0, 1, 2 with categorical_crossentropy crash with a shape error, because that loss expects one-hot rows. Integer labels want sparse_categorical_crossentropy. This one error message confuses everyone the first time.
Calling fit twice and thinking it restarts. A second fit continues from the current weights. That is a feature (it is how fine-tuning works), but if you wanted a fresh start, rebuild the model.
Try it yourself
Add validation_split=0.25 to the fit call and print history.history["val_loss"] at the same four epochs. With eight samples the validation set is two students — watch how jumpy the numbers are, and connect that to why tiny validation sets mislead.
What to learn next
- Callbacks — reacting to training while it runs, including stopping at the right time.
- Loss functions — choosing the definition of wrong.
- Overfitting and underfitting — why training accuracy lies.
Researcher — Mathematics and papers.
What fit runs per step
For parameters $\theta$, one optimizer step over a mini-batch $B$ minimises the empirical risk estimate:
$$ \theta_{t+1} = \theta_t - \eta \cdot A!\left(\nabla_\theta \frac{1}{|B|} \sum_{i \in B} \ell(f_\theta(x_i), y_i)\right) $$
Symbols: $\theta$ — all trainable weights; $\eta$ — the learning rate; $\ell$ — the per-sample loss; $f_\theta$ — the model; $|B|$ — batch size; $A$ — the optimizer's transformation of the raw gradient (identity for SGD; bias-corrected first/second moment scaling for Adam, Kingma and Ba 2015).
Binary crossentropy for label $y \in {0,1}$ and prediction $p$: $\ell = -y \log p - (1-y)\log(1-p)$. Keras computes it from logits internally when the last layer is sigmoid, folding sigmoid and log together for numerical stability (the log-sum-exp trick), which is why from_logits=True with no activation is exactly equivalent and slightly better conditioned.
Cost accounting
Per epoch: $\lceil N/|B| \rceil$ steps, each costing one forward and one backward pass — roughly 3× the FLOPs of inference-only. Memory is dominated by stored activations, $O(|B| \sum_l d_l)$, which is why the practical fix for out-of-memory is lowering batch size, not model size. Epoch loop overhead in Keras is amortised by compiling the train step into a graph — the mechanism is covered in tf.function and retracing.
The three-way contract
compile fixes the objective, fit performs stochastic optimisation, evaluate estimates generalisation. The statistical claim behind evaluate is that held-out loss is an unbiased estimate of expected loss under the data distribution — a claim that survives only if the test set never influenced any choice: not the weights, not the architecture, not the stopping epoch. Repeated evaluation during development erodes it (adaptive overfitting; Dwork et al. 2015, The reusable holdout).
References
- Kingma and Ba (2015), Adam: a method for stochastic optimization.
- Bottou (2010), Large-scale machine learning with stochastic gradient descent — why mini-batches beat full-batch.
- Dwork et al. (2015), The reusable holdout: preserving validity in adaptive data analysis.
What to learn next
- Callbacks — reacting to training while it runs, including stopping at the right time.
- Loss functions — choosing the definition of wrong.
- Overfitting and underfitting — why training accuracy lies.