Saving and loading Keras models
One .keras file carries the architecture, the learned weights and the optimizer state — load it back and continue exactly where you stopped.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Saving a Keras model writes everything it knows into one file, so you can close your laptop and load the same trained model back tomorrow.
Think of a chess game you must pause halfway. You photograph the board, note whose turn it is, and pack up. Next week you rebuild the exact position from the photo and play on — no moves lost.
A trained model is a paused game. The weights — all the numbers learning has tuned — are the position of the pieces. The architecture — the layer structure — is the board itself. The optimizer state — the trainer's memory of recent progress — is whose turn it is and the plans in flight. The .keras file stores all three.
Why it exists
Training costs hours, sometimes weeks. Without saving, every prediction would require retraining from zero — like replaying a full chess game from move one to make the next move.
Saving splits life into two halves. Train once, on a big machine, slowly. Then load and predict many times, anywhere, instantly. Every model serving real users lives on the loading side of this split.
How it works
save: architecture + weights + optimizer state ──▶ one .keras file
load: one .keras file ──▶ the same model,
ready to predict — or keep training
lighter option: weights only ──▶ needs the matching architecture in codeThe weights-only option stores the piece positions without the board. It is smaller and flexible. But loading requires you to rebuild the identical architecture in code first: a photo of pieces is useless on the wrong board.
A real example you have seen
Every app update that says "improved camera" or "smarter keyboard" is shipping you a saved model file. The training happened months earlier on company servers. Your phone only ever loads the frozen result. Nobody retrains anything on your device when you install an update.
Remember this
.kerasfile = architecture + weights + optimizer state, one file, resume anything.- Weights-only = smaller file, but you must rebuild the same architecture in code.
- Train once, save, load many times — this split is how all real systems ship.
What to learn next
- Stopping TensorFlow taking the whole GPU — the last practical hurdle before real training runs.
- Callbacks — ModelCheckpoint automates the saving you now understand.
- TensorFlow Lite — shrinking a saved model to run on phones.
Developer — Code and libraries.
Setup
pip install tensorflow numpyOutputs verified with TensorFlow 2.21 (Keras 3.12), CPU.
Save, load, verify
Never trust a save/load cycle you have not verified with a prediction comparison. The habit costs three lines.
import numpy as np
import keras
from keras import layers
keras.utils.set_random_seed(0)
model = keras.Sequential([
keras.Input(shape=(3,)),
layers.Dense(8, activation="relu"),
layers.Dense(2, activation="softmax"),
])
model.compile(optimizer="adam", loss="sparse_categorical_crossentropy")
x = np.random.rand(4, 3).astype("float32")
before = model.predict(x, verbose=0)
model.save("sensor_model.keras") # ONE file, everything inside
loaded = keras.models.load_model("sensor_model.keras")
after = loaded.predict(x, verbose=0)
print("same predictions:", np.allclose(before, after))
print("optimizer restored:", type(loaded.optimizer).__name__)
model.save_weights("sensor.weights.h5") # numbers only, no architecture
twin = keras.models.clone_model(model)
twin.load_weights("sensor.weights.h5")
print("weights-only twin matches:", np.allclose(before, twin.predict(x, verbose=0)))same predictions: True optimizer restored: Adam weights-only twin matches: True
The walkthrough
model.save("sensor_model.keras") writes a zip archive holding the architecture as JSON, the weights, and the compiled optimizer with its accumulated state. Because optimizer state travels too, loaded.fit(...) continues training as if never interrupted — the resume-from-checkpoint workflow.
The extension is enforced. Keras 3 refuses a bare filename: ValueError: Invalid filepath extension for saving. Please add either a .keras extension for the native Keras format (recommended).... The older HDF5 format (.h5) still loads for legacy files; write new files as .keras.
save_weights requires .weights.h5 as its extension, and stores only the numbers. Loading them into clone_model(model) — a fresh copy of the architecture — reproduces the predictions. This path shines when the architecture lives in code (it usually does) and you want small, frequent checkpoints.
When to use which: .keras for "give me back my model" — resuming, sharing, archiving. Weights-only for checkpoints inside a training script you control. For serving without Python, model.export("dirname") writes a TensorFlow SavedModel for TF Serving — a deployment topic beyond this lesson; the wider decisions live in TensorFlow Lite and its serving cousins.
Common mistakes
Skipping the np.allclose check. Most save/load bugs — wrong file, custom object missing, architecture drift — surface instantly under this check and stay hidden without it. Save, load, compare, every time you build a new save path.
Custom pieces without registration. A model containing a custom layer or loss reloads only if Keras can rebuild those objects. Decorate the class with @keras.saving.register_keras_serializable() and give it a get_config(); otherwise load_model raises Unknown layer errors. Alternative: pass custom_objects={"ScaledDense": ScaledDense} at load time.
Assuming pickle works. pickle.dump(model) may appear to succeed and produces files that break across versions and machines. The .keras format exists precisely to be the stable, portable path.
Loading old files across big version jumps. Keras 2 .h5 files usually load in Keras 3; exotic layers and custom code sometimes do not. Version-pin serving environments, and re-export archives when you upgrade — silent weight mismatches are worse than crashes.
Try it yourself
Train the model above for 5 epochs on random labels, save, reload, and call loaded.fit for 5 more. Compare the loss curve against training 10 epochs straight through in one process — near-identical, because optimizer state travelled with the file. Then repeat using only save_weights and watch the resumed loss jump at the boundary: the fresh optimizer lost its momentum memory.
What to learn next
- Stopping TensorFlow taking the whole GPU — the last practical hurdle before real training runs.
- Callbacks — ModelCheckpoint automates the saving you now understand.
- TensorFlow Lite — shrinking a saved model to run on phones.
Researcher — Mathematics and papers.
What must be serialised, exactly
A resumable training state is the triple $(A, \theta_t, s_t)$: the architecture $A$ (a program, serialised as a config graph of layer classes and constructor arguments), the parameters $\theta_t \in \mathbb{R}^p$ at step $t$, and the optimizer state $s_t$. For Adam, $s_t = (m_t, v_t, t)$ — first-moment and second-moment accumulators plus the step counter — which is another $2p$ floats; dropping it and restarting sets $m=v=0$, and bias correction makes the first resumed steps behave like a warm restart with a briefly inflated effective learning rate. That is the loss bump in the try-it-yourself exercise, and why serious checkpointing always includes $s_t$.
Weight files map cleanly onto a name→tensor dictionary (HDF5 groups keyed by layer path). The failure mode is name drift: rebuild the architecture with layers created in a different order and auto-generated names (dense_3) no longer align. Keras matches by structural path, so clone_model is safe; hand-rebuilding "the same" model is where mismatches breed.
Formats and their trust boundaries
The .keras archive holds config.json (architecture), model.weights.h5, and metadata.json. Rebuilding from config executes the constructors of whatever classes the config names — which is why deserialisation of custom classes is gated behind explicit registration: loading a model file is code execution by proxy, the same threat model as pickle, narrowed to registered constructors. Keras 3 additionally ships "safe mode" (default on) refusing arbitrary lambda deserialisation.
SavedModel (model.export) is a different artefact class: it stores the traced graph (concrete functions, see tf.function) rather than Python-rebuildable config — executable without the defining code, hence its role as the serving interchange format, and its weakness for further training.
Cost: all formats are $O(p)$ bytes, dominated by weights — float32 is 4 bytes per parameter; the file for a 100M-parameter model is ~400 MB, and ~1.2 GB with Adam state. Post-training quantisation to int8 quarters the weight payload, one branch of the model compression toolbox.
References
- Kingma and Ba (2015), Adam: a method for stochastic optimization — defines the state being checkpointed.
- The HDF5 Group (1997–), Hierarchical Data Format 5 — the container underlying weight files.
- Olston et al. (2017), TensorFlow-Serving: flexible, high-performance ML serving — the consumer of exported SavedModels.
What to learn next
- Stopping TensorFlow taking the whole GPU — the last practical hurdle before real training runs.
- Callbacks — ModelCheckpoint automates the saving you now understand.
- TensorFlow Lite — shrinking a saved model to run on phones.