TensorFlow Lite
TensorFlow Lite, now called LiteRT, converts a trained model into a small flat file that a few-megabyte runtime can execute on Android, iOS and microcontrollers.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
TensorFlow Lite turns a trained model into a small file. It also gives you a small program that runs that file on a phone.
The analogy you have already lived
Think about your kitchen at home, and the tiffin you carry to work.
The kitchen has the gas, the grinder, the pressure cooker, twenty spice jars and a sink. You need all of it to make the food.
The tiffin has the food. It has none of the equipment, because eating does not need the equipment.
TensorFlow is the kitchen. A .tflite file is the tiffin.
One tool, two names
Google renamed it LiteRT in 2024. The files are still called .tflite and everything works the same way. You will meet both names, and they mean the same thing.
Why it exists
Training a model needs machinery: automatic differentiation, optimisers, data pipelines, checkpointing. That machinery is enormous. TensorFlow takes over a gigabyte of space once installed.
Running a trained model needs none of it. The weights are already decided. Something has to read the file and do the arithmetic in order.
So the tooling was split in two. A converter runs once on your laptop. A runtime of a few megabytes ships to the device.
How it works
your laptop the phone
------------------------ ---------------------
train the model ─┐
convert it ─┤──► model.tflite ──► tiny runtime ──► answers
(needs the big tool) ┘ (a few hundred KB) (a few MB)The conversion happens once, at build time. The phone never converts anything.
What the converter does while converting
It does not only change the format. It makes the model smaller and faster on the way through.
- It folds away work that can be done in advance.
- It joins neighbouring steps into single combined steps.
- It can quantise the model, storing each number in one byte instead of four. That is covered in quantisation in practice.
That last one is the big win, and it is one line of code.
Where you have already used it
- Almost every Android app that recognises something in a photo.
- Google Lens, Live Caption and Android's on-device speech.
- Doorbell cameras and smart displays.
- Tiny boards with no operating system at all, through LiteRT for Microcontrollers, in models measured in kilobytes.
The honest part
Not everything converts. TensorFlow has thousands of operations, and LiteRT supports a few hundred of them.
When your model uses something unsupported, one of two things happens. The converter fails with a long error message. Or it pulls in a chunk of full TensorFlow, which makes the app much larger.
The second honest thing: the converter needs the full, large TensorFlow install. So this route is heavy to set up even though what it produces is light. If your training already uses PyTorch, ONNX is the shorter path.
Remember this
- Convert once on a laptop, run everywhere on a small runtime.
- The converter also optimises and can quantise the model.
- Some operations do not convert, and finding out early saves days.
What to learn next
- Core ML — the equivalent route on Apple devices.
- ONNX — the format to use when your training is in PyTorch.
- Power and thermal limits — what happens after the model is on the phone.
Developer — Code and libraries.
Setup
Two installs, and the difference between them is the whole point of this lesson.
# on your laptop, to convert. Large.
pip install tensorflow scikit-learn
# on the device, to run. Small.
pip install ai-edge-litertMeasured on one machine: tensorflow occupies about 1400 MB installed, ai_edge_litert about 49 MB. On Android the runtime is a few megabytes, because it is a native library rather than a Python package.
Step 1: train and convert, on the laptop
import os, numpy as np, tensorflow as tf
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
tf.random.set_seed(0); np.random.seed(0)
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split((X / 16.0).astype("float32"), y,
test_size=0.3, random_state=0, stratify=y)
model = tf.keras.Sequential([
tf.keras.layers.Input(shape=(64,)),
tf.keras.layers.Dense(256, activation="relu"),
tf.keras.layers.Dense(256, activation="relu"),
tf.keras.layers.Dense(10),
])
model.compile(optimizer="adam",
loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=["accuracy"])
model.fit(Xtr, ytr, epochs=30, batch_size=64, verbose=0)
print("keras parameters :", model.count_params())
print("keras accuracy : %.4f" % model.evaluate(Xte, yte, verbose=0)[1])
model.export("saved_model", verbose=False) # the converter reads a SavedModel, not a .keras file
def convert(name, optimize=False, representative=None):
c = tf.lite.TFLiteConverter.from_saved_model("saved_model")
if optimize:
c.optimizations = [tf.lite.Optimize.DEFAULT] # this alone gives 8-bit weights
if representative is not None:
c.representative_dataset = representative
c.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
c.inference_input_type = tf.int8 # integers all the way in and out
c.inference_output_type = tf.int8
open(name, "wb").write(c.convert())
print("%-16s %6.1f KB" % (name, os.path.getsize(name) / 1024))
def sample_inputs():
for i in range(100): # 100 real inputs, so the converter can measure activation ranges
yield [Xtr[i:i + 1]]
convert("float.tflite")
convert("dynamic_int8.tflite", optimize=True)
convert("full_int8.tflite", optimize=True, representative=sample_inputs)keras parameters : 85002 keras accuracy : 0.9759 float.tflite 333.9 KB dynamic_int8.tflite 92.6 KB full_int8.tflite 98.9 KB
float.tflite at 333.9 KB reproduces exactly, because its size depends only on the shape of the network. The other three numbers do not. TensorFlow's CPU kernels are not fully deterministic, so the Keras accuracy moves by two or three tenths of a point between runs — yours will be near 0.976, not equal to it. The two quantised sizes ride on those weights, because the converter stores a scale and a zero point derived from them, so expect roughly 92–93 KB and 99–100 KB rather than these exact figures. Checked on TensorFlow 2.21; a different release can shift them by a few hundred bytes as well.
Step 2: run it, with only the small package
This script never imports TensorFlow. It is what you would port to a phone.
import os, time, numpy as np
from ai_edge_litert.interpreter import Interpreter
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True)
_, Xte, _, yte = train_test_split((X / 16.0).astype("float32"), y,
test_size=0.3, random_state=0, stratify=y)
def evaluate(path):
interp = Interpreter(model_path=path, num_threads=1)
interp.allocate_tensors() # reserve every buffer up front, once
inp, out = interp.get_input_details()[0], interp.get_output_details()[0]
scale, zero = inp["quantization"] # (0.0, 0) when the input is float
def predict(row):
v = row[None, :]
if inp["dtype"] == np.int8: # a full-integer model wants int8 in and out
v = np.clip(np.round(v / scale + zero), -128, 127).astype(np.int8)
interp.set_tensor(inp["index"], v)
interp.invoke()
return int(interp.get_tensor(out["index"])[0].argmax())
preds = [predict(r) for r in Xte]
for _ in range(50):
predict(Xte[0])
t0 = time.perf_counter()
for _ in range(500):
predict(Xte[0])
ms = (time.perf_counter() - t0) / 500 * 1000
return np.mean(np.array(preds) == yte), np.dtype(inp["dtype"]).name, ms
print("file size input accuracy ms/image")
for path in ["float.tflite", "dynamic_int8.tflite", "full_int8.tflite"]:
acc, dt, ms = evaluate(path)
print("%-19s %6.1f KB %-9s %.4f %.3f" % (path, os.path.getsize(path) / 1024, dt, acc, ms))file size input accuracy ms/image float.tflite 333.9 KB float32 0.9759 0.006 dynamic_int8.tflite 92.6 KB float32 0.9759 0.004 full_int8.tflite 98.9 KB int8 0.9759 0.007
You may also see INFO: Created TensorFlow Lite XNNPACK delegate for CPU. on stderr. XNNPACK is the optimised CPU kernel library, and that line means it is being used.
Reading the three rows honestly
Dynamic int8 is the sweet spot here. 3.6 times smaller than float, identical accuracy, and faster. If you take one thing from this lesson, it is that optimizations = [tf.lite.Optimize.DEFAULT] is one line and usually free.
Full int8 is larger than dynamic int8. 98.9 KB against 92.6 KB. The full-integer model carries extra quantisation parameters and explicit conversion nodes at the boundaries. It is also slightly slower here, because converting the input and output costs more than the integer arithmetic saves on a model this small.
So why would anyone use it? Because it is the only one of the three that runs on hardware with no floating-point unit at all — an NPU, a DSP, an Edge TPU, a microcontroller. On such hardware the full-integer model is not slower, it is the only option that works.
All three accuracies are identical. Your three numbers will sit a fraction of a point away from these, because they inherit the non-deterministic training run above — what should reproduce is that the three agree with each other. On this task, quantisation cost nothing measurable. Do not carry that conclusion to your own model without checking; see the bit-width cliff in quantisation in practice.
The representative dataset
sample_inputs() is the piece people get wrong. Full-integer quantisation needs to know the range of every activation inside the network, and the only way to learn that is to run real data through and watch.
I tested four different representative sets on the same trained model, converting each to a full-integer .tflite:
| Representative set | Input scale chosen | Test accuracy |
|---|---|---|
| 100 real training images | 0.00392 | 0.9796 |
| 100 uniform random vectors in [0, 1) | 0.00392 | 0.9796 |
| 100 Gaussian vectors, spread five times wider | 0.13832 | 0.9796 |
| 100 vectors of all zeros | 0.00000 | 0.1019 |
Read that carefully, because it is not the result people expect.
Uniform noise did no damage at all. It happens to cover the same range as the real pixels, so the converter picked the same scale. The rule is not "the data must be realistic". The rule is the data must cover the range you will actually see.
Even badly over-wide noise survived here, at a much coarser input scale, because this task has enormous margin. On a harder task it would not.
All zeros destroyed the model. Every range collapsed to a point, the input scale became zero, and accuracy fell to 0.1019 — random guessing among ten classes.
Practical rules that follow:
- Use real inputs from your training or validation set. It is the only choice guaranteed to cover the right range.
- 100 to 500 samples is usually enough. More rarely helps.
- Cover the variety you expect. If half your users photograph in dim light, dim images belong in there — otherwise dim inputs get clipped at run time and nothing warns you.
- A degenerate set, such as one repeated input, is the failure that actually happens in practice.
Common mistakes
Passing a .keras file to the converter. It reads a SavedModel directory. model.export("saved_model") produces one; model.save("model.keras") does not.
Forgetting allocate_tensors(). The interpreter raises before any inference happens.
Feeding float data to a full-integer model. The interpreter expects int8. The scaling line in predict is the conversion, and getting the sign of zero wrong yields confident nonsense rather than an error.
Assuming int8 always runs faster. Measured above, it did not, on a small model on a laptop CPU. Measure on the device you are shipping to.
Not checking operator support until the end. Convert a stub of your architecture on day one. Finding out in week six that one layer is unsupported is an expensive way to learn it.
Try it yourself
Change sample_inputs to yield np.zeros((1, 64), dtype="float32") and rerun both scripts. The file size does not change by a single byte, the conversion prints no warning, and accuracy falls to about 0.10. Then print inp["quantization"] in run.py and see the zero scale that caused it. A silent, byte-identical, completely broken model is the failure mode this whole step exists to prevent.
What to learn next
- Core ML — the equivalent route on Apple devices.
- ONNX — the format to use when your training is in PyTorch.
- Power and thermal limits — what happens after the model is on the phone.
Researcher — Mathematics and papers.
The format
A .tflite file is a FlatBuffer, not a protobuf. That choice is load-bearing on a device: FlatBuffers can be read in place with no parsing and no deserialisation allocation, so the model can be memory-mapped straight from flash. On a phone this removes both a copy and a chunk of peak memory at startup.
The schema holds a list of subgraphs; each subgraph has tensors, an operator list, and explicit input and output indices. Weights are stored as buffers referenced by index, which is what allows the same buffer to back several tensors.
The 2024 rename to LiteRT accompanied a widening of scope: the runtime now targets PyTorch models via ai-edge-torch as well as TensorFlow ones, and the Python runtime package split out as ai-edge-litert. The file format and the .tflite extension were kept.
Delegates
The runtime executes operators on the CPU by default and offers delegates that claim subgraphs for accelerators:
| Delegate | Target | Note |
|---|---|---|
| XNNPACK | CPU, all platforms | Default; highly optimised float and quantised kernels |
| GPU | OpenGL ES / OpenCL / Metal | Good for large float convolutional models |
| NNAPI | Android accelerators | Deprecated from Android 15 in favour of vendor delegates |
| Hexagon, QNN | Qualcomm DSP and NPU | Integer only |
| Edge TPU | Coral | Full-integer models only, and a separate compile step |
Delegation is partial and silent. A delegate claims the largest contiguous subgraphs it supports; the rest stays on the CPU. Each hand-off between CPU and accelerator costs a synchronisation and often a memory copy, so a model split into many alternating segments can run slower than the pure CPU path. The diagnostic is the delegate's own logging plus a per-operator profile, not the end-to-end number.
Quantisation modes, precisely
| Mode | Weights | Activations | Calibration data | Runs on integer-only hardware |
|---|---|---|---|---|
| Float16 | fp16 | fp32 | none | no |
| Dynamic range | int8 | fp32, quantised per call | none | no |
| Full integer | int8 | int8, fixed ranges | required | yes |
| Int16 activations, int8 weights | int8 | int16 | required | on supporting hardware |
The 16x8 mode exists because int8 activations are the binding constraint for models with wide dynamic range — audio front ends and some recurrent networks. It roughly doubles activation memory and typically recovers most of the accuracy lost by int8 activations.
Full-integer conversion computes per-tensor activation ranges $[a_{\min}, a_{\max}]$ from the representative set. Because these are fixed at conversion time, deployment inputs that fall outside the calibrated range are clipped. This is a distribution-shift failure mode with no runtime warning attached, and it is why the representative set must resemble deployment data rather than only being real.
Quantisation-aware training
When post-training quantisation loses too much, the TensorFlow Model Optimization Toolkit inserts fake-quantise operations during training so the network learns weights robust to the rounding. The forward pass simulates quantisation; the backward pass uses the straight-through estimator (Bengio et al., 2013) to push a gradient through a function whose true derivative is zero almost everywhere.
The empirical ordering, consistent across the literature: float32 ≥ QAT int8 ≥ post-training int8 ≥ post-training int4. The gap between the middle two is small for convolutional networks with batch normalisation, and large for compact architectures such as MobileNetV1 with depthwise convolutions, where per-channel ranges vary by orders of magnitude across channels. Per-channel weight quantisation, now the default, closed most of that particular gap.
Microcontrollers
LiteRT for Microcontrollers is a separate C++ runtime, roughly 16 KB of core, with no dynamic memory allocation, no operating system dependency and no file system. The model is compiled into the binary as a C array. The developer supplies a fixed tensor arena, and an OpResolver listing exactly the operators to link, so unused kernels never enter the binary.
The constraints that follow are absolute rather than advisory: no dynamic shapes, no operator not present at link time, and a peak activation footprint that must be known statically. Models in this class are typically 20–200 KB, and wake-word and anomaly-detection tasks dominate the practical deployments.
Reading
- LiteRT documentation — ai.google.dev/edge/litert
- Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, CVPR 2018 — arxiv.org/abs/1712.05877
- David et al., TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems, MLSys 2021 — arxiv.org/abs/2010.08678
- Lai, Suda and Chandra, CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M, 2018 — arxiv.org/abs/1801.06601
- Banbury et al., MLPerf Tiny Benchmark, 2021 — arxiv.org/abs/2106.07597
What to learn next
- Core ML — the equivalent route on Apple devices.
- ONNX — the format to use when your training is in PyTorch.
- Power and thermal limits — what happens after the model is on the phone.