Edge and On-device AI

TensorFlow Lite

TensorFlow Lite, now called LiteRT, converts a trained model into a small flat file that a few-megabyte runtime can execute on Android, iOS and microcontrollers.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. One tool, two names
  4. Why it exists
  5. How it works
  6. What the converter does while converting
  7. Where you have already used it
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

TensorFlow Lite turns a trained model into a small file. It also gives you a small program that runs that file on a phone.

The analogy you have already lived

Think about your kitchen at home, and the tiffin you carry to work.

The kitchen has the gas, the grinder, the pressure cooker, twenty spice jars and a sink. You need all of it to make the food.

The tiffin has the food. It has none of the equipment, because eating does not need the equipment.

TensorFlow is the kitchen. A .tflite file is the tiffin.

One tool, two names

Google renamed it LiteRT in 2024. The files are still called .tflite and everything works the same way. You will meet both names, and they mean the same thing.

Why it exists

Training a model needs machinery: automatic differentiation, optimisers, data pipelines, checkpointing. That machinery is enormous. TensorFlow takes over a gigabyte of space once installed.

Running a trained model needs none of it. The weights are already decided. Something has to read the file and do the arithmetic in order.

So the tooling was split in two. A converter runs once on your laptop. A runtime of a few megabytes ships to the device.

How it works

  your laptop                          the phone
  ------------------------             ---------------------
  train the model      ─┐
  convert it           ─┤──►  model.tflite  ──►  tiny runtime  ──►  answers
  (needs the big tool)  ┘      (a few hundred KB)   (a few MB)

The conversion happens once, at build time. The phone never converts anything.

What the converter does while converting

It does not only change the format. It makes the model smaller and faster on the way through.

  • It folds away work that can be done in advance.
  • It joins neighbouring steps into single combined steps.
  • It can quantise the model, storing each number in one byte instead of four. That is covered in quantisation in practice.

That last one is the big win, and it is one line of code.

Where you have already used it

  • Almost every Android app that recognises something in a photo.
  • Google Lens, Live Caption and Android's on-device speech.
  • Doorbell cameras and smart displays.
  • Tiny boards with no operating system at all, through LiteRT for Microcontrollers, in models measured in kilobytes.

The honest part

Not everything converts. TensorFlow has thousands of operations, and LiteRT supports a few hundred of them.

When your model uses something unsupported, one of two things happens. The converter fails with a long error message. Or it pulls in a chunk of full TensorFlow, which makes the app much larger.

The second honest thing: the converter needs the full, large TensorFlow install. So this route is heavy to set up even though what it produces is light. If your training already uses PyTorch, ONNX is the shorter path.

Remember this

  • Convert once on a laptop, run everywhere on a small runtime.
  • The converter also optimises and can quantise the model.
  • Some operations do not convert, and finding out early saves days.

What to learn next

  • Core ML — the equivalent route on Apple devices.
  • ONNX — the format to use when your training is in PyTorch.
  • Power and thermal limits — what happens after the model is on the phone.

Developer — Code and libraries.

Setup

Two installs, and the difference between them is the whole point of this lesson.

bash
# on your laptop, to convert. Large.
pip install tensorflow scikit-learn

# on the device, to run. Small.
pip install ai-edge-litert

Measured on one machine: tensorflow occupies about 1400 MB installed, ai_edge_litert about 49 MB. On Android the runtime is a few megabytes, because it is a native library rather than a Python package.

Step 1: train and convert, on the laptop

convert.py
import os, numpy as np, tensorflow as tf
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

tf.random.set_seed(0); np.random.seed(0)
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split((X / 16.0).astype("float32"), y,
                                      test_size=0.3, random_state=0, stratify=y)

model = tf.keras.Sequential([
    tf.keras.layers.Input(shape=(64,)),
    tf.keras.layers.Dense(256, activation="relu"),
    tf.keras.layers.Dense(256, activation="relu"),
    tf.keras.layers.Dense(10),
])
model.compile(optimizer="adam",
              loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
              metrics=["accuracy"])
model.fit(Xtr, ytr, epochs=30, batch_size=64, verbose=0)
print("keras parameters :", model.count_params())
print("keras accuracy   : %.4f" % model.evaluate(Xte, yte, verbose=0)[1])

model.export("saved_model", verbose=False)     # the converter reads a SavedModel, not a .keras file

def convert(name, optimize=False, representative=None):
    c = tf.lite.TFLiteConverter.from_saved_model("saved_model")
    if optimize:
        c.optimizations = [tf.lite.Optimize.DEFAULT]      # this alone gives 8-bit weights
    if representative is not None:
        c.representative_dataset = representative
        c.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
        c.inference_input_type = tf.int8                  # integers all the way in and out
        c.inference_output_type = tf.int8
    open(name, "wb").write(c.convert())
    print("%-16s %6.1f KB" % (name, os.path.getsize(name) / 1024))

def sample_inputs():
    for i in range(100):        # 100 real inputs, so the converter can measure activation ranges
        yield [Xtr[i:i + 1]]

convert("float.tflite")
convert("dynamic_int8.tflite", optimize=True)
convert("full_int8.tflite", optimize=True, representative=sample_inputs)
Output
keras parameters : 85002
keras accuracy   : 0.9759
float.tflite      333.9 KB
dynamic_int8.tflite   92.6 KB
full_int8.tflite   98.9 KB

float.tflite at 333.9 KB reproduces exactly, because its size depends only on the shape of the network. The other three numbers do not. TensorFlow's CPU kernels are not fully deterministic, so the Keras accuracy moves by two or three tenths of a point between runs — yours will be near 0.976, not equal to it. The two quantised sizes ride on those weights, because the converter stores a scale and a zero point derived from them, so expect roughly 92–93 KB and 99–100 KB rather than these exact figures. Checked on TensorFlow 2.21; a different release can shift them by a few hundred bytes as well.

Step 2: run it, with only the small package

This script never imports TensorFlow. It is what you would port to a phone.

run.py
import os, time, numpy as np
from ai_edge_litert.interpreter import Interpreter
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
_, Xte, _, yte = train_test_split((X / 16.0).astype("float32"), y,
                                  test_size=0.3, random_state=0, stratify=y)

def evaluate(path):
    interp = Interpreter(model_path=path, num_threads=1)
    interp.allocate_tensors()                    # reserve every buffer up front, once
    inp, out = interp.get_input_details()[0], interp.get_output_details()[0]
    scale, zero = inp["quantization"]            # (0.0, 0) when the input is float

    def predict(row):
        v = row[None, :]
        if inp["dtype"] == np.int8:              # a full-integer model wants int8 in and out
            v = np.clip(np.round(v / scale + zero), -128, 127).astype(np.int8)
        interp.set_tensor(inp["index"], v)
        interp.invoke()
        return int(interp.get_tensor(out["index"])[0].argmax())

    preds = [predict(r) for r in Xte]
    for _ in range(50):
        predict(Xte[0])
    t0 = time.perf_counter()
    for _ in range(500):
        predict(Xte[0])
    ms = (time.perf_counter() - t0) / 500 * 1000
    return np.mean(np.array(preds) == yte), np.dtype(inp["dtype"]).name, ms

print("file                  size    input      accuracy   ms/image")
for path in ["float.tflite", "dynamic_int8.tflite", "full_int8.tflite"]:
    acc, dt, ms = evaluate(path)
    print("%-19s %6.1f KB  %-9s  %.4f     %.3f" % (path, os.path.getsize(path) / 1024, dt, acc, ms))
Output
file                  size    input      accuracy   ms/image
float.tflite         333.9 KB  float32    0.9759     0.006
dynamic_int8.tflite   92.6 KB  float32    0.9759     0.004
full_int8.tflite      98.9 KB  int8       0.9759     0.007

You may also see INFO: Created TensorFlow Lite XNNPACK delegate for CPU. on stderr. XNNPACK is the optimised CPU kernel library, and that line means it is being used.

Reading the three rows honestly

Dynamic int8 is the sweet spot here. 3.6 times smaller than float, identical accuracy, and faster. If you take one thing from this lesson, it is that optimizations = [tf.lite.Optimize.DEFAULT] is one line and usually free.

Full int8 is larger than dynamic int8. 98.9 KB against 92.6 KB. The full-integer model carries extra quantisation parameters and explicit conversion nodes at the boundaries. It is also slightly slower here, because converting the input and output costs more than the integer arithmetic saves on a model this small.

So why would anyone use it? Because it is the only one of the three that runs on hardware with no floating-point unit at all — an NPU, a DSP, an Edge TPU, a microcontroller. On such hardware the full-integer model is not slower, it is the only option that works.

All three accuracies are identical. Your three numbers will sit a fraction of a point away from these, because they inherit the non-deterministic training run above — what should reproduce is that the three agree with each other. On this task, quantisation cost nothing measurable. Do not carry that conclusion to your own model without checking; see the bit-width cliff in quantisation in practice.

The representative dataset

sample_inputs() is the piece people get wrong. Full-integer quantisation needs to know the range of every activation inside the network, and the only way to learn that is to run real data through and watch.

I tested four different representative sets on the same trained model, converting each to a full-integer .tflite:

Representative setInput scale chosenTest accuracy
100 real training images0.003920.9796
100 uniform random vectors in [0, 1)0.003920.9796
100 Gaussian vectors, spread five times wider0.138320.9796
100 vectors of all zeros0.000000.1019

Read that carefully, because it is not the result people expect.

Uniform noise did no damage at all. It happens to cover the same range as the real pixels, so the converter picked the same scale. The rule is not "the data must be realistic". The rule is the data must cover the range you will actually see.

Even badly over-wide noise survived here, at a much coarser input scale, because this task has enormous margin. On a harder task it would not.

All zeros destroyed the model. Every range collapsed to a point, the input scale became zero, and accuracy fell to 0.1019 — random guessing among ten classes.

Practical rules that follow:

  • Use real inputs from your training or validation set. It is the only choice guaranteed to cover the right range.
  • 100 to 500 samples is usually enough. More rarely helps.
  • Cover the variety you expect. If half your users photograph in dim light, dim images belong in there — otherwise dim inputs get clipped at run time and nothing warns you.
  • A degenerate set, such as one repeated input, is the failure that actually happens in practice.

Common mistakes

Passing a .keras file to the converter. It reads a SavedModel directory. model.export("saved_model") produces one; model.save("model.keras") does not.

Forgetting allocate_tensors(). The interpreter raises before any inference happens.

Feeding float data to a full-integer model. The interpreter expects int8. The scaling line in predict is the conversion, and getting the sign of zero wrong yields confident nonsense rather than an error.

Assuming int8 always runs faster. Measured above, it did not, on a small model on a laptop CPU. Measure on the device you are shipping to.

Not checking operator support until the end. Convert a stub of your architecture on day one. Finding out in week six that one layer is unsupported is an expensive way to learn it.

Try it yourself

Change sample_inputs to yield np.zeros((1, 64), dtype="float32") and rerun both scripts. The file size does not change by a single byte, the conversion prints no warning, and accuracy falls to about 0.10. Then print inp["quantization"] in run.py and see the zero scale that caused it. A silent, byte-identical, completely broken model is the failure mode this whole step exists to prevent.

What to learn next

  • Core ML — the equivalent route on Apple devices.
  • ONNX — the format to use when your training is in PyTorch.
  • Power and thermal limits — what happens after the model is on the phone.

Researcher — Mathematics and papers.

The format

A .tflite file is a FlatBuffer, not a protobuf. That choice is load-bearing on a device: FlatBuffers can be read in place with no parsing and no deserialisation allocation, so the model can be memory-mapped straight from flash. On a phone this removes both a copy and a chunk of peak memory at startup.

The schema holds a list of subgraphs; each subgraph has tensors, an operator list, and explicit input and output indices. Weights are stored as buffers referenced by index, which is what allows the same buffer to back several tensors.

The 2024 rename to LiteRT accompanied a widening of scope: the runtime now targets PyTorch models via ai-edge-torch as well as TensorFlow ones, and the Python runtime package split out as ai-edge-litert. The file format and the .tflite extension were kept.

Delegates

The runtime executes operators on the CPU by default and offers delegates that claim subgraphs for accelerators:

DelegateTargetNote
XNNPACKCPU, all platformsDefault; highly optimised float and quantised kernels
GPUOpenGL ES / OpenCL / MetalGood for large float convolutional models
NNAPIAndroid acceleratorsDeprecated from Android 15 in favour of vendor delegates
Hexagon, QNNQualcomm DSP and NPUInteger only
Edge TPUCoralFull-integer models only, and a separate compile step

Delegation is partial and silent. A delegate claims the largest contiguous subgraphs it supports; the rest stays on the CPU. Each hand-off between CPU and accelerator costs a synchronisation and often a memory copy, so a model split into many alternating segments can run slower than the pure CPU path. The diagnostic is the delegate's own logging plus a per-operator profile, not the end-to-end number.

Quantisation modes, precisely

ModeWeightsActivationsCalibration dataRuns on integer-only hardware
Float16fp16fp32noneno
Dynamic rangeint8fp32, quantised per callnoneno
Full integerint8int8, fixed rangesrequiredyes
Int16 activations, int8 weightsint8int16requiredon supporting hardware

The 16x8 mode exists because int8 activations are the binding constraint for models with wide dynamic range — audio front ends and some recurrent networks. It roughly doubles activation memory and typically recovers most of the accuracy lost by int8 activations.

Full-integer conversion computes per-tensor activation ranges $[a_{\min}, a_{\max}]$ from the representative set. Because these are fixed at conversion time, deployment inputs that fall outside the calibrated range are clipped. This is a distribution-shift failure mode with no runtime warning attached, and it is why the representative set must resemble deployment data rather than only being real.

Quantisation-aware training

When post-training quantisation loses too much, the TensorFlow Model Optimization Toolkit inserts fake-quantise operations during training so the network learns weights robust to the rounding. The forward pass simulates quantisation; the backward pass uses the straight-through estimator (Bengio et al., 2013) to push a gradient through a function whose true derivative is zero almost everywhere.

The empirical ordering, consistent across the literature: float32 ≥ QAT int8 ≥ post-training int8 ≥ post-training int4. The gap between the middle two is small for convolutional networks with batch normalisation, and large for compact architectures such as MobileNetV1 with depthwise convolutions, where per-channel ranges vary by orders of magnitude across channels. Per-channel weight quantisation, now the default, closed most of that particular gap.

Microcontrollers

LiteRT for Microcontrollers is a separate C++ runtime, roughly 16 KB of core, with no dynamic memory allocation, no operating system dependency and no file system. The model is compiled into the binary as a C array. The developer supplies a fixed tensor arena, and an OpResolver listing exactly the operators to link, so unused kernels never enter the binary.

The constraints that follow are absolute rather than advisory: no dynamic shapes, no operator not present at link time, and a peak activation footprint that must be known statically. Models in this class are typically 20–200 KB, and wake-word and anomaly-detection tasks dominate the practical deployments.

Reading

What to learn next

  • Core ML — the equivalent route on Apple devices.
  • ONNX — the format to use when your training is in PyTorch.
  • Power and thermal limits — what happens after the model is on the phone.

What to learn next

These follow on from what you just read.

  • Edge and On-device AI

    Core ML

    Core ML is Apple's on-device model format, which automatically routes each part of a model to the CPU, GPU or Neural Engine — and only runs on Apple hardware.

  • Edge and On-device AI

    Running models in the browser

    A browser can run a trained model with no install and no server, using WebAssembly or WebGPU — at the cost of a few megabytes of runtime on the first visit.

  • Edge and On-device AI

    AI on a Raspberry Pi

    A Raspberry Pi runs real models on four ARM cores with no GPU, so the work is choosing a small model, quantising it, and being honest about frames per second.