The MobileNet family
MobileNets split one convolution into two cheaper steps, one that filters each channel and one that mixes channels, which cuts the cost by about nine times.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
MobileNets split one convolution into two cheaper steps: one that looks at shape, and one that mixes channels.
Think about making a spice blend at home. The costly way is to grind every raw spice from scratch for every blend you want to make.
The sensible way is to grind each spice once and store the powders. Then measure out different proportions for each blend. Grinding happens once. Mixing is quick.
An ordinary convolution grinds and mixes in a single expensive step. A MobileNet grinds each channel separately, then mixes the ground channels with a cheap second step.
Why it exists
A ResNet needs a fast computer. That is fine in a data centre and impossible on a five thousand rupee phone.
Plenty of things must run on the device itself. Face unlock cannot wait for a network round trip. A camera preview must keep up with thirty frames a second. Anything private should never leave the handset at all.
So a team at Google asked a narrow question. What is the smallest network that still recognises things properly?
The two steps
ORDINARY: one filter looks at a patch AND at every channel at once
cost: big
SPLIT:
step 1 each channel gets its own small filter, no mixing (cheap)
step 2 a one-pixel filter mixes the channels together (cheap)
cost: about nine times smaller for a three-by-three filterThe two steps together do nearly the same job. The measurement further down shows a saving of 8.69 times on a typical layer. The arithmetic predicts exactly that number.
Nine times cheaper is not a small win. It is the difference between a model that runs on your phone and one that does not.
Two dials for size
MobileNet gave you two dials to turn.
The first dial thins every layer. Ask for seventy-five percent and every layer gets three quarters as many channels.
The second dial shrinks the input picture. A smaller picture means less work at every single layer.
Turn either dial and you trade accuracy for speed, smoothly, without redesigning anything.
Where you have already seen this
- Face unlock on a mid-range Android phone.
- Live background blur in a video call.
- Barcode and text scanning inside a camera app.
- Offline translation of a signboard held up to your camera.
The honest part
Fewer calculations does not always mean faster. This surprises people every time.
A MobileNet does far less arithmetic than a ResNet, and on some ordinary computers it is not much quicker. The reason is that its work is spread thinly. The machine spends its time fetching numbers from memory rather than multiplying them.
These networks are tuned for phone chips. On a desktop processor the advantage shrinks, and sometimes disappears. Measure on the device you will ship on, not on your laptop.
Remember this
- A depthwise separable convolution splits filtering from channel mixing.
- The saving is roughly nine times for a three-by-three filter.
- Fewer operations does not automatically mean lower latency, so measure on the target device.
What to learn next
- EfficientNet and compound scaling — how to grow a small network without guessing.
- Quantization in practice — the other half of making a model fit on a phone.
- Latency and throughput — why counting operations is not enough.
Developer — Code and libraries.
Setup
pip install "torch==2.5.1" "torchvision==0.20.1"Run against PyTorch 2.5.1 and torchvision 0.20.1 on CPU. Nothing is downloaded.
Split the convolution and count what it saves
import torch
import torch.nn as nn
p = lambda m: sum(q.numel() for q in m.parameters())
C_IN, C_OUT, K, HW = 128, 256, 3, 28
standard = nn.Conv2d(C_IN, C_OUT, K, padding=1, bias=False)
separable = nn.Sequential(
nn.Conv2d(C_IN, C_IN, K, padding=1, groups=C_IN, bias=False), # one filter per channel
nn.Conv2d(C_IN, C_OUT, 1, bias=False), # mixes the channels
)
x = torch.zeros(1, C_IN, HW, HW)
print("same input, same output shape")
print(" standard :", tuple(standard(x).shape), f"{p(standard):>9,} weights")
print(" separable:", tuple(separable(x).shape), f"{p(separable):>9,} weights")
print(f" ratio : {p(standard) / p(separable):.2f}x fewer")
predicted = 1 / C_OUT + 1 / (K * K)
print(f" formula 1/C_out + 1/k^2 = {predicted:.5f} -> {1/predicted:.2f}x, and measured is the same")
class InvertedResidual(nn.Module):
"""MobileNetV2: widen with 1x1, filter cheaply with depthwise, narrow with 1x1."""
def __init__(self, c, expand=6):
super().__init__()
hidden = c * expand
self.block = nn.Sequential(
nn.Conv2d(c, hidden, 1, bias=False), nn.BatchNorm2d(hidden), nn.ReLU6(),
nn.Conv2d(hidden, hidden, 3, padding=1, groups=hidden, bias=False),
nn.BatchNorm2d(hidden), nn.ReLU6(),
nn.Conv2d(hidden, c, 1, bias=False), nn.BatchNorm2d(c), # no activation here
)
def forward(self, x):
return x + self.block(x)
blk = InvertedResidual(64)
print(f"\ninverted residual, 64 channels, expansion 6: {p(blk):,} parameters")
print(" shape in/out:", tuple(torch.zeros(1, 64, 14, 14).shape),
tuple(blk(torch.zeros(1, 64, 14, 14)).shape))
from torchvision.models import mobilenet_v2, mobilenet_v3_small, mobilenet_v3_large, resnet50
print("\nwhole networks, 1000 ImageNet classes")
for name, fn in [("mobilenet_v3_small", mobilenet_v3_small), ("mobilenet_v2", mobilenet_v2),
("mobilenet_v3_large", mobilenet_v3_large), ("resnet50", resnet50)]:
m = fn(weights=None)
print(f" {name:20s} {p(m):>12,} parameters")
print("\nwidth multiplier: MobileNetV2 with thinner channels")
for w in (0.35, 0.5, 0.75, 1.0, 1.4):
m = mobilenet_v2(weights=None, width_mult=w)
print(f" width_mult={w:<5} {p(m):>12,} parameters")same input, same output shape standard : (1, 256, 28, 28) 294,912 weights separable: (1, 256, 28, 28) 33,920 weights ratio : 8.69x fewer formula 1/C_out + 1/k^2 = 0.11502 -> 8.69x, and measured is the same inverted residual, 64 channels, expansion 6: 54,272 parameters shape in/out: (1, 64, 14, 14) (1, 64, 14, 14) whole networks, 1000 ImageNet classes mobilenet_v3_small 2,542,856 parameters mobilenet_v2 3,504,872 parameters mobilenet_v3_large 5,483,032 parameters resnet50 25,557,032 parameters width multiplier: MobileNetV2 with thinner channels width_mult=0.35 1,677,128 parameters width_mult=0.5 1,968,680 parameters width_mult=0.75 2,636,424 parameters width_mult=1.0 3,504,872 parameters width_mult=1.4 6,108,776 parameters
Reading the output
groups=C_IN is the entire depthwise trick. With groups equal to the input channel count, each channel gets its own filter and no channel talks to any other. PyTorch has no separate DepthwiseConv2d layer; this argument is it.
8.69x measured, 8.69x predicted. The formula is 1/C_out + 1/k². With a 3x3 kernel the 1/9 term dominates unless the output is very narrow, which is why the saving is quoted as "about 9x" for 3x3 and "about 25x" for 5x5.
The inverted residual goes wide in the middle, not narrow. A ResNet bottleneck squeezes 256 channels to 64 and back. MobileNetV2 expands 64 to 384 and back. That inversion is the name. The expensive 3x3 runs at the wide point, but it is depthwise, so wide costs little.
There is no activation after the final 1x1. That is deliberate and is the "linear bottleneck" in the paper's title. A ReLU on a narrow layer destroys information it cannot recover, because anything negative becomes zero and the layer has too few channels to encode it elsewhere.
MobileNetV3-Small is ten times smaller than ResNet-50. 2.5 million parameters against 25.6 million, at 67.7% ImageNet top-1 against 76.1% for the original ResNet-50 weights. That is the trade, stated plainly.
The width multiplier does not scale parameters linearly. Halving the width from 1.0 to 0.5 takes 3.50 million down to 1.97 million, not to 1.75 million. Convolution weights scale with the square of width, but the final classifier scales linearly and the channel counts get rounded to multiples of eight.
Common mistakes
Forgetting groups and wondering where the saving went. nn.Conv2d(128, 128, 3, padding=1) is a dense convolution costing 147,456 weights. Adding groups=128 makes it 1,152.
Putting a ReLU after the projection. It undoes the linear-bottleneck design and costs accuracy. In V2 the last operation of a block is a batch norm, then the residual addition, and nothing else.
Applying depthwise convolution when the channel count is tiny. With 8 channels there is almost nothing to save and the layer becomes memory-bound. MobileNets use ordinary convolutions in the first layer for this reason.
Assuming small means fast. Measured on one desktop CPU at four threads, mobilenet_v3_small was about twice as fast as resnet18 on a 224x224 image. ResNet-18 does roughly 30 times more arithmetic. Depthwise layers have low arithmetic intensity, so memory traffic limits them rather than multiplication. Absolute timings move between machines and between runs, so measure your own.
Quantising depthwise layers per-tensor. Depthwise weight distributions vary sharply between channels, so a single scale factor for the whole tensor loses accuracy. Use per-channel quantisation, as covered in quantization in practice.
Try it yourself
Set K = 5 in the first block and predict the new ratio before running it. Then set C_OUT = 8 with K = 3 and watch the saving fall well below 9x, because the 1/C_out term stops being negligible.
What to learn next
- EfficientNet and compound scaling — how to grow a small network without guessing.
- Quantization in practice — the other half of making a model fit on a phone.
- Latency and throughput — why counting operations is not enough.
Researcher — Mathematics and papers.
The factorisation
A standard convolution maps $C_{\text{in}} \times H \times W$ to $C_{\text{out}} \times H \times W$ at a cost of
$$ C_{\text{in}} C_{\text{out}} k^2 H W $$
multiply-accumulates. The depthwise separable form splits it into a depthwise stage, $C_{\text{in}} k^2 H W$, and a pointwise stage, $C_{\text{in}} C_{\text{out}} H W$. The ratio is
$$ \frac{C_{\text{in}} k^2 + C_{\text{in}} C_{\text{out}}}{C_{\text{in}} C_{\text{out}} k^2} = \frac{1}{C_{\text{out}}} + \frac{1}{k^2}. $$
For $k = 3$ and $C_{\text{out}} \geq 256$ this is close to $1/9$. The factorisation is exact only in cost; it is a restriction of the hypothesis space, since a separable convolution cannot express an arbitrary dense kernel.
Howard et al. (2017), MobileNets, arxiv.org/abs/1704.04861, introduced the width multiplier $\alpha$ and the resolution multiplier $\rho$, which scale cost by roughly $\alpha^2$ and $\rho^2$. Both later became axes in EfficientNet's compound scaling.
The linear bottleneck argument
Sandler et al. (2018), MobileNetV2: Inverted Residuals and Linear Bottlenecks, arxiv.org/abs/1801.04381, argue that the information of interest lies on a low-dimensional manifold inside the activation space, and that ReLU preserves complete information about the input only when that input lies in a low-dimensional subspace of the ReLU's output space.
The consequence: applying ReLU in a narrow layer destroys information irreversibly, while applying it in a wide layer is close to harmless. So the block expands by a factor of 6, applies its non-linearities there, and projects back down with no activation at all. The residual connection joins the narrow ends, which is the inversion relative to ResNet and also keeps the tensors crossing block boundaries small, an important memory property on device.
ReLU6, $\min(\max(x,0),6)$, is used instead of ReLU because the bounded range is friendly to low-precision fixed-point arithmetic.
The versions
| Version | Year | What it added |
|---|---|---|
| V1 | 2017 | Depthwise separable convolutions; width and resolution multipliers |
| V2 | 2018 | Inverted residuals; linear bottlenecks; ReLU6 |
| V3 | 2019 | Platform-aware NAS plus NetAdapt; squeeze-and-excitation; hard-swish; hand-tuned head |
| V4 | 2024 | Universal Inverted Bottleneck; Mobile MQA attention; distillation recipe |
Howard et al. (2019), Searching for MobileNetV3, arxiv.org/abs/1905.02244, report V3-Large as 3.2% more accurate than V2 at 15% lower latency, and V3-Small as 4.6% more accurate at 5% lower latency, on Pixel phones. Hard-swish, $x \cdot \mathrm{ReLU6}(x+3)/6$, replaces swish because it avoids a sigmoid, which is expensive in fixed point.
Qin et al. (2024), MobileNetV4: Universal Models for the Mobile Ecosystem, arxiv.org/abs/2404.10518, introduce the Universal Inverted Bottleneck, a search space unifying the inverted bottleneck, the ConvNeXt block, the feed-forward block and a new ExtraDW variant. Mobile MQA, a multi-query attention variant, gives a reported 39% speed-up on mobile accelerators. MNv4-Hybrid-Large reaches 87% ImageNet-1K top-1 at a Pixel 8 EdgeTPU runtime of 3.8 ms. The paper's framing is Pareto optimality across CPUs, DSPs, GPUs, the Apple Neural Engine and EdgeTPUs, rather than a single accuracy number.
Why FLOPs mislead here
Arithmetic intensity, the ratio of arithmetic operations to bytes moved, is the relevant quantity. A dense 3x3 convolution reuses each loaded weight across many positions and each loaded activation across many filters. A depthwise convolution reuses each weight across positions only, so its intensity is roughly $k^2$ times lower.
On a roofline model, dense convolutions sit in the compute-bound region while depthwise convolutions sit in the memory-bound region. Cutting FLOPs by 9x therefore does not cut latency by 9x, and on hardware with high compute-to-bandwidth ratios it may not cut latency at all.
This is why V3 and V4 were designed against measured latency on specific chips rather than against FLOPs, and why V4's contribution is framed as a Pareto frontier across devices. Ma et al. (2018), ShuffleNet V2, made the same argument earlier and stated it as four practical guidelines, including "equal channel width minimises memory access cost" and "avoid excessive group convolution".
Deployment notes
- Fuse conv, batch norm and activation before export. On mobile runtimes the fused graph can be several times faster than the unfused one.
- Prefer per-channel weight quantisation for depthwise layers, and evaluate INT8 accuracy rather than assuming it.
- Batch size 1 is the deployment case, so measure there. Throughput at batch 64 is a data-centre metric and tells you little about a camera preview.
What to learn next
- EfficientNet and compound scaling — how to grow a small network without guessing.
- Quantization in practice — the other half of making a model fit on a phone.
- Latency and throughput — why counting operations is not enough.