CNN Backbones and Pretraining

EfficientNet and compound scaling

EfficientNet grows depth, width and image size together by fixed ratios instead of guessing which one to increase, and gets far more accuracy per unit of compute.

Read these first

On this page 7
  1. What people did before
  2. The idea in one sentence
  3. Why the ratios are not equal
  4. Where you have seen the result
  5. The honest part
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

EfficientNet grows a network by increasing its depth, its width and the picture size together, in fixed proportions.

Think about cooking for a bigger crowd. Four guests turn into sixteen, so you need four times the rice.

Add four times the rice to the same small pot, with the same water. You get a burnt mess. The rice, the water and the pot all have to grow together.

Networks have the same three ingredients. Depth is how many layers. Width is how many channels each layer holds. Resolution is how big the input picture is. Grow one alone and you waste the extra effort.

What people did before

Everyone scaled one ingredient at a time, by feel.

ResNet-50 became ResNet-101 by adding layers. Wide ResNet added channels. Some papers trained on bigger images. Each choice was defended with intuition and a few experiments.

Two things went wrong with that. Deeper networks stopped improving after a point. Wider ones stopped picking up complicated patterns. Bigger pictures without more layers gave a network no way to use the extra detail.

The idea in one sentence

Pick one number for how much bigger you want the model. Then grow all three ingredients by fixed ratios tied to that single number.

   want twice the compute?

   depth      x 1.2
   width      x 1.1
   resolution x 1.15

   ... and those three multiply out to about 2x the work

Those three ratios were not guessed. The team searched for them once on a small model, then reused them for every larger size.

That is why the family runs B0, B1, B2 and onwards. Each step is one turn of the same dial.

Why the ratios are not equal

Doubling the layers doubles the work. Doubling the channels roughly quadruples it, because every channel talks to every other channel.

Doubling the picture width also quadruples it, because a picture has both height and width. So width and resolution get smaller ratios than depth. The arithmetic works out to roughly a doubling overall, which is exactly what was wanted.

Where you have seen the result

  • EfficientNet-B0 is a common default inside phone apps that classify photos.
  • It became the standard "small but accurate" model in Kaggle competitions for years.
  • The bigger members were state of the art on image recognition when they appeared in 2019.

The honest part

EfficientNet is efficient in arithmetic, and that is not the same as fast.

The larger members need big input pictures, which need a lot of memory during training. Several members were slower to train than a plain ResNet with a similar score. The follow-up version was built to fix exactly that.

Being clever about counting operations does not automatically make a wall clock move faster. Read that twice, because it is the most common misunderstanding about this family.

Remember this

  • Depth, width and resolution should grow together, not one at a time.
  • One dial controls all three, using ratios found by a one-time search.
  • Fewer operations is not the same as less time, and this family is a good example.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install "torch==2.5.1" "torchvision==0.20.1"

Run against PyTorch 2.5.1 and torchvision 0.20.1 on CPU. Nothing is downloaded; the accuracy and preprocessing figures come from metadata that ships with torchvision.

Do the scaling arithmetic, then check it against real models

compound_scaling.py
import torch
from torch.utils.flop_counter import FlopCounterMode
from torchvision.models import (efficientnet_b0, efficientnet_b1, efficientnet_b2,
                                efficientnet_b3, efficientnet_b4, resnet50,
                                EfficientNet_B0_Weights, EfficientNet_B4_Weights)

ALPHA, BETA, GAMMA = 1.2, 1.1, 1.15          # depth, width, resolution, from the paper

print("the constraint the authors searched under:")
print(f"  alpha * beta^2 * gamma^2 = {ALPHA * BETA**2 * GAMMA**2:.3f}   (they wanted about 2)")

print("\nfour ways to spend twice the compute")
print(f"{'plan':28s} {'depth':>7s} {'width':>7s} {'res':>7s} {'cost':>7s}")
plans = [
    ("depth only", 2.0, 1.0, 1.0),
    ("width only", 1.0, 2 ** 0.5, 1.0),
    ("resolution only", 1.0, 1.0, 2 ** 0.5),
    ("compound, phi = 1", ALPHA, BETA, GAMMA),
]
for name, d, w, r in plans:
    print(f"{name:28s} {d:7.2f} {w:7.2f} {r:7.2f} {d * w * w * r * r:7.2f}")

print("\ncompound scaling as phi grows")
print(f"{'phi':>4s} {'depth':>7s} {'width':>7s} {'res mult':>9s} {'224 ->':>7s} {'cost':>7s}")
for phi in range(0, 7):
    d, w, r = ALPHA ** phi, BETA ** phi, GAMMA ** phi
    print(f"{phi:4d} {d:7.2f} {w:7.2f} {r:9.2f} {round(224 * r):7d} {2 ** phi:7d}")


def gmacs(model, res):
    model.eval()
    counter = FlopCounterMode(display=False)
    with counter, torch.no_grad():
        model(torch.zeros(1, 3, res, res))
    return counter.get_total_flops() / 2 / 1e9        # /2 turns FLOPs back into MACs


print("\nwhat torchvision actually ships (measured, not quoted)")
rows = [("efficientnet_b0", efficientnet_b0), ("efficientnet_b1", efficientnet_b1),
        ("efficientnet_b2", efficientnet_b2), ("efficientnet_b3", efficientnet_b3),
        ("efficientnet_b4", efficientnet_b4), ("resnet50", resnet50)]
res_by_name = {"efficientnet_b0": 224, "efficientnet_b1": 240, "efficientnet_b2": 260,
               "efficientnet_b3": 300, "efficientnet_b4": 380, "resnet50": 224}
print(f"{'model':18s} {'res':>5s} {'params':>12s} {'GMACs':>8s}")
for name, fn in rows:
    m = fn(weights=None)
    res = res_by_name[name]
    print(f"{name:18s} {res:5d} {sum(p.numel() for p in m.parameters()):12,} {gmacs(m, res):8.2f}")

print("\nthe preprocessing that ships with the weights (no download needed to read it)")
print(" b0:", EfficientNet_B0_Weights.IMAGENET1K_V1.transforms())
print(" b4:", EfficientNet_B4_Weights.IMAGENET1K_V1.transforms())
Output
the constraint the authors searched under:
  alpha * beta^2 * gamma^2 = 1.920   (they wanted about 2)

four ways to spend twice the compute
plan                           depth   width     res    cost
depth only                      2.00    1.00    1.00    2.00
width only                      1.00    1.41    1.00    2.00
resolution only                 1.00    1.00    1.41    2.00
compound, phi = 1               1.20    1.10    1.15    1.92

compound scaling as phi grows
 phi   depth   width  res mult  224 ->    cost
   0    1.00    1.00      1.00     224       1
   1    1.20    1.10      1.15     258       2
   2    1.44    1.21      1.32     296       4
   3    1.73    1.33      1.52     341       8
   4    2.07    1.46      1.75     392      16
   5    2.49    1.61      2.01     451      32
   6    2.99    1.77      2.31     518      64

what torchvision actually ships (measured, not quoted)
model                res       params    GMACs
efficientnet_b0      224    5,288,548     0.39
efficientnet_b1      240    7,794,184     0.69
efficientnet_b2      260    9,109,994     0.99
efficientnet_b3      300   12,233,232     1.83
efficientnet_b4      380   19,341,616     4.39
resnet50             224   25,557,032     4.09

the preprocessing that ships with the weights (no download needed to read it)
 b0: ImageClassification(
    crop_size=[224]
    resize_size=[256]
    mean=[0.485, 0.456, 0.406]
    std=[0.229, 0.224, 0.225]
    interpolation=InterpolationMode.BICUBIC
)
 b4: ImageClassification(
    crop_size=[380]
    resize_size=[384]
    mean=[0.485, 0.456, 0.406]
    std=[0.229, 0.224, 0.225]
    interpolation=InterpolationMode.BICUBIC
)

Reading the output

The constraint comes out at 1.920, not exactly 2. The paper's grid search settled on the ratios 1.2, 1.1 and 1.15 subject to their product being about 2. Cost per step is therefore a little under a doubling. That is why the family's real FLOP counts drift from the powers of two.

The four plans all cost 2x and are not equally good. Cost is the same by construction; accuracy is not. The paper's finding is that single-axis scaling saturates: adding depth alone stops helping, and adding resolution alone gives the network detail it has no capacity to use.

Resolution grows the fastest in absolute terms. By the sixth step the input is 518 pixels wide. That is why the large members are memory-hungry: activation memory scales with the square of resolution, and gradients must be kept for every layer.

The measured numbers land close to the published ones. EfficientNet-B0 measures 0.39 GMACs, exactly the paper's 0.39B figure. Note the halving inside gmacs: PyTorch's counter reports two FLOPs per multiply-accumulate, while papers count one.

The headline comparison is the last two rows. EfficientNet-B4 does about the same arithmetic as ResNet-50, 4.39 against 4.09 GMACs, with 19.3 million parameters against 25.6 million, and scores several points higher on ImageNet. Same budget, more accuracy.

Each member has its own preprocessing. B0 crops to 224 after resizing to 256; B4 crops to 380 after resizing to 384. Feeding B4 a 224-pixel image runs without error and loses a lot of accuracy. Always call weights.transforms().

Common mistakes

Feeding every member 224x224. The most common EfficientNet bug. The model accepts it, the pooling layer adapts, and your accuracy is quietly wrong.

Comparing FLOPs across libraries without checking the convention. MACs and FLOPs differ by a factor of two. Some counters include batch norm and activations, and others do not. Compare only numbers produced by the same tool.

Expecting the FLOP saving to become a speed saving. B0 does one tenth of ResNet-50's arithmetic. On one desktop CPU at four threads, B0 came out consistently slower than ResNet-18, which does about five times more arithmetic. Depthwise and squeeze-excitation layers are memory-bound. Absolute timings move between machines and between runs, so measure yours.

Fine-tuning with the default dropout. EfficientNet uses stochastic depth plus dropout that increases with model size. On a small target dataset those defaults are often too strong; check model.classifier[0].p before blaming the learning rate.

Reaching for B7 because it tops a table. B7 needs 600-pixel inputs. On one consumer GPU it will train at a batch size of a handful of images. Start at B0 and move up only while the validation curve keeps rewarding you.

Try it yourself

Change the ratios to ALPHA, BETA, GAMMA = 1.4, 1.05, 1.1 and check that the product of the constraint stays near 2. That is a valid alternative point on the same budget curve, weighted towards depth. Then measure gmacs(efficientnet_b0(weights=None), 380) and see how much of B4's cost is resolution alone.

What to learn next

Researcher — Mathematics and papers.

The scaling formulation

Tan and Le (2019), EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, ICML, arxiv.org/abs/1905.11946, pose scaling as a constrained optimisation. With depth $d$, width $w$ and resolution $r$ multipliers:

$$ d = \alpha^{\phi}, \quad w = \beta^{\phi}, \quad r = \gamma^{\phi} $$ $$ \text{subject to} \quad \alpha \cdot \beta^2 \cdot \gamma^2 \approx 2, \quad \alpha \geq 1,\ \beta \geq 1,\ \gamma \geq 1 . $$

$\phi$ is the single user-chosen coefficient. The exponents on $\beta$ and $\gamma$ are 2 because convolution FLOPs scale linearly with depth, quadratically with width and quadratically with resolution. The constraint therefore fixes total FLOPs at approximately $2^{\phi}$ times the baseline.

A small grid search on the baseline gives $\alpha = 1.2$, $\beta = 1.1$, $\gamma = 1.15$. These are then held fixed while $\phi$ is varied, which avoids a search at every scale, the practical contribution of the paper.

The baseline EfficientNet-B0 itself came from a multi-objective neural architecture search over MobileNetV2-style inverted bottlenecks with squeeze-and-excitation, optimising accuracy and target latency jointly. So the family has two separate ingredients: a searched baseline and a scaling rule. They are often conflated; the scaling rule is the transferable part and has been applied to ResNets and to transformers.

Reported result: EfficientNet-B7 at 84.3% top-1, described as 8.4x smaller and 6.1x faster on inference than the best existing ConvNet at the time.

Why single-axis scaling saturates

The paper's empirical study shows accuracy versus FLOPs flattening for each axis alone:

  • Depth. Deeper networks capture richer features, but gains diminish sharply beyond a point even with residual connections and normalisation.
  • Width. Wide, shallow networks capture fine-grained features easily and struggle with high-level ones.
  • Resolution. More pixels help only if receptive field and capacity grow to consume them, which is the argument that ties this lesson back to receptive fields.

The interaction argument is the paper's core claim: higher resolution demands more layers to enlarge the receptive field and more channels to represent finer patterns. Scaling one axis while holding the others fixed leaves the added capacity unusable.

EfficientNetV2, and the honest correction

Tan and Le (2021), EfficientNetV2: Smaller Models and Faster Training, ICML, arxiv.org/abs/2104.00298, diagnose three problems with V1:

  1. Training with very large images is slow and memory-bound.
  2. Depthwise convolutions are slow in early layers, where channel counts are low and the hardware is under-used. Fused-MBConv, a plain 3x3 convolution replacing the expand-plus-depthwise pair, is faster there despite more FLOPs.
  3. Equally scaling every stage is suboptimal; V2 adds stages non-uniformly and caps the maximum image size.

They also introduce progressive learning: train at small resolution with weak regularisation, then increase both together. Reported: up to 6.8x smaller than prior models, and 87.3% top-1 on ImageNet with ImageNet-21k pretraining, "outperforming the recent ViT by 2.0% accuracy while training 5x-11x faster".

Point 2 deserves emphasis. It is a public admission that the FLOP-minimising design choice was the wrong choice on real hardware. The measurement in the developer block is the same phenomenon on a CPU. B0 is slower than ResNet-18 at a fifth of the arithmetic.

Scaling laws elsewhere

The compound idea reappears repeatedly:

  • Bello et al. (2021), Revisiting ResNets, apply modern training and scaling rules to ResNets and match EfficientNets at equal TPU training time, arguing again that recipe and scaling explain much of the reported gap.
  • Dollár, Singh and Girshick (2021), Fast and Accurate Model Scaling, propose scaling that keeps activation memory, rather than FLOPs, as the constrained quantity, since activations predict runtime better than FLOPs do.
  • Zhai et al. (2022), Scaling Vision Transformers, derive a saturating power law relating compute, data and accuracy for ViTs, a direct descendant of the same question.

Practical guidance

Use B0 to B2 when you have a fixed inference budget and modest data. Use V2-S or V2-M when training throughput matters, since they were designed for it. Prefer timm's implementations for the wider set of pretrained recipes, and be explicit about which weights you are loading: the same architecture name spans several recipes with several accuracy levels, exactly as it does for ResNet-50.

What to learn next

What to learn next

These follow on from what you just read.

  • CNN Backbones and Pretraining

    ConvNeXt

    ConvNeXt takes a ResNet and swaps in one transformer design choice at a time, ending with a pure convolutional network that matches Swin at the same compute.

  • CNN Backbones and Pretraining

    Swin transformer

    Swin runs attention inside small windows and shifts the windows every other layer, which makes attention affordable on large images while keeping the CNN's four-stage shape.

  • CNN Backbones and Pretraining

    Contrastive learning for images

    Contrastive learning trains a backbone with no labels by showing it two altered copies of the same photo and asking it to tell them apart from everybody else's.