OCR and Document Vision

CRNN text recognition

A CRNN reads a whole word strip left to right without ever cutting it into letters, by turning image columns into a sequence and letting a recurrent layer decide where letters begin.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Why the CNN squeezes height but not width
  6. Where you have already seen it
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A CRNN reads a strip of text left to right, without deciding in advance where each letter starts.

CRNN stands for convolutional recurrent neural network. Convolutional is the image part. Recurrent is the part that reads in order.

The analogy you have already lived

Think of reading a friend's cursive handwriting. The letters run into each other, and you cannot point at where the m stops and the e begins.

You do not cut the word up. You sweep across it, and the letters resolve as you go, using what came before to settle what comes next.

That sweep is the whole design of a CRNN. It refuses to cut.

Why it exists

Older OCR cut a word into single letters, then classified each one. That works on a printed government form and collapses everywhere else.

Printed letters touch. In small print, r and n merge and look like m. Cut wrongly once and everything after is wrong.

Handwriting has no gaps at all. There is nothing to cut on.

A wrong cut is unrecoverable. The classifier never sees the whole letter, so it cannot notice that the cut was bad.

A CRNN removes the cutting step. It looks at the strip in thin vertical slices and outputs a guess per slice. Later, repeats and blanks are collapsed into a word.

How it works

   strip of pixels:  "C A T"  (one image, 96 pixels wide)
         |
   [ CNN ]   squeezes height away, keeps width
         |
   a row of 48 feature columns, left to right
         |
   [ LSTM ]  sweeps left-to-right and right-to-left
         |
   per column: which letter is here? or nothing?
         |
   C C - A A A - - T T   ->  collapse  ->  "CAT"

The - is a special extra symbol called blank — it means "no letter finishes here". Blanks are what let the model say LL for a double letter and L - L for two separate ones.

Why the CNN squeezes height but not width

This is the neat trick. The CNN pools the strip until it is only a few pixels tall, but keeps its width.

Height carries no order information — a letter's top and bottom belong to the same letter. Width does carry order — left comes before right.

So height gets collapsed into "what does this column look like", and width becomes time.

Where you have already seen it

  • Number plate readers at a toll booth.
  • Bank cheque amount readers.
  • Reading the expiry date off a card in a payments app.
  • The text layer inside most open-source OCR engines, including PaddleOCR.

The honest part

A CRNN reads one line at a time. It has no idea what the line means, and no memory of the line above.

So l and 1, O and 0, 5 and S stay genuinely ambiguous. At the pixel level they often are. Real systems fix this outside the model, with rules like "this field holds six digits".

Remember this

  • A CRNN never cuts the word into letters.
  • Image height is squeezed away; image width becomes a sequence of timesteps.
  • A blank symbol marks "nothing finishes here", which is what makes double letters possible.

What to learn next

  • CTC loss — the alignment-free loss that makes this model trainable.
  • TrOCR and transformer OCR — what replaced the CRNN, and what it gave up.
  • LSTM — the recurrent layer doing the left-to-right sweep.

Developer — Code and libraries.

Setup

bash
pip install torch==2.5.1 numpy==1.26.4

CPU only. The model below trains in about a minute on an ordinary laptop, and it genuinely learns to read — no pretrained checkpoint, no download.

A CRNN that trains and reads

crnn.py
import numpy as np
import torch
import torch.nn as nn

torch.manual_seed(0)
rng = np.random.default_rng(0)

FONT = {
    "A": ["..#..", ".#.#.", "#...#", "#...#", "#####", "#...#", "#...#"],
    "C": [".###.", "#...#", "#....", "#....", "#....", "#...#", ".###."],
    "E": ["#####", "#....", "#....", "####.", "#....", "#....", "#####"],
    "L": ["#....", "#....", "#....", "#....", "#....", "#....", "#####"],
    "N": ["#...#", "##..#", "#.#.#", "#.#.#", "#..##", "#...#", "#...#"],
    "R": ["####.", "#...#", "#...#", "####.", "#.#..", "#..#.", "#...#"],
    "T": ["#####", "..#..", "..#..", "..#..", "..#..", "..#..", "..#.."],
    "1": ["..#..", ".##..", "..#..", "..#..", "..#..", "..#..", ".###."],
    "9": [".###.", "#...#", "#...#", ".####", "....#", "#...#", ".###."],
}
ALPHABET = "".join(sorted(FONT))          # index 0 is reserved for the CTC blank
GLYPH = {c: np.array([[x == "#" for x in r] for r in FONT[c]], np.float32) for c in FONT}
IMG_H, IMG_W, SCALE = 24, 96, 2


def draw(word, jitter=True):
    """Render a word into a fixed 24x96 strip, with a random gap between letters."""
    strip = np.zeros((7, 0), np.float32)
    for ch in word:
        gap = rng.integers(1, 4) if jitter else 2
        strip = np.hstack([strip, GLYPH[ch], np.zeros((7, gap), np.float32)])
    big = np.kron(strip, np.ones((SCALE, SCALE), np.float32))
    out = np.zeros((IMG_H, IMG_W), np.float32)
    top = (IMG_H - big.shape[0]) // 2
    left = rng.integers(0, 6) if jitter else 3
    w = min(big.shape[1], IMG_W - left)
    out[top:top + big.shape[0], left:left + w] = big[:, :w]
    return out


def batch(n, lo=3, hi=6):
    words = ["".join(rng.choice(list(ALPHABET), rng.integers(lo, hi + 1))) for _ in range(n)]
    images = torch.tensor(np.stack([draw(w) for w in words]))[:, None]     # N,1,H,W
    targets = torch.tensor([ALPHABET.index(c) + 1 for w in words for c in w])
    lengths = torch.tensor([len(w) for w in words])
    return words, images, targets, lengths


class CRNN(nn.Module):
    def __init__(self, n_classes):
        super().__init__()
        self.cnn = nn.Sequential(
            nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
            nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d((2, 1)),
            nn.Conv2d(32, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d((2, 1)),
        )
        self.rnn = nn.LSTM(32 * 3, 64, bidirectional=True, batch_first=False)
        self.head = nn.Linear(128, n_classes)

    def forward(self, x):
        f = self.cnn(x)                       # N, C, H', W'
        f = f.permute(3, 0, 1, 2).flatten(2)  # W', N, C*H'  -- width becomes time
        out, _ = self.rnn(f)
        return self.head(out)                 # W', N, classes


model = CRNN(len(ALPHABET) + 1)
print("input strip:      ", tuple(batch(1)[1].shape))
with torch.no_grad():
    probe = model.cnn(batch(1)[1])
    print("after the CNN:    ", tuple(probe.shape), " (batch, channels, height, width)")
    print("fed to the LSTM:  ", tuple(model(batch(1)[1]).shape), " (timesteps, batch, classes)")
print(f"so one 96-pixel-wide strip becomes {probe.shape[3]} timesteps, "
      f"about {IMG_W // probe.shape[3]} pixels each")
print("alphabet:", ALPHABET, " classes:", len(ALPHABET) + 1, "(9 letters + 1 blank)\n")

loss_fn = nn.CTCLoss(blank=0, zero_infinity=True)
opt = torch.optim.Adam(model.parameters(), lr=3e-3)

for step in range(1, 601):
    _, images, targets, lengths = batch(16)
    logits = model(images)
    logp = logits.log_softmax(2)
    input_lengths = torch.full((16,), logits.shape[0], dtype=torch.long)
    loss = loss_fn(logp, targets, input_lengths, lengths)
    opt.zero_grad()
    loss.backward()
    opt.step()
    if step % 100 == 0:
        print(f"step {step:4d}   ctc loss {loss.item():.3f}")


def greedy(logits):
    best = logits.argmax(2).T                  # batch, time
    out = []
    for row in best.tolist():
        chars, prev = [], -1
        for k in row:
            if k != prev and k != 0:           # drop repeats, then drop blanks
                chars.append(ALPHABET[k - 1])
            prev = k
        out.append("".join(chars))
    return out


model.eval()
words, images, _, _ = batch(12)
with torch.no_grad():
    read = greedy(model(images))
right = sum(a == b for a, b in zip(words, read))
print(f"\nexact word matches on 12 unseen strips: {right}/12")
for a, b in list(zip(words, read))[:6]:
    print(f"   truth {a:<7} read {b:<7} {'ok' if a == b else 'WRONG'}")
Output
input strip:       (1, 1, 24, 96)
after the CNN:     (1, 32, 3, 48)  (batch, channels, height, width)
fed to the LSTM:   (48, 1, 10)  (timesteps, batch, classes)
so one 96-pixel-wide strip becomes 48 timesteps, about 2 pixels each
alphabet: 19ACELNRT  classes: 10 (9 letters + 1 blank)

step  100   ctc loss 2.605
step  200   ctc loss 2.572
step  300   ctc loss 2.463
step  400   ctc loss 2.068
step  500   ctc loss 0.243
step  600   ctc loss 0.033

exact word matches on 12 unseen strips: 12/12
   truth AE11R   read AE11R   ok
   truth CATE    read CATE    ok
   truth ETE     read ETE     ok
   truth ATEAR   read ATEAR   ok
   truth AE1     read AE1     ok
   truth LEEA9   read LEEA9   ok

Reading the output carefully

(1, 32, 3, 48) is the whole architectural idea. The input was 24 pixels tall and 96 wide. Height fell to 3, width stayed at 48. Look at the pooling layers: the first is MaxPool2d(2), which halves both. The next two are MaxPool2d((2, 1)), which halve height and leave width alone. That asymmetry is deliberate and it is what makes this a sequence model.

32 * 3 = 96 is the LSTM's input size. Flattening channels and remaining height gives one 96-number vector per column. Each column becomes a timestep.

48 timesteps for at most 6 characters is not wasteful, it is required. CTC needs at least one timestep per output character, plus room for blanks between repeated characters. A common rule is to keep timesteps at roughly four times the maximum expected string length.

The loss curve is the interesting part. It sits near 2.5 for 300 steps, barely moving, and then falls off a cliff. That plateau is real and it confuses people the first time. Early in training the model emits blank everywhere, because blank is the safest single guess under CTC. It stays there until the CNN features become good enough for a letter to beat blank, and then everything unlocks at once.

The greedy decoder is nine lines and it is what most engines ship. Take the best class per timestep, drop runs of repeats, then drop blanks. Order matters — collapsing repeats before removing blanks is what preserves LL in LEEA9.

Why the blank symbol exists at all

Consider the model's per-timestep output for a word with a double letter.

   timesteps:   1  2  3  4  5  6
   argmax:      E  E  E  E  E  E     ->  collapse repeats  ->  "E"
   argmax:      E  E  -  -  E  E     ->  collapse repeats  ->  "EE"

Without a blank, there is no way to write "two E's in a row". The blank is a separator that survives the repeat-collapsing rule. This is explained fully in CTC loss.

Common mistakes

Feeding logits instead of log_softmax output to nn.CTCLoss. PyTorch's CTC expects log probabilities. Passing raw logits trains to a plausible-looking but wrong loss and converges badly.

Getting the tensor order wrong. nn.CTCLoss wants (time, batch, classes) by default, while almost everything else in PyTorch is batch-first. The permute(3, 0, 1, 2) in forward is doing that work.

Using class index 0 for a real character. blank=0 is the default. If your alphabet starts at index 0, your first letter is silently treated as blank. Note the + 1 when building targets.

Downsampling width too aggressively. Three MaxPool2d(2) layers would leave 12 timesteps, which cannot represent a 6-character word with any repeated letters. If loss refuses to fall below about 2, check this before anything else.

Squashing every crop to a fixed width. Real strips vary from 3 to 30 characters. Resize by height and pad by width, then pass true input_lengths. Stretching a long word into a short box destroys it.

Try it yourself

Change the last two pooling layers to nn.MaxPool2d(2) so width is divided by 8 as well. Rerun. The loss will stall high and the reads will be truncated, because 12 timesteps cannot express the longer words. That single experiment teaches the time-budget rule better than any formula.

What to learn next

  • CTC loss — the alignment-free loss that makes this model trainable.
  • TrOCR and transformer OCR — what replaced the CRNN, and what it gave up.
  • LSTM — the recurrent layer doing the left-to-right sweep.

Researcher — Mathematics and papers.

The architecture

CRNN comes from Shi, Bai and Yao (2017), An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition, TPAMI 39(11). Three stacked components:

  1. Convolutional feature extractor. Input is normalised to fixed height $H$ with variable width $W$. Pooling is anisotropic after the first stages — $2 \times 1$ windows — so the final map is $C \times 1 \times W'$ or $C \times h \times W'$ with small $h$.
  2. Map-to-sequence. Column $t$ of the feature map becomes frame $x_t \in \mathbb{R}^{Ch}$. Crucially, frame $t$ corresponds to a fixed receptive field in the input, so the sequence is spatially ordered by construction.
  3. Bidirectional LSTM plus a linear projection to $|\mathcal{A}| + 1$ classes, where $\mathcal{A}$ is the alphabet and the extra class is the blank.

The receptive field of frame $t$ spans considerably more than $W/W'$ input pixels. In the code above, each of the 48 frames advances 2 pixels but sees roughly 15 pixels of context through the stacked $3\times3$ convolutions — which is why a frame can carry information about a whole narrow character.

Why the sequence axis must not be over-pooled

CTC requires $T \geq \ell + r$, where $T$ is the number of frames, $\ell$ the label length, and $r$ the number of adjacent repeated labels in the target. Violating it makes the target unreachable and the loss infinite; zero_infinity=True masks that to zero, which silently discards those samples rather than crashing. If a fraction of your batch is being zeroed, training looks fine and accuracy plateaus for no visible reason. Check input_lengths against target_lengths explicitly.

Alternatives to the recurrent core

The R in CRNN is the least essential letter.

  • Fully convolutional. Replace the BiLSTM with stacked dilated or 1D convolutions. Same accuracy on printed text, far better parallelism, and a bounded context that sometimes helps.
  • Transformer encoder. Self-attention over frames, unbounded context. Standard in newer recognisers; see attention.
  • SVTR (Du et al., IJCAI 2022) drops the CNN/RNN split entirely for a vision-transformer-style patch model with local and global mixing blocks, and is what PP-OCRv4 and v5 recognisers are built on.
  • Attention decoders (Shi et al., ASTER, TPAMI 2019) replace CTC with an autoregressive decoder. They handle character-level dependencies and irregular text better, and they can hallucinate — a property CTC structurally lacks.

The rectification question

Curved and perspective-distorted text remains the standing weakness. Two responses:

  • Explicit rectification. A spatial transformer network predicts a thin-plate-spline warp that flattens the crop before recognition (ASTER, MORAN). Adds parameters, adds a second failure mode.
  • Robust features. Train on heavily-augmented synthetic data — MJSynth (Jaderberg et al., 2014, 9M images) and SynthText (Gupta et al., CVPR 2016, 800K images) are the standard pair — and let the model absorb the variation.

Baek et al. (2019), What is wrong with scene text recognition model comparisons?, is the single most useful paper here. It decomposes recognisers into transformation, feature extraction, sequence modelling and prediction stages, and re-benchmarks all combinations under identical data. The headline finding is that a large share of published gains came from differences in training data rather than architecture — a caution worth carrying into any recogniser comparison you read.

Cost

For a strip of width $W$ and height $H$, the CNN dominates at $O(HWC^2k^2)$ and the BiLSTM adds $O(W' d^2)$. Because $H$ is fixed (typically 32) and $W' \ll W$, inference is close to linear in strip width. Batched CPU inference at a few milliseconds per word is routine, which is why pipeline OCR remains dramatically cheaper than any VLM page reader — a comparison drawn out in TrOCR and transformer OCR.

What to learn next

  • CTC loss — the alignment-free loss that makes this model trainable.
  • TrOCR and transformer OCR — what replaced the CRNN, and what it gave up.
  • LSTM — the recurrent layer doing the left-to-right sweep.

What to learn next

These follow on from what you just read.

  • OCR and Document Vision

    CTC loss

    CTC lets you train a reader on "this strip says CAT" without ever labelling which pixels are the C, by scoring every alignment at once instead of picking one.

  • OCR and Document Vision

    Tesseract in practice

    Tesseract is free, offline and everywhere, and it repays a good scan and the right page segmentation mode far more than it repays clever code.

  • OCR and Document Vision

    TrOCR and transformer OCR

    TrOCR reads a text crop with an image transformer and writes the answer with a language decoder, which makes it strong on messy text and capable of inventing words that were never there.