Text detection models
Text detection finds where the words are before anything reads them, and the modern winners do it by painting a shrunken map of every word and growing it back.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Text detection draws a box around every word in a picture, without reading any of them.
The analogy you have already lived
Think of scanning a noticeboard at a railway station for your train number. Before you read a single word, your eyes have already found the patches that are writing. The posters, the timetable, the scribbled note.
You located the text before you understood it. Detection is that first sweep, done by a model.
Why it exists
Reading a whole photograph at once is far harder than reading one word. Words come in different sizes, angles and colours, scattered across the picture.
Split the job and both halves get easier. One model answers "where is writing?" Another answers "what does this small strip say?"
The split also lets you fix problems separately. A missed word is a detection bug. A misread word is a recognition bug. Different problem, different fix.
Why an ordinary object detector struggles
You may know detectors that find dogs and cars, like the one in object detection. Text breaks their assumptions.
Text has extreme shapes. A word can be twenty times wider than it is tall. Dogs are not.
Text is not always horizontal. Signboards curve. Photos are taken at an angle.
Text crowds together. Two words on a receipt can sit six pixels apart. Two dogs rarely overlap that tightly.
So the field went a different way. Instead of predicting boxes, paint the picture.
Painting instead of boxing
A segmentation model labels every pixel rather than drawing rectangles. It outputs a map where each pixel says "I am inside a word" or "I am not".
photo text-probability map
┌──────────────┐ ┌──────────────┐
│ PLATFORM 12 │ -> │ ▓▓▓▓▓▓▓▓ ▓▓ │
│ │ │ │
└──────────────┘ └──────────────┘
then: find each connected blob -> one box per blobThis handles curves and slants for free, because a blob has no fixed shape.
The trap, and the clever fix
There is a catch that took the field years to solve neatly.
The model does not output a map at full picture size. To stay fast it works on a shrunken grid, roughly one output cell per four input pixels.
At that coarseness, two words sitting six pixels apart share a cell. Their blobs touch, they fuse, and you get one box covering both. Every word in the line merges into one long smear.
The fix is unexpectedly plain. Train the model to paint a shrunken version of each word, pulled in from its edges. The gaps become wide enough to survive. Afterwards, grow each blob back outward by the same amount.
real words: [PLATFORM][12] gap too small, blobs fuse
shrunk targets: [PLATFOR][2] gap now clear, blobs separate
grown back: [PLATFORM][12] two boxes, correctThat shrink-and-regrow idea is the heart of DBNet, which is the detector inside a great many production OCR systems today.
Where you have already seen it
- Google Lens putting a highlight over every line of a menu you point at.
- A translation app boxing text on a signboard before replacing it.
- Number plate systems finding the plate in a wide camera view.
- Your phone letting you long-press and copy text out of a photo.
The honest part
Detection is usually the weakest link, and it fails silently. A missed word produces no error message. It produces a document that quietly lacks a line.
So the first thing to do when OCR output looks wrong is not to change the recognition model. Draw the detected boxes on the image and look at them.
Remember this
- Detection finds where words are; recognition reads them. Separate jobs.
- Painting a per-pixel map beats drawing boxes, because text bends and stretches.
- Nearby words fuse on a coarse map, so models learn shrunken words and grow them back.
What to learn next
- CRNN text recognition — the model that reads each detected strip.
- Image segmentation — the per-pixel labelling idea this whole family is built on.
- Document layout analysis — detection one level up, at paragraphs and tables.
Developer — Code and libraries.
Setup
pip install opencv-python==4.10.0.84 numpy==1.26.4Running DBNet needs a trained checkpoint. The idea inside it does not. Below we reproduce the exact failure DBNet was built to fix, and the exact fix, using nothing but morphology — so the mechanism is visible rather than described.
The merge problem and the shrink fix
import cv2
import numpy as np
H, W, STRIDE = 60, 200, 4
# Two words printed close together, the way they sit on a real receipt line.
page = np.zeros((H, W), np.uint8)
cv2.rectangle(page, (20, 22), (88, 40), 255, -1) # word 1
cv2.rectangle(page, (95, 22), (176, 40), 255, -1) # word 2, a 6-pixel gap away
def boxes_from(mask):
n, _, stats, _ = cv2.connectedComponentsWithStats(mask.astype(np.uint8), 8)
return [tuple(int(v) for v in stats[i, :4]) for i in range(1, n)] # x, y, w, h
def predict(target):
"""Stand in for a detector: it emits a probability map at 1/4 resolution."""
small = cv2.resize(target, (W // STRIDE, H // STRIDE), interpolation=cv2.INTER_AREA)
small = cv2.GaussianBlur(small, (5, 5), 1.2) / 255.0 # networks give soft edges
return cv2.resize(small, (W, H), interpolation=cv2.INTER_LINEAR)
print("the two words really are:", boxes_from(page > 0))
print(f"gap between them: {95 - 89} pixels, at a detector stride of {STRIDE}\n")
print("ATTEMPT 1 - train the network on the word shapes themselves")
mask_a = predict(page) > 0.5
print(" boxes recovered:", boxes_from(mask_a))
print(" a 6-pixel gap is 1.5 cells wide at stride 4, so the words fused\n")
print("ATTEMPT 2 - train it on SHRUNK words instead (this is what DB does)")
shrunk = cv2.erode(page, np.ones((9, 9), np.uint8))
mask_b = predict(shrunk) > 0.5
print(" boxes recovered:", boxes_from(mask_b))
print(" the gap is now wide enough to survive the stride\n")
print("STEP 3 - grow each shrunk blob back out to the real word")
grown = cv2.dilate(mask_b.astype(np.uint8), np.ones((9, 9), np.uint8))
print(" final boxes:", boxes_from(grown))
print(" true boxes:", boxes_from(page > 0))
print(" close, though each edge is a pixel or two out\n")
print("STEP 4 - why DB learns the threshold instead of fixing it")
print(" B = 1 / (1 + exp(-k * (P - T))) with k = 50")
for p in [0.30, 0.45, 0.49, 0.50, 0.51, 0.55, 0.70]:
print(f" P = {p:.2f} T = 0.50 -> B = {1 / (1 + np.exp(-50 * (p - 0.5))):.4f}")
print(" it is a step in all but name, yet it still has a usable gradient")the two words really are: [(20, 22, 69, 19), (95, 22, 82, 19)]
gap between them: 6 pixels, at a detector stride of 4
ATTEMPT 1 - train the network on the word shapes themselves
boxes recovered: [(20, 22, 157, 19)]
a 6-pixel gap is 1.5 cells wide at stride 4, so the words fused
ATTEMPT 2 - train it on SHRUNK words instead (this is what DB does)
boxes recovered: [(26, 26, 57, 11), (101, 26, 70, 11)]
the gap is now wide enough to survive the stride
STEP 3 - grow each shrunk blob back out to the real word
final boxes: [(22, 22, 65, 19), (97, 22, 78, 19)]
true boxes: [(20, 22, 69, 19), (95, 22, 82, 19)]
close, though each edge is a pixel or two out
STEP 4 - why DB learns the threshold instead of fixing it
B = 1 / (1 + exp(-k * (P - T))) with k = 50
P = 0.30 T = 0.50 -> B = 0.0000
P = 0.45 T = 0.50 -> B = 0.0759
P = 0.49 T = 0.50 -> B = 0.3775
P = 0.50 T = 0.50 -> B = 0.5000
P = 0.51 T = 0.50 -> B = 0.6225
P = 0.55 T = 0.50 -> B = 0.9241
P = 0.70 T = 0.50 -> B = 1.0000
it is a step in all but name, yet it still has a usable gradientReading the output carefully
[(20, 22, 157, 19)] is the whole problem in one tuple. One box, 157 pixels wide, spanning both words and the gap. The recogniser will now be handed a strip containing two words and will produce something like PLATFORM12. Nothing downstream can undo this.
The shrunk boxes are visibly smaller: (26, 26, 57, 11) against a true (20, 22, 69, 19). That is deliberate. The network's job changed from "paint the word" to "paint the core of the word", and the core is easier to keep separate.
The regrown boxes are close but not exact: (22, 22, 65, 19) against (20, 22, 69, 19). Two pixels short on each side. This is honest and it matters — DB-style boxes are approximately right, not exactly right. Production pipelines add a margin of a few pixels before cropping, because a recogniser handed a box that clips the first stroke of a letter reads it wrong.
The B column is the differentiable binarisation curve. Between P = 0.45 and P = 0.55 the output swings from 0.076 to 0.924. Outside that range it is flat at 0 or 1. It behaves as a hard threshold in the forward pass while still passing a gradient backwards, which is what lets DB learn its own threshold map instead of using a hand-set 0.5.
The families of text detector
| Approach | Representative | Predicts | Handles curves | Notes |
|---|---|---|---|---|
| Regression | EAST (2017) | Rotated box per pixel | No | Fast, still fine for straight signage |
| Character-level | CRAFT (2019) | Character heat + affinity | Yes | Groups characters into words itself |
| Segmentation + shrink | DBNet (2020), DBNet++ (2022) | Shrunken word map + threshold map | Yes | The common production default |
| Query-based | DETR-style text detectors | Set of polygons | Yes | Cleaner design, heavier to train |
If you are choosing today rather than researching: PaddleOCR ships a DB detector and is the pragmatic default for printed documents in many languages. docTR offers DBNet and LinkNet backbones with a friendly PyTorch and TensorFlow API. Both give you polygons, and both take a page in and give boxes out without you training anything.
Common mistakes
Cropping a box exactly. Detected boxes clip strokes. Pad by 2–5% of the box height before sending a crop to the recogniser, and measure the effect.
Ignoring the polygon and keeping only the axis-aligned box. For slanted text this pulls in the neighbouring words' pixels. Warp the polygon to a straight rectangle first — cv2.getPerspectiveTransform plus cv2.warpPerspective, which is covered in homographies and perspective warp.
Resizing a page down before detection. Detectors have a smallest readable text height, usually around 8–10 pixels in the network's input. Shrinking a 300 DPI page to 640 pixels wide deletes footnotes entirely, and the failure is silent.
Measuring detection with a classification metric. Report recall separately from precision. A detector at 99% precision and 90% recall loses one word in ten, and the average of the two hides it.
Assuming boxes come out in reading order. They do not. Sorting is a separate stage.
Try it yourself
Change the erosion kernel from 9 to 3 and rerun. The shrink is now too small to open the gap, and Attempt 2 fails the same way Attempt 1 did. That is the shrink ratio hyperparameter, and this is what happens when it is set too low for the text spacing in your documents.
What to learn next
- CRNN text recognition — the model that reads each detected strip.
- Image segmentation — the per-pixel labelling idea this whole family is built on.
- Document layout analysis — detection one level up, at paragraphs and tables.
Researcher — Mathematics and papers.
The three formulations
Regression. EAST (Zhou et al., CVPR 2017, An efficient and accurate scene text detector) predicts, at each pixel inside a text region, the four distances to the box edges plus a rotation angle. It removed the region-proposal stage entirely and reached 13 FPS at 720p. It cannot represent curved text, since the output space is a rotated rectangle.
Character-level grouping. CRAFT (Baek et al., CVPR 2019, Character region awareness for text detection) predicts two maps: a character-region score and an affinity score between adjacent characters. Words are recovered by thresholding both and taking connected components on the union. Character-level ground truth does not exist at scale, so CRAFT bootstraps pseudo-labels from a word-level model — an elegant weak-supervision loop. It remains strong on curved and arbitrarily-shaped text.
Segmentation with adaptive binarisation. DB (Liao et al., AAAI 2020, Real-time scene text detection with differentiable binarization) is the design reproduced above. The network emits a probability map $P$ and a threshold map $T$ at stride 4, and combines them:
$$ \hat{B}{i,j} = \frac{1}{1 + e^{-k(P{i,j} - T_{i,j})}} $$
Where $k = 50$ in the paper. As $k \to \infty$ this approaches a step function, so $\hat{B}$ is an approximate binary map that is still differentiable. Because $T$ is predicted per pixel, the network learns to raise the threshold precisely at word boundaries, which sharpens separation further.
The gradient of the binary cross-entropy loss with respect to $x = P - T$ carries a factor $k \cdot \sigma(kx)(1 - \sigma(kx))$, which is largest exactly where $P \approx T$. That amplification at the boundary is the mechanism the paper credits for the improvement, rather than the extra output head on its own.
Label generation. Ground-truth polygons are shrunk by the Vatti clipping algorithm with offset
$$ D = \frac{A(1 - r^2)}{L} $$
Where $A$ is the polygon area, $L$ its perimeter, and $r = 0.4$ the shrink ratio. At inference the predicted region is expanded with $D' = A' r' / L'$, using $r' = 1.5$. The 2-pixel discrepancy in the code output is the discrete analogue of the fact that erosion and dilation by the same structuring element are not inverses.
DBNet++ (Liao et al., TPAMI 2022) adds an Adaptive Scale Fusion module — spatial attention over the feature-pyramid levels — improving robustness to text size without materially changing speed.
Cost and metrics
Post-processing cost is dominated by connected components, $O(HW/s^2)$ at stride $s$, which is why DB reports real-time throughput while contour-based competitors do not.
Detection is scored by IoU-thresholded precision, recall and F-measure on ICDAR 2015 (incidental scene text), Total-Text and CTW1500 (curved). Two evaluation traps recur in the literature:
- The one-to-one matching rule. ICDAR's protocol penalises a single detection covering two ground-truth words, and also two detections splitting one word. Reported F-measures are not comparable across papers using DetEval versus the IoU protocol.
- "Do not care" regions. ICDAR 2015 marks illegible text as ignorable. Implementations differ in whether a detection landing there counts as a false positive, and the difference is worth one to two F points.
Where the field has moved
Two shifts matter for someone building today.
Unified spotters. Detection and recognition trained jointly — ABCNet, Mask TextSpotter, TESTR, DeepSolo — remove the crop-and-hand-off boundary and let recognition gradients improve localisation. They win on benchmarks and cost more to train.
Detection-free page reading. Donut, Nougat and general vision-language models read a page without ever emitting a box. This is a genuine capability change for layout-heavy documents, and a genuine regression for anything needing provenance: without boxes you cannot highlight the source of an extracted field, and you cannot bound hallucination. Regulated document pipelines in 2026 still overwhelmingly keep an explicit detector, sometimes purely to supply evidence boxes for a VLM's answer.
What to learn next
- CRNN text recognition — the model that reads each detected strip.
- Image segmentation — the per-pixel labelling idea this whole family is built on.
- Document layout analysis — detection one level up, at paragraphs and tables.