Preprocessing scans for OCR
Straightening the page, removing the shadow and choosing a threshold per neighbourhood buys more OCR accuracy than any model change, and costs milliseconds.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Preprocessing means cleaning the picture before anything tries to read it.
The analogy you have already lived
Think of reading a bus timetable through a dirty window at dusk. You move your head to cut the glare, you tilt the sheet straight, you get closer.
You did not become a better reader. You made the page easier to read. Preprocessing is exactly that, done by code.
Why it exists
A phone photo of a bill is nothing like a clean scan. Three things go wrong almost every time.
The light is uneven. One side of the page catches the tube light, the other sits in your own shadow.
The page is crooked. Nobody holds a phone perfectly square to a piece of paper.
There is grain. Camera sensors add speckle, and JPEG saving adds smears around the letters.
Any of these can turn a readable page into garbage output. Fixing them is cheap. Retraining a model is not.
Turning grey into black and white
An OCR engine wants to know one thing per pixel: ink, or paper. Turning a grey image into that yes-or-no picture is called binarisation — making every pixel one of two values.
The naive way picks one brightness level for the whole page. Anything darker is ink, anything lighter is paper.
That works on a scanner and fails on a photo:
bright side of page dim side of page
paper = very light paper = quite dark
ink = very dark ink = nearly black
one cut for the whole page:
|------ paper ------|--- ink ---|
^ cut here
and the dim side's PAPER falls on the ink side of the cutThe fix is to stop asking one question for the whole page. Ask it per neighbourhood instead. "Is this pixel darker than its own surroundings?"
A shadow makes a whole region darker. The letters inside it are still darker than that region.
This is called adaptive thresholding — choosing the ink-or-paper cut separately for each small area.
Straightening the page
A crooked page breaks the step that groups letters into lines. Every OCR engine assumes text runs along image rows.
The trick for finding the tilt is pleasing. Rotate the page a little, count how many ink pixels sit in each row, and look at the pattern.
crooked text: ink is smeared across MANY rows
straight text: ink is packed into FEW rows, with clean empty rows betweenTry many small rotations, keep the one where the ink is packed tightest. That is called a projection profile — a count of ink per row.
Where you have already seen it
- The "scan document" mode in your phone camera that snaps the page flat and white.
- Adobe Scan and Microsoft Lens cleaning a photographed page.
- A photocopier's "lighten/darken" dial doing the crude version of the same job.
The honest part
Preprocessing can hurt. Turn the cleaning up too high and thin strokes vanish, dots on letters disappear, and Devanagari matras get eaten.
There is no universal setting. Good practice is to keep a folder of your own worst pages, and score every change against them.
Remember this
- Clean the picture first; it is the cheapest accuracy in the whole pipeline.
- One brightness cut for a whole photo fails. Choose a cut per neighbourhood.
- Straighten the page by finding the rotation that packs ink into the fewest rows.
What to learn next
- Text detection models — the stage that consumes this cleaned image.
- Tesseract in practice — the engine that gains the most from good preprocessing.
- Homographies and perspective warp — flattening a page photographed at an angle.
Developer — Code and libraries.
Setup
pip install opencv-python==4.10.0.84 numpy==1.26.4We build a scan that is deliberately bad — uneven lamp, sensor noise, six degrees of skew — and we keep the true ink mask, so every cleaning step can be scored instead of admired.
The whole cleaning stage, scored
import cv2
import numpy as np
FONT = {
"A": ["..#..", ".#.#.", "#...#", "#...#", "#####", "#...#", "#...#"],
"C": [".###.", "#...#", "#....", "#....", "#....", "#...#", ".###."],
"E": ["#####", "#....", "#....", "####.", "#....", "#....", "#####"],
"I": ["#####", "..#..", "..#..", "..#..", "..#..", "..#..", "#####"],
"L": ["#....", "#....", "#....", "#....", "#....", "#....", "#####"],
"N": ["#...#", "##..#", "#.#.#", "#.#.#", "#..##", "#...#", "#...#"],
"R": ["####.", "#...#", "#...#", "####.", "#.#..", "#..#.", "#...#"],
"T": ["#####", "..#..", "..#..", "..#..", "..#..", "..#..", "..#.."],
}
def ink_mask(text, scale=5):
cells = []
for ch in text:
g = np.zeros((7, 3)) if ch == " " else np.array(
[[c == "#" for c in r] for r in FONT[ch]], float)
cells += [g, np.zeros((7, 2))]
return np.kron(np.hstack(cells), np.ones((scale, scale))) > 0
letters = ink_mask("CENTRAL RAIL")
truth = np.zeros((120, 460), bool)
truth[30:30 + letters.shape[0], 30:30 + letters.shape[1]] = letters
# Build a realistically bad scan: dark ink on light paper, a lamp on the left,
# sensor noise, and the sheet laid down crooked.
page = np.where(truth, 45.0, 232.0)
page = page * np.linspace(1.05, 0.42, 460)[None, :] # bright left edge, dim right
page = np.clip(page + np.random.default_rng(0).normal(0, 6, page.shape), 0, 255)
TRUE_SKEW = 6.0
M = cv2.getRotationMatrix2D((230, 60), TRUE_SKEW, 1.0)
scan = cv2.warpAffine(page.astype(np.uint8), M, (460, 120), borderValue=200)
gt = cv2.warpAffine(truth.astype(np.uint8), M, (460, 120), flags=cv2.INTER_NEAREST) > 0
left, right = gt.copy(), gt.copy()
left[:, 230:] = False
right[:, :230] = False
print("paper grey level: left", int(scan[5, 20]), " right", int(scan[5, 440]))
print("ink grey level: left", int(np.median(scan[left])),
" right", int(np.median(scan[right])))
print("ink is", round(100 * gt.mean(), 1), "percent of the pixels on this page\n")
def overlap(pred, true):
return (pred & true).sum() / (pred | true).sum()
fixed = scan < 128
cut, otsu_img = cv2.threshold(scan, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
print(f"Otsu chose a single cut for the whole page at {int(cut)}\n")
smoothed = cv2.medianBlur(scan, 3) # kills single-pixel specks
adaptive = cv2.adaptiveThreshold(smoothed, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
cv2.THRESH_BINARY_INV, 21, 16)
clean = cv2.morphologyEx(adaptive, cv2.MORPH_OPEN, np.ones((3, 3), np.uint8))
print("overlap with the true ink (1.0 would be perfect):")
for name, mask in [("one fixed cut at 128", fixed),
("Otsu, one cut chosen for the page", otsu_img > 0),
("adaptive, a cut per neighbourhood", adaptive > 0),
("adaptive + median blur + opening", clean > 0)]:
print(f" {name:36s} {overlap(mask, gt):.3f}")
# Two ways to measure the skew, both reading it off the ink itself.
ys, xs = np.nonzero(clean)
box_angle = cv2.minAreaRect(np.column_stack([xs, ys]).astype(np.float32))[2]
box_angle = box_angle if box_angle < 45 else box_angle - 90
scores = []
for candidate in np.arange(-15, 15.01, 0.25):
R = cv2.getRotationMatrix2D((230, 60), candidate, 1.0)
rotated = cv2.warpAffine(clean, R, (460, 120), flags=cv2.INTER_NEAREST)
scores.append((rotated.sum(axis=1).var(), candidate)) # straight text = spiky profile
profile_angle = max(scores)[1]
print(f"\nskew applied when the page was scanned: {TRUE_SKEW:+.2f} deg")
print(f" measured by the tightest rotated box: {box_angle:+.2f} deg")
print(f" measured by the row-projection search: {profile_angle:+.2f} deg")
R = cv2.getRotationMatrix2D((230, 60), profile_angle, 1.0)
straight = cv2.warpAffine(clean, R, (460, 120), flags=cv2.INTER_NEAREST)
print("\nimage rows containing ink:",
((clean > 0).sum(axis=1) > 0).sum(), "crooked ->",
((straight > 0).sum(axis=1) > 0).sum(), "straight")paper grey level: left 200 right 103 ink grey level: left 41 right 27 ink is 7.5 percent of the pixels on this page Otsu chose a single cut for the whole page at 152 overlap with the true ink (1.0 would be perfect): one fixed cut at 128 0.310 Otsu, one cut chosen for the page 0.194 adaptive, a cut per neighbourhood 0.820 adaptive + median blur + opening 0.855 skew applied when the page was scanned: +6.00 deg measured by the tightest rotated box: -6.82 deg measured by the row-projection search: -6.00 deg image rows containing ink: 101 crooked -> 46 straight
Reading the output carefully
Look at the four grey levels first. Paper runs from 200 down to 103 across the page. Ink runs from 41 down to 27. The lamp changes the paper by 97 levels, which is far more than the 14 levels separating bright ink from dim ink.
Otsu picked 152, and that is worse than picking nothing. Otsu's rule looks for two peaks in the brightness histogram and cuts between them. Here ink is 7.5% of the pixels, so the dominant structure in the histogram is bright paper versus dim paper — and the cut lands there. Everything below 152 is called ink, which includes the whole right-hand side of the page. Score 0.194.
The fixed cut at 128 scores 0.310, better than the automatic one. That ordering is worth sitting with. An automatic method is not automatically better; it is better only when its assumption holds.
Adaptive thresholding jumps to 0.820. cv2.adaptiveThreshold with blockSize=21 compares each pixel to a Gaussian-weighted average of the 21x21 pixels around it, minus C=16. A shadow moves the pixel and its neighbourhood together, so it cancels out. This is the single most valuable line in the file.
Median blur and morphological opening add 0.035 more. medianBlur replaces each pixel with the median of its 3x3 patch, which erases single-pixel speckle without blurring edges. MORPH_OPEN is an erosion then a dilation, which deletes blobs too small to survive a 3x3 erosion. Neither touches the strokes, which are 5 pixels wide.
The two skew estimates disagree, and the projection search wins. minAreaRect returns -6.82, off by nearly a degree, because the tightest rotated rectangle is dragged by leftover specks and by the ragged ends of the text. The projection search returns -6.00 exactly, because it optimises the property we actually want: ink packed into few rows. Row count falls from 101 to 46.
The sign flip is not a bug. The page was rotated by +6, so undoing it needs -6.
Choosing blockSize and C
These two numbers are the whole game, and they are not universal.
blockSizemust be comfortably larger than a stroke and smaller than the shadow. If it is smaller than a stroke, the middle of a thick letter looks like its own background and hollows out. Odd numbers only.Cis a constant subtracted from the local mean. Raise it to demand more contrast, which removes noise and thins letters. Lower it to keep faint ink and admit speckle.
At 300 DPI scans of body text, blockSize in the 15 to 35 range is a sensible starting bracket. Sweep both on your own worst pages and score them, exactly as the loop above does.
Common mistakes
Binarising before deskewing, then deskewing with interpolation. Rotating a 0/1 image with bilinear interpolation invents grey values along every stroke edge. Use cv2.INTER_NEAREST on masks, or deskew the greyscale image and binarise afterwards.
Upscaling a small image with plain resizing. Tesseract wants roughly 30 pixels of x-height. If your capital letters are 12 pixels tall, cv2.resize with INTER_CUBIC before binarising helps more than any threshold tuning.
Aggressive denoising on Indic scripts. Opening with a 3x3 kernel removes speckle and also removes the dot of a nukta or a small matra. Test on Devanagari or Tamil pages before shipping a kernel size.
Deskewing a page with more than one text block. A single global angle is wrong for a photograph of a curved book spine. Those need per-line dewarping, not a rotation.
Passing an already-binarised image to a modern engine. PaddleOCR and TrOCR were trained on natural greyscale or colour crops. Hard binarisation can lower their accuracy, while it usually raises Tesseract's. Measure, per engine.
Try it yourself
Change blockSize from 21 to 5 and rerun. Watch the overlap collapse. Print the mask for a single letter and you will see it hollowed into an outline, because a 5x5 neighbourhood inside a 5-pixel-wide stroke is entirely ink.
What to learn next
- Text detection models — the stage that consumes this cleaned image.
- Tesseract in practice — the engine that gains the most from good preprocessing.
- Homographies and perspective warp — flattening a page photographed at an angle.
Researcher — Mathematics and papers.
Binarisation as a decision rule
Global thresholding assigns pixel $I(x,y)$ to ink when $I(x,y) < t$. Otsu (1979) chooses $t$ by maximising between-class variance:
$$ t^{*} = \arg\max_{t} \; \omega_0(t)\,\omega_1(t)\,\bigl(\mu_0(t) - \mu_1(t)\bigr)^2 $$
Where $\omega_0, \omega_1$ are the pixel fractions below and above $t$, and $\mu_0, \mu_1$ their mean intensities. The criterion is exactly equivalent to minimising within-class variance.
Two assumptions are baked in: the histogram is bimodal, and the classes are of comparable mass. Document images violate both — ink is typically 3–10% of pixels, and illumination widens the paper mode until it dominates. The failure in the code above is not an implementation artefact; it is the assumption being false.
Local methods
Niblack (1986) replaces the global constant with a per-window statistic:
$$ t(x,y) = \mu_w(x,y) + k \, \sigma_w(x,y) $$
Where $\mu_w$ and $\sigma_w$ are mean and standard deviation over a window $w$, and $k \approx -0.2$. Niblack is noisy in blank regions, where $\sigma_w \to 0$ and the threshold clings to the mean.
Sauvola and Pietikäinen (2000) fix that by scaling the correction with contrast:
$$ t(x,y) = \mu_w(x,y)\left[1 + k\left(\frac{\sigma_w(x,y)}{R} - 1\right)\right] $$
Where $R$ is the dynamic range of $\sigma$ (128 for 8-bit) and $k \approx 0.2$–$0.5$. In blank regions $\sigma_w \approx 0$, so $t \to \mu_w(1-k) < \mu_w$, and nothing is marked as ink. Sauvola remains the default for degraded historical documents; scikit-image ships it as threshold_sauvola.
OpenCV's adaptiveThreshold implements the simpler $t = \mu_w - C$, which is Sauvola with the contrast term removed. It is fast and adequate for evenly-inked modern documents.
Learned binarisation now beats all of the above on damaged manuscripts. The DIBCO / H-DIBCO competition series (2009–2019) is the benchmark, and U-Net style segmentation models and conditional GANs have topped it since roughly 2017. The relevant metrics are F-measure, pseudo-F-measure, PSNR and DRD (distance reciprocal distortion).
Skew estimation
Three families, with different cost and different failure modes.
Projection profile (Postl, 1986). Define $p_\theta(y)$ as the ink count in row $y$ after rotating by $\theta$. Maximise a spikiness objective, commonly
$$ E(\theta) = \sum_y p_\theta(y)^2 \quad \text{or} \quad \operatorname{Var}y\bigl[p\theta(y)\bigr] $$
Cost is $O(n_\theta \cdot HW)$, which is why coarse-to-fine search is standard. It is the most accurate method on ordinary text pages and the one used above.
Hough transform on the binarised image, taking the dominant line orientation. Robust to sparse text, expensive, and sensitive to ruled lines and table borders — which frequently outvote the text.
Principal component or minimum-area-rectangle on ink coordinates. Effectively $O(n)$ and appealing, but it estimates the orientation of the point cloud, not of the baselines. A single block of text whose aspect ratio is near square gives a nearly arbitrary answer, and outliers dominate. The 0.82-degree error in the output above is this effect in miniature.
Angles beyond roughly $\pm 15^\circ$ need a different treatment. Photographed pages suffer perspective, not rotation, and the correct fix is a homography from the detected page quadrilateral — see homographies and perspective warp. Curved book pages need dewarping models such as DewarpNet (Das et al., ICCV 2019) or DocTr (Feng et al., ACM MM 2021).
Does preprocessing still pay?
For classical engines, decisively yes. Tesseract 5 with LSTM still expects near-binary input and documents a preferred x-height around 30 pixels.
For modern neural recognisers the answer inverts. CRNN and TrOCR recognisers are trained on augmented natural crops, including blur, low contrast and shadow. Hard binarisation discards intensity information those models use, and published ablations on scene-text benchmarks generally show it hurting. The reliable modern preprocessing set is geometric (deskew, dewarp, crop, resample), not photometric.
The one photometric step that survives everywhere is resolution. No architecture recovers strokes that were never sampled.
What to learn next
- Text detection models — the stage that consumes this cleaned image.
- Tesseract in practice — the engine that gains the most from good preprocessing.
- Homographies and perspective warp — flattening a page photographed at an angle.