Evaluating Vision Models

Measuring OCR accuracy

Character and word error rates tell you how close the text is, and neither of them tells you whether the invoice number was right.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. Counting edits
  4. The surprise about word error rate
  5. Cleaning up before comparing
  6. The number that actually matters
  7. Where you have seen this
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

OCR accuracy is measured by how many single-character edits it takes to turn the model's text into the correct text.

The analogy you have lived

Someone reads a phone number aloud and you write it on a slip of paper. Ten digits. You get nine of them right.

Ninety percent. It sounds respectable. The call does not connect.

The number is not nine-tenths useful. It is completely useless, and no averaging over digits will change that. This gap between "mostly right" and "actually works" is the entire subject of this lesson.

Counting edits

Start with the honest small measurement. Put the true text beside what the model read. Count the smallest number of single-character changes that fix it.

There are three kinds of change: insert a character, delete a character, or replace one with another.

   truth       G S T I N
   model read  G S T l N
                     ^
   one replacement -> one edit

That count, divided by the length of the true text, is the character error rate, always written CER. Zero is perfect. Higher is worse.

Do the same thing treating whole words as the units, and you get the word error rate, or WER.

The surprise about word error rate

WER can be greater than one. That sounds like a bug and it is not.

Suppose the true text is one word: a date written without spaces. The model reads it with a space after every character. The model produced nine words where there was one. Fixing it takes more edits than the original had words.

An error rate above one means the output has more junk in it than the original had content. It happens constantly with OCR, because deciding where a word ends is part of the problem.

Cleaning up before comparing

Should uppercase count as an error? What about a full stop the model missed, or two spaces instead of one?

For reading a scanned book, no. For reading a bank account number, absolutely yes.

So most teams normalise before comparing: lowercase everything, strip punctuation, squeeze runs of spaces. That is a reasonable choice as long as you say you made it.

Watch out for something sneaky. Normalising also shortens the true text, and the true text is what you divide by. A smaller divisor can push the error rate up even though you removed errors. Cleaning is not a guaranteed win.

The number that actually matters

Now go back to the phone number.

Almost nobody wants a character error rate. They want the invoice number, the total, the date, the tax id. Each of those is right or it is wrong. There is no partial credit for a total that is off by one digit.

So the metric that matters is field-level exact match. What fraction of the fields came out exactly right?

It is always a harsher number than CER, and it is the one your users experience.

   character error rate   0.07     "looks good"
   field exact match      0.50     "half the invoices are wrong"
   same model, same page

Where you have seen this

  • Depositing a cheque by photographing it in a banking app.
  • Uploading an ID document and having the fields fill in for you.
  • A parking gate reading your number plate.
  • Scanning a receipt to file an expense claim.

The parking gate is the clearest case. One wrong character and the wrong person is billed. Nobody in that system cares about character error rate.

The honest part

There is a second measurement problem hiding underneath, and it has nothing to do with the model.

Somebody had to type the correct answer for every test page. That person also makes mistakes, and their mistakes are counted against the model.

On real documents, a slice of your reported errors will be wrong ground truth. Before you spend a month improving a model, read fifty of its worst pages yourself. It is common to find the model was right and the reference was not.

Remember this

  • CER counts single-character edits. WER counts whole-word edits, and can go above one.
  • Normalising the text changes the divisor too, so it can make the score look worse.
  • The number your users feel is field-level exact match, not CER.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install --upgrade pip     # the example below needs nothing else

The standard library is enough. jiwer and torchmetrics both implement these, but writing the twelve lines once is what stops you misreading the results forever.

Edit distance, CER, WER and field match

ocr_metrics.py
import re, unicodedata

PAIRS = [
    ("Invoice No: INV-2024/0917",   "Invoice No: lNV-2024/0917"),
    ("Total: Rs. 14,250.00",        "Total: Rs. 14,25O.OO"),
    ("Shri Ramesh Kulkarni",        "Shri Ramesh Kulkarni"),
    ("GSTIN 27AABCU9603R1ZM",       "GSTlN 27AABCU96O3R1ZM"),
    ("Date 07/11/2025",             "Date 0 7 / 1 1 / 2 0 2 5"),
]


def edit_distance(a, b):
    """Levenshtein: how many insert/delete/substitute steps turn a into b."""
    prev = list(range(len(b) + 1))
    for i, ca in enumerate(a, 1):
        cur = [i]
        for j, cb in enumerate(b, 1):
            cur.append(min(prev[j] + 1,          # delete a character
                           cur[j - 1] + 1,       # insert a character
                           prev[j - 1] + (ca != cb)))   # substitute
        prev = cur
    return prev[-1]


def cer(ref, hyp):
    return edit_distance(ref, hyp) / max(len(ref), 1)


def wer(ref, hyp):
    r, h = ref.split(), hyp.split()
    return edit_distance(r, h) / max(len(r), 1)


print(f"{'CER':>6} {'WER':>6}  truth -> prediction")
for ref, hyp in PAIRS:
    print(f"{cer(ref, hyp):>6.3f} {wer(ref, hyp):>6.3f}  {ref!r} -> {hyp!r}")

total_chars = sum(len(r) for r, _ in PAIRS)
total_errs = sum(edit_distance(r, h) for r, h in PAIRS)
print(f"\ncorpus CER = {total_errs}/{total_chars} = {total_errs/total_chars:.4f}")
print(f"average-of-lines CER = {sum(cer(r,h) for r,h in PAIRS)/len(PAIRS):.4f}")


def normalise(s):
    s = unicodedata.normalize("NFKC", s)          # fold look-alike unicode forms
    s = s.lower()
    s = re.sub(r"[^\w\s]", "", s)                 # drop punctuation
    s = re.sub(r"\s+", " ", s).strip()            # squeeze runs of whitespace
    return s


print(f"\n{'raw CER':>8} {'normalised CER':>15}  line")
for ref, hyp in PAIRS:
    print(f"{cer(ref, hyp):>8.3f} {cer(normalise(ref), normalise(hyp)):>15.3f}  {ref!r}")

# What a document pipeline actually gets paid for: whole fields, right or wrong.
FIELDS = {"invoice_no": ("INV-2024/0917", "lNV-2024/0917"),
          "total":      ("14250.00",      "14250.00"),
          "gstin":      ("27AABCU9603R1ZM", "27AABCU96O3R1ZM"),
          "date":       ("07/11/2025",    "07/11/2025")}
exact = sum(1 for k, (r, h) in FIELDS.items() if r == h)
print(f"\nfield-level exact match: {exact}/{len(FIELDS)} = {exact/len(FIELDS):.2f}")
for k, (r, h) in FIELDS.items():
    print(f"  {k:<11} {'OK ' if r == h else 'BAD'}  {r} -> {h}   (CER {cer(r,h):.3f})")
Output
   CER    WER  truth -> prediction
 0.040  0.333  'Invoice No: INV-2024/0917' -> 'Invoice No: lNV-2024/0917'
 0.150  0.333  'Total: Rs. 14,250.00' -> 'Total: Rs. 14,25O.OO'
 0.000  0.000  'Shri Ramesh Kulkarni' -> 'Shri Ramesh Kulkarni'
 0.095  1.000  'GSTIN 27AABCU9603R1ZM' -> 'GSTlN 27AABCU96O3R1ZM'
 0.600  5.000  'Date 07/11/2025' -> 'Date 0 7 / 1 1 / 2 0 2 5'

corpus CER = 15/101 = 0.1485
average-of-lines CER = 0.1770

 raw CER  normalised CER  line
   0.040           0.045  'Invoice No: INV-2024/0917'
   0.150           0.188  'Total: Rs. 14,250.00'
   0.000           0.000  'Shri Ramesh Kulkarni'
   0.095           0.095  'GSTIN 27AABCU9603R1ZM'
   0.600           0.538  'Date 07/11/2025'

field-level exact match: 2/4 = 0.50
  invoice_no  BAD  INV-2024/0917 -> lNV-2024/0917   (CER 0.077)
  total       OK   14250.00 -> 14250.00   (CER 0.000)
  gstin       BAD  27AABCU9603R1ZM -> 27AABCU96O3R1ZM   (CER 0.067)
  date        OK   07/11/2025 -> 07/11/2025   (CER 0.000)

Reading the output carefully

A WER of 5.000 is not a bug. The date line has three true words. The model produced ten tokens by inserting spaces between digits, and repairing that takes fifteen word-level edits. WER above one means the output contains more error than the reference contained content. Any OCR evaluation that clips WER at 1.0 is hiding this failure mode.

Line four shows why you need both. CER is a mild 0.095 while WER is a flat 1.000. Two characters are wrong in a twenty-one character tax id, but both wrong characters land in the same long token, so every word is wrong. For identifier-shaped text, WER is close to useless and CER is close to meaningless.

Corpus CER 0.1485, average-of-lines CER 0.1770. Same data, two aggregation rules, and a difference of three points. Corpus CER pools all edits over all characters, which weights long lines more. Averaging per line weights every line equally, so one short broken line dominates. Say which one you used; papers report both and rarely label them.

Normalisation made three lines look worse. Line one went from 0.040 to 0.045 and line two from 0.150 to 0.188. Stripping punctuation removed characters from the reference, shrinking the denominator faster than it removed errors. Cleaning your text is a policy decision, not an improvement.

Field exact match is 0.50 against a corpus CER of 0.1485. Two of the four fields are wrong, each by a single character. This is the number the business feels, and it is four times harsher than the character number that looked fine.

The two failures are the classic OCR confusions: capital I read as lowercase l, and digit 0 read as capital O. Neither is random. Both are fixable with a per-field character whitelist, since a tax id cannot contain a lowercase letter and a total cannot contain O.

Constrain the field, do not blame the model

The highest-value change in most document pipelines is not a better recogniser. It is telling the pipeline what each field is allowed to look like. But there are two strengths of constraint, and the weaker one is not enough.

field_repair.py
# A GSTIN is 15 characters with a fixed shape: 2 digits, 5 letters, 4 digits,
# 1 letter, then 3 characters that may be either.  d = digit, a = letter.
GSTIN_SHAPE = "ddaaaaadddda???"
TO_DIGIT = {"O": "0", "o": "0", "l": "1", "I": "1", "S": "5", "B": "8"}
TO_LETTER = {"0": "O", "1": "I", "5": "S", "8": "B"}


def repair_shape(text, shape):
    out = []
    for want, c in zip(shape, text):
        if want == "d" and not c.isdigit():
            c = TO_DIGIT.get(c, c)
        elif want == "a" and not c.isalpha():
            c = TO_LETTER.get(c, c)
        out.append(c)
    return "".join(out)


def repair_numeric(text, allowed=set("0123456789.,")):
    return "".join(c if c in allowed else TO_DIGIT.get(c, c) for c in text)


alpha = set("0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ")
naive = "".join(c if c in alpha else TO_DIGIT.get(c, c) for c in "27AABCU96O3R1ZM")

print("gstin, alphabet filter only :", naive, "  correct:", naive == "27AABCU9603R1ZM")
fixed = repair_shape("27AABCU96O3R1ZM", GSTIN_SHAPE)
print("gstin, positional shape     :", fixed, "  correct:", fixed == "27AABCU9603R1ZM")
tot = repair_numeric("14,25O.OO")
print("total, numeric alphabet     :", tot, "       correct:", tot == "14,250.00")
Output
gstin, alphabet filter only : 27AABCU96O3R1ZM   correct: False
gstin, positional shape     : 27AABCU9603R1ZM   correct: True
total, numeric alphabet     : 14,250.00        correct: True

The alphabet filter fixed the total and did nothing for the tax id. A total may contain only digits and separators, so O has nowhere to hide and becomes 0. A GSTIN may contain both letters and digits, so O is a legal character in that field and the filter leaves it alone.

Only the positional shape fixes it. Character nine of a GSTIN must be a digit. Knowing that turns an ambiguous glyph into an unambiguous one. Field exact match goes from 0.50 to 1.00 with no model change, but the weak constraint would have got you halfway and left you puzzled.

The general rule is worth carrying: alphabet constraints help only where the alphabets of the confusable characters do not overlap. Where they do, you need structure — a positional shape, a checksum, or a lookup against a list of valid values.

Common mistakes

Reporting accuracy as 1 - CER. Since CER can exceed 1, this produces negative "accuracy". Report the error rate and let it be a rate.

Evaluating detection and recognition together and calling it recognition. Modern OCR is two stages: find the text regions, then read each one. If the detector missed a line, the recogniser never saw it. Score them separately, or you will tune the wrong stage. The same reasoning as detection error analysis.

Ignoring reading order. A page read correctly but assembled in the wrong order has a low CER on each line and produces a nonsense document. Multi-column layouts and tables fail here constantly, and no character metric will tell you.

Comparing across scripts without checking normalisation. Devanagari and other Indic scripts have multiple valid Unicode encodings of the same visible text. Without NFC or NFKC normalisation, identical-looking strings compare as different, and your Hindi CER will look terrible for no reason.

Assuming a language model always helps. Decoding with a dictionary fixes ordinary words and destroys identifiers, names and product codes — the exact fields you are being paid to read.

Try it yourself

Add a line where the model reads a correct total but formats it as 14250.00 against a reference of 14,250.00. Watch CER report a failure that a human would call a success. Then decide, and write down, whether your field comparison should normalise separators. That decision is the metric.

What to learn next

Researcher — Mathematics and papers.

Levenshtein distance

For strings $a$ and $b$, the edit distance $D(i,j)$ over prefixes satisfies

$$ D(i,j) = \min \begin{cases} D(i-1,j) + 1 & \text{deletion} \ D(i,j-1) + 1 & \text{insertion} \ D(i-1,j-1) + \mathbb{1}[a_i \ne b_j] & \text{substitution} \end{cases} $$

with $D(i,0) = i$ and $D(0,j) = j$. Cost is $O(|a||b|)$ time and $O(\min(|a|,|b|))$ space with the rolling-row form used above. Then

$$ \mathrm{CER} = \frac{S + D + I}{N}, \qquad \mathrm{WER} = \frac{S_w + D_w + I_w}{N_w} $$

where $S, D, I$ are substitution, deletion and insertion counts on the optimal alignment and $N$ is the reference length. Because $I$ is unbounded relative to $N$, both rates are unbounded above. This is inherited directly from speech recognition, where WER was standardised by NIST.

Damerau-Levenshtein adds transposition as a unit-cost operation. It fits typing errors well and OCR errors poorly, since OCR substitutes visually similar glyphs rather than swapping adjacent ones.

Aggregation

$$ \mathrm{CER}_{\text{corpus}} = \frac{\sum_k e_k}{\sum_k N_k} \qquad\text{versus}\qquad \mathrm{CER}_{\text{macro}} = \frac{1}{K}\sum_k \frac{e_k}{N_k} $$

These differ whenever line lengths vary, and the macro form has unbounded variance from short lines. ICDAR competitions report the corpus form. Most in-house dashboards report the macro form without labelling it.

Beyond flat text

Real document evaluation needs three further families.

Detection. Text-line and word boxes are scored with IoU-based precision, recall and H-mean, historically at IoU 0.5 (ICDAR 2013/2015). The DetEval protocol additionally handles one-to-many and many-to-one matches, which flat IoU cannot, because a single ground-truth line legitimately splits into several predicted words.

End-to-end spotting. Detection and recognition scored jointly: a detection counts only if its transcription is also correct, usually under a normalised comparison. ICDAR 2015 "end-to-end" and Total-Text use this. It is the closest published analogue of field exact match.

Structure. For tables and forms, TEDS (Tree-Edit-Distance-based Similarity, Zhong et al., 2020, PubTabNet) computes edit distance over the HTML tree rather than over the string, so a cell placed in the wrong column is penalised even when its text is perfect. GriTS (Smock et al., 2023) refines this with a two-dimensional alignment.

Key-information extraction

For the field-level task, the standard reporting is per-field precision, recall and $F_1$ over exact string match after a declared normalisation, as used by the FUNSD, CORD and SROIE benchmarks. Two protocol choices dominate the reported numbers and are frequently unstated:

  • Whether matching is exact or fuzzy. SROIE's official script is exact after whitespace stripping. Papers reporting fuzzy match are not comparable.
  • How multi-value fields are scored. A line-items table with ten rows can be scored as one field or ten, changing the headline by tens of points.

Confidence and selective prediction

Character-level posterior probabilities from CTC or attention decoders are typically overconfident, but their ranking is usable. The practical construction is a risk-coverage curve: sort fields by minimum character confidence, route the lowest fraction to a human, and plot error against the fraction automated. El-Yaniv and Wiener (2010) give the selective-classification framing. In document processing this curve, not CER, is what determines cost per document, and it is the correct thing to optimise.

CTC posteriors also carry a specific pathology: the blank symbol absorbs probability mass, so per-character confidences are not comparable across sequence lengths. Normalising by the number of emitted characters before thresholding is necessary and often skipped.

Ground-truth error

Reported CER floors on historical-document benchmarks are frequently within the range of transcription disagreement between human annotators. Clausner et al. (2020) and the successive ICDAR competition reports document ground-truth revisions between editions of the same dataset. The correct protocol is double transcription of a held-out slice, reporting inter-annotator CER as the measurement floor alongside the model score. Almost nobody does this, which is why sub-one-percent CER claims on messy scans should be read sceptically.

Papers

  • Levenshtein, Binary codes capable of correcting deletions, insertions, and reversals, 1966
  • Karatzas et al., ICDAR 2015 Competition on Robust Reading — rrc.cvc.uab.es
  • Zhong et al., Image-based Table Recognition (PubTabNet, TEDS), 2020 — arxiv.org/abs/1911.10683
  • Smock et al., GriTS: Grid Table Similarity Metric, 2023 — arxiv.org/abs/2203.12555
  • Huang et al., ICDAR 2019 Competition on Scanned Receipt OCR and Information Extraction (SROIE)
  • El-Yaniv and Wiener, On the Foundations of Noise-free Selective Classification, JMLR 2010

What to learn next