Messy Real-World Text

Punctuating a speech transcript

Speech-to-text output has no punctuation or capitals, and restoring them means finding sentence boundaries from cues other than punctuation itself.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Restoring punctuation and casing takes a flat wall of speech-to-text output and turns it back into readable sentences.

Think about a text message with no punctuation: "can you call me back im at the station waiting". You can still work out where one thought ends and the next begins. You use the words, and the pauses you imagine.

Speech-to-text engines routinely output text exactly like that: no periods, no capitals, no commas. Restoring them is what turns raw transcript text into something a person can comfortably read.

Why it exists

Many speech-to-text systems focus entirely on getting the words right. Punctuation and capitalisation were spoken as pauses and emphasis, not as symbols. Nothing in the audio directly says "insert a period here."

Meeting transcripts, voicemail-to-text, and auto-generated video captions all produce this same flat, unpunctuated text. Restoring structure back into it is a separate step, run after the words are already known.

How it works

  "so the first thing we need to do is open the file then we check the headers"
       |
       v
  find likely sentence-boundary cues (pause words: "so", "then", "after")
       |
       v
  "So the first thing we need to do is open the file. Then we check the headers."

Where you have already seen it

  • Auto-generated meeting notes, cleaned up from a raw transcript into readable paragraphs.
  • YouTube auto-captions, which start as flat text before any punctuation gets added.
  • Voice assistant transcripts, shown as a message history after you spoke a command aloud.

Remember this

  • Speech-to-text output typically has no punctuation or capitalisation at all — that structure has to be added back separately.
  • Simple cue-word rules can approximate sentence boundaries, but they are a rough heuristic, not a reliable general solution.
  • Real systems treat this as a labelling task, deciding for each word whether a punctuation mark follows it.

What to learn next

Developer — Code and libraries.

Setup

Nothing to install. Pure Python standard library.

A rough, rule-based punctuation restorer

restore_punctuation.py
# A transcript straight off a speech-to-text engine: no punctuation, no capitals.
transcript = "so the first thing we need to do is open the file then we check the headers after that we run the script and if it fails we look at the log"

# Cue words that usually start a new clause in spoken English.
BOUNDARY_WORDS = {"so", "then", "after", "now", "but", "and", "because"}

def restore_punctuation(text):
    words = text.split()
    sentences = []
    current = []
    for w in words:
        if current and w in BOUNDARY_WORDS and len(current) > 3:
            sentences.append(current)
            current = []
        current.append(w)
    if current:
        sentences.append(current)

    out = []
    for sent in sentences:
        sent[0] = sent[0].capitalize()
        out.append(" ".join(sent) + ".")
    return " ".join(out)

print("BEFORE:", transcript)
print()
print("AFTER: ", restore_punctuation(transcript))
Output
BEFORE: so the first thing we need to do is open the file then we check the headers after that we run the script and if it fails we look at the log

AFTER:  So the first thing we need to do is open the file. Then we check the headers. After that we run the script. And if it fails we look at the log.

Line by line

BOUNDARY_WORDS is a small, hand-picked list of words that often start a new clause in spoken English. This is a genuine heuristic, not a rule with any formal grounding — it happens to work reasonably on this particular transcript because the speaker's clauses really do start with these words.

len(current) > 3 stops the function from splitting on every single boundary word. Without it, two boundary words appearing close together would produce a fragment sentence far too short to be meaningful.

Every resulting sentence gets its first word capitalised and a period appended. This handles casing and one punctuation mark. Real speech-to-text punctuation also needs commas, question marks, and exclamation marks — none of which this simple version attempts.

Common mistakes

Trusting cue-word heuristics on text without natural cue words. This transcript happened to use "so," "then," "after" and "and" at real clause boundaries. A transcript that says "we open the file we check the headers we run the script," with no cue words at all, gets none of its boundaries found by this method.

Assuming every sentence ends in a period. Real speech is full of questions and exclamations. A rule-based restorer with only period-insertion will mark a spoken question as a flat statement, changing its meaning to a reader.

Deploying this kind of heuristic in production without disclosure. A rough restoration is genuinely useful for a first pass, but should not be presented as equivalent to a careful transcript — meetings with legal or medical stakes need a stronger method, or human review.

Try it yourself

Feed the function a transcript with no boundary words at all — for example, a run of short declarative statements strung together with no connecting words. Watch it fail to split anything, returning one giant sentence, and consider what additional signal — a pause duration from the speech-to-text engine's timestamps, for instance — could help where word-based cues cannot.

What to learn next

  • Splitting text into sentences — segmenting text that already has punctuation, a related but different problem.
  • Named entity recognition — another per-word labelling task, useful for comparing labelling-based approaches.
  • BERT — the kind of model a real punctuation-restoration system is usually built on.

Researcher — Mathematics and papers.

Framing punctuation restoration as sequence labelling

Real punctuation restoration systems treat this as token classification, not clause-boundary heuristics: for each word w_i in the unpunctuated transcript, predict a label y_i from a small punctuation-class vocabulary.

text
y_i in { NONE, PERIOD, COMMA, QUESTION_MARK, EXCLAMATION_MARK }
P(y_1, ..., y_n | w_1, ..., w_n) = product over i of P(y_i | w_1, ..., w_n)
  • The label is predicted per word, conditioned on the full surrounding context — typically via a bidirectional encoder such as BERT, reading both preceding and following words, since punctuation placement depends on what comes after a word as much as what comes before it.
  • Capitalisation is usually modelled as a second, related labelling task — TRUE_CASE versus LOWER_CASE per word — either jointly with punctuation or as a separate pass.

This is structurally identical to the token-classification framing used for named entity recognition: one label per token, predicted jointly across the sequence rather than by independent per-word rules.

Why cue-word heuristics are a weak baseline

The developer demo's rule set only fires on a fixed, small vocabulary of boundary words, giving it zero recall on any sentence boundary not marked by one of them — a hard ceiling no amount of tuning the rule list removes entirely, since natural spoken language does not reliably mark every clause boundary with a recognisable cue word.

A trained sequence-labelling model instead learns boundary signal from the full range of syntactic and lexical context — subject-verb patterns that typically start a new sentence, prosodic proxies where available (pause duration, if the speech-to-text system exposes word-level timestamps), and statistical regularities no fixed word list could enumerate by hand.

Using timing information from the ASR system

Where the upstream automatic speech recognition (ASR) system exposes word-level timestamps, pause duration between words is a strong, independent signal for sentence and clause boundaries — genuinely absent from the text-only heuristic in the developer demo. Production systems commonly fuse acoustic pause features with the text-based labelling model above, since the two signals are complementary: text captures syntactic structure, timing captures the prosodic boundary the speaker actually produced.

Key references

  • Tilk, O. & Alumäe, T. (2016). Bidirectional Recurrent Neural Network with Attention Mechanism for Punctuation Restoration. Interspeech. — an early neural sequence-labelling approach to this task.
  • Nagy, A. et al. (2021). Automatic Punctuation Restoration with BERT Models. — bidirectional transformer-based punctuation restoration, the current standard architecture pattern.
  • Chen, Q. et al. (2020). Discriminative Self-Training for Punctuation Prediction. Interspeech. — addresses the scarcity of clean training data for this task through self-training on unlabelled transcripts.

Current state and open problems

Fine-tuned transformer models are now the standard approach for production punctuation and casing restoration, trained on large corpora of (unpunctuated transcript, correctly punctuated original) pairs, cheaply constructed by stripping punctuation from any already-correctly-punctuated text corpus — a form of self-supervised label generation requiring no manual annotation.

Genuine ambiguity remains at true fork points — a spoken sentence that a human transcriber could reasonably punctuate two different ways depending on intended meaning, with no acoustic or lexical signal resolving it either way. No current system, rule-based or learned, resolves ambiguity that was not resolved by the speaker's own delivery in the first place; this is a hard ceiling on the task, not a current model limitation.

What to learn next

  • Named entity recognition — the closest published lesson to this task's real token-classification framing.
  • BERT — the architecture family production punctuation restoration is typically built from.
  • Splitting text into sentences — the inverse-direction problem, over text that already carries punctuation.