Text Preprocessing

Splitting text into sentences

Splitting on full stops shreds "Dr. Rao" and "U.S." into fragments — real sentence segmentation has to know which dots actually end a sentence.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Sentence segmentation cuts a wall of text into its individual sentences.

The hard part is that a full stop does not always mean "sentence over".

Think of cutting a long roll of fabric into shirts. You cannot cut at every printed dot on the cloth; some dots are part of the pattern. Cut blindly at every dot, and you shred the shirts. The tailor first learns which marks are real cutting lines.

Why it exists

Many text jobs work one sentence at a time. Translation systems translate sentence by sentence. Summarisers pick the most important sentences. Systems that feed documents to chatbots split them into pieces first — and sentence boundaries are the natural seams.

The tempting rule — "cut at every full stop" — fails immediately. English dots do triple duty:

  • Abbreviations: "Dr. Rao", "U.S.", "etc."
  • Numbers: "₹4.50", "version 3.2"
  • Actual sentence endings.

A sentence segmenter is a tool that tells these apart, using clues like the next word's capitalisation and a memory of common abbreviations.

How it works

"Dr. Rao moved to the U.S. in 1998. He paid $4.50 for coffee."

cut at every dot:      ["Dr", "Rao moved to the U", "S", "in 1998", ...]   ✗ shredded
proper segmenter:      ["Dr. Rao moved to the U.S. in 1998.",
                        "He paid $4.50 for coffee."]                        ✓ two sentences

The segmenter walks the text dot by dot, asking: is this dot inside a known abbreviation? Is it followed by a capital letter starting a plausible new sentence? Is it part of a number? Each answer nudges the cut/no-cut decision.

A real example you have seen

Subtitles on streaming apps arrive one sentence (or clause) at a time, cut from a full transcript. Voice assistants reading news aloud pause at sentence ends — a wrong split makes the voice stop mid-name: "Dr." (pause) "Rao". When translation apps translate a paragraph, the per-sentence pipeline underneath depends on these cuts being right.

Remember this

  • Sentence segmentation = cutting text at real sentence ends, not at every dot.
  • Abbreviations, initials and decimal numbers are the classic traps.
  • Use a proper segmenter; the one-line split is nearly always wrong on real text.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install spacy
python -m spacy download en_core_web_sm

The model is a ~12 MB one-time download. Verified with spaCy 3.8, CPU only.

The naive split, caught in the act

sentences.py
import spacy

text = ("Dr. Rao moved to the U.S. in 1998. He paid $4.50 for coffee. "
        "Was it worth it? He thinks so.")

# the tempting one-liner, and why it fails
naive = [s.strip() for s in text.split(".") if s.strip()]
print("naive split on '.':", len(naive), "pieces")
for s in naive[:4]:
    print("  ", repr(s))

nlp = spacy.load("en_core_web_sm")
sents = [s.text for s in nlp(text).sents]
print("\nspaCy:", len(sents), "sentences")
for s in sents:
    print("  ", repr(s))
Output
naive split on '.': 7 pieces
   'Dr'
   'Rao moved to the U'
   'S'
   'in 1998'

spaCy: 4 sentences
   'Dr. Rao moved to the U.S. in 1998.'
   'He paid $4.50 for coffee.'
   'Was it worth it?'
   'He thinks so.'

The walkthrough

Seven shreds versus four sentences. The naive split cut inside "Dr.", twice inside "U.S.", and inside "$4.50" — and it silently deleted every question mark and full stop, so the pieces cannot even be glued back. str.split throws the separator away.

.sents comes from the parser. In en_core_web_sm, sentence boundaries fall out of the dependency parse — the model understands enough grammar to know "Rao" continues the "Dr." sentence. That grammar knowledge is what you are downloading in those 12 MB.

Need speed without the full parse? spaCy ships a lightweight statistical splitter you can run instead:

senter.py
fast = spacy.load("en_core_web_sm", exclude=["parser"])
fast.enable_pipe("senter")

for sent in fast(text).sents:
    print(repr(sent.text))
Output
'Dr. Rao moved to the U.S. in 1998.'
'He paid $4.50 for coffee.'
'Was it worth it?'
'He thinks so.'

Same four sentences, no dependency parse. The senter component is trained purely for boundaries — several times faster, nearly as accurate. Component juggling like this is the subject of spaCy pipelines.

Alternative: NLTK's Punkt (nltk.sent_tokenize) — an unsupervised, language-agnostic splitter that requires nltk.download("punkt_tab") first, a ~4 MB download. Solid for clean prose; weaker on messy web text.

Common mistakes

Splitting on ". " (dot-space) as the "smarter" one-liner. It survives "$4.50" but still shreds "Dr. Rao" and "U.S. in" — both have a space after the dot. Patching the pattern case by case ends with you re-inventing a worse segmenter.

Forgetting the other enders. Question marks, exclamation marks, ellipses, and — in real web text — line breaks that act as sentence ends without any punctuation. Chat messages and bullet lists barely use full stops; for those, split on newlines first, then segment within each line.

Assuming English rules travel. Hindi ends sentences with "।" (the danda), Chinese and Japanese use "。", and Thai marks boundaries with spaces. Language-specific models or rules are mandatory; see language detection.

Segmenting before cleaning encoding. Mojibake from the Unicode lesson can mangle the very punctuation the segmenter reads. Fix bytes first, cut sentences second.

Try it yourself

Feed the splitter three messy inputs: a WhatsApp-style message with no full stops, a paragraph containing "i.e." and "e.g.", and a price list with decimals. Compare naive, Punkt and spaCy outputs, and note which trap catches which tool.

What to learn next

Researcher — Mathematics and papers.

Problem statement and difficulty

Sentence boundary detection (SBD) is binary classification over candidate boundary positions (conventionally, each ., ?, !, and in the general case every inter-token gap). English news text is deceptively easy — the period is unambiguous ~90% of the time (Kiss and Strunk, 2006) — so error rates on WSJ underestimate difficulty on the open web, transcripts (no punctuation at all), legal text (heavy citation abbreviation), and biomedical prose ("E. coli" mid-sentence). Reported accuracies of 98–99% on newswire fall to the low 90s on user-generated content (Read et al., 2012, Sentence boundary detection: a long solved problem? — the title's answer is no).

Algorithms

  • Punkt (Kiss and Strunk, 2006): unsupervised. Detects abbreviations via collocational statistics — a type ending in "." whose truncated form co-occurs with the period far above chance is an abbreviation; log-likelihood ratios over type frequencies drive the decision, plus heuristics on following-token capitalisation. Trains from raw text in any alphabetic language; this corpus-adaptivity is why NLTK ships per-language Punkt parameter files.
  • Supervised sequence models: maximum-entropy classifiers over local features (Reynar and Ratnaparkhi, 1997); CRFs; and neural token classifiers. spaCy's senter is a lightweight token-level classifier; its parser-derived boundaries come from joint dependency parsing, where boundary and structure disambiguate each other.
  • Multilingual neural SBD: recent systems (Ersatz — Wicks and Post, 2021; wtpsplit/SaT — Frohmann et al., 2024) train subword transformers over many languages, handle missing punctuation, and segment transcripts — the current state of the art for web-scale multilingual pipelines.

Complexity is linear in tokens for all practical systems; the parser-based route costs the full parse (~10× the senter cost) but you were often parsing anyway.

Downstream sensitivity

Segmentation errors propagate multiplicatively: machine translation quality drops measurably with wrong boundaries since decoder context is truncated or polluted; summarisation extractors mis-rank shredded fragments; and in retrieval-augmented generation, chunk boundaries that cut mid-sentence measurably hurt retrieval precision — sentence-aligned chunking is the standard recommendation. Evaluation uses precision/recall on boundary positions against gold treebanks (UD's sentence segmentation task); when comparing tools, verify the same candidate set is being scored — some report only on period positions, inflating scores.

What to learn next