spaCy pipelines
One nlp() call runs a whole assembly line — tagger, parser, lemmatiser, entity recogniser — and knowing the stations lets you skip the ones you are not using.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A spaCy pipeline is an assembly line for text: your sentence enters at one end, and each station stamps on another layer of understanding.
Think of a car factory line. The chassis rolls past stations: one fits the engine, one the doors, one the paint. Each station needs the previous one's work in place. In spaCy, the raw sentence is the chassis, and the stations stamp on word types, grammar links, dictionary forms and names.
Why it exists
Understanding text is not one job but a stack of jobs. Which words are nouns? How do the words connect? What is each word's dictionary form? Which words name people, places, companies?
You could wire up separate tools for each and pray their formats agree. spaCy's answer is one object: the pipeline. Its stations come pre-assembled in the right order, trained to work together, and shipped as a single download. One call runs the whole line:
"Infosys hired 500 engineers." → [tagger] → [parser] → [lemmatiser] → [NER] → docThe result is called a doc, and it is the annotated car. Every token carries its part of speech and its lemma, the dictionary form. It also carries an entity label if it names something.
How it works
Two facts make the line practical:
- Stations depend on earlier stations. The lemmatiser reads the tagger's output — "meeting" lemmatises differently as noun or verb. Order is not decoration.
- You can switch stations off. Skipping the stations you do not need is the main speed lever. Only lemmatising? Switch off the entity recogniser and the line runs much faster.
A real example you have seen
Customer-support systems read incoming complaints with exactly such lines. Find the product name, then find what happened to it. Normalise the words, then route the ticket. News apps that highlight company names in articles run the same NER station you will run below.
Remember this
- A pipeline = ordered stations, each adding one layer of annotation.
- Stations depend on earlier ones — order matters.
- Speed comes from switching off stations you do not use.
What to learn next
- Getting clean text out of PDFs — feeding the assembly line from the messiest common source.
- Named entity recognition — the NER station as a topic of its own.
- Sentence segmentation — the
senterstation and its rivals.
Developer — Code and libraries.
Setup
pip install spacy
python -m spacy download en_core_web_smen_core_web_sm is a ~12 MB one-time download. Larger siblings: md (~40 MB) and lg (~400 MB) add word vectors; en_core_web_trf (~450 MB) swaps in a transformer for top accuracy at real GPU-wanting cost. Verified with spaCy 3.8, CPU only.
The line, opened up
import spacy
nlp = spacy.load("en_core_web_sm")
print("components:", nlp.pipe_names)
doc = nlp("Infosys hired 500 engineers in Bengaluru last March.")
print("\ntoken pos lemma entity")
for tok in doc:
ent = tok.ent_type_ or "-"
print(f"{tok.text:9s} {tok.pos_:6s} {tok.lemma_:10s} {ent}")
# processing many texts: nlp.pipe batches, and you can switch parts off
texts = ["Great phone.", "Battery died in a day.", "Camera is superb."]
with nlp.select_pipes(enable=["tok2vec", "tagger", "attribute_ruler", "lemmatizer"]):
for d in nlp.pipe(texts):
print([t.lemma_ for t in d])components: ['tok2vec', 'tagger', 'parser', 'attribute_ruler', 'lemmatizer', 'ner'] token pos lemma entity Infosys PROPN Infosys ORG hired VERB hire - 500 NUM 500 CARDINAL engineers NOUN engineer - in ADP in - Bengaluru PROPN Bengaluru GPE last ADJ last DATE March PROPN March DATE . PUNCT . - ['great', 'phone', '.'] ['battery', 'die', 'in', 'a', 'day', '.'] ['camera', 'be', 'superb', '.']
The walkthrough
nlp.pipe_names lists the stations. tok2vec computes shared word representations the other statistical stations consume — switch it off and the tagger goes blind. attribute_ruler applies small rule fixes between tagging and lemmatising. Tokenization itself is station zero, always on, not listed.
One pass filled three columns at once. Infosys came out PROPN (proper noun) and ORG (organisation entity); "hired" became lemma "hire"; "last March" was grouped as a DATE — two tokens, one entity. Whole-entity spans live in doc.ents, token-level labels in tok.ent_type_.
nlp.pipe(texts) is the batching API. It processes a stream far faster than calling nlp() in a loop, and takes n_process for multiprocess runs. For large corpora this is the difference between minutes and hours.
select_pipes switches stations off cleanly. Inside the with block, the parser and NER never run — the printed lemmas needed only the four enabled components. The block restores everything on exit. For a permanent trim, load with exclude=["ner", "parser"] and the excluded stations are never even read from disk.
Common mistakes
Loading the model inside a loop or per request. spacy.load reads tens of megabytes and builds the whole line. Load once at startup, reuse the nlp object everywhere. It is the single most common spaCy performance bug.
Expecting sm to know word similarity. en_core_web_sm ships no word vectors; doc.similarity() warns and falls back to weak context tensors. Use en_core_web_md or dedicated embeddings for similarity work.
Mutating text and expecting the doc to follow. A doc is a snapshot of the exact input string. Clean your text before the pipeline — the normalisation passes come first, always.
Trusting NER on out-of-domain text. The small model was trained on web and news text. On medical notes, legal filings or Hinglish chat, entity quality drops steeply. Test on your data before believing the labels; fine-tuning or a domain model (scispacy for biomedical) is often needed.
Try it yourself
Run nlp("Apple bought a startup for $2 billion in March.") and print doc.ents with labels. Watch "Apple" get ORG — then lowercase the sentence and see what survives. Case is a feature the entity station leans on.
What to learn next
- Getting clean text out of PDFs — feeding the assembly line from the messiest common source.
- Named entity recognition — the NER station as a topic of its own.
- Sentence segmentation — the
senterstation and its rivals.
Researcher — Mathematics and papers.
Architecture
spaCy v3 pipelines are configuration-defined DAGs of components sharing a Vocab and operating in-place on Doc objects — an annotate-in-place design that avoids serialisation between stations (Honnibal and Montani, 2017). Statistical components in the core models share the tok2vec encoder: a CNN with hash embeddings ("embed, encode, attend, predict" — Bloom embeddings hash multiple lexical attributes into a fixed table, trading collisions for a small memory footprint). Listeners let downstream heads (tagger, parser, NER) backprop into the shared encoder during training; _trf models replace the CNN with a transformer via spacy-transformers, aligning wordpieces to linguistic tokens.
The parser and NER share a transition-based imitation-learning core (arc-eager dependency parsing and BILOU chunking as action sequences), with dynamic oracles (Goldberg and Nivre, 2012) and beam or greedy decoding. Sentence boundaries fall out of the parser's transitions, which is why doc.sents needs either parser or the dedicated senter — as used in sentence segmentation.
Speed and accuracy envelope
Published figures for en_core_web_sm on CPU are on the order of 10⁴ words/second for the full line; trf models run ~10² –10³ on CPU and need GPU for production throughput, for roughly +2–4 points of UAS/NER F1 (spaCy's model cards publish per-component metrics; sm NER sits near 0.85 F1 on OntoNotes-style eval, trf near 0.90). The practical levers, in order: exclude unused components; nlp.pipe batching (amortises per-doc overhead); n_process for CPU parallelism; and doc_bin serialisation for corpus caching. Memory: the Vocab grows with observed strings — long-running services processing unbounded unique tokens should recycle the nlp object periodically.
Extension and the ecosystem
Custom stations register via @Language.component (stateless) or @Language.factory (stateful, serialisable into the config system); EntityRuler injects pattern-based entities (the regex lesson's world) that can pre-empt or supplement statistical NER; Matcher/PhraseMatcher give token-level rule engines over linguistic attributes — a capability plain regex lacks. Training uses the config/spacy train CLI with per-component losses. Positioning against the field: spaCy occupies the "fast, production, pipeline" niche; Stanza (Qi et al., 2020) trades speed for accuracy across 60+ languages with a similar station design; and transformer fine-tunes via HuggingFace win pure-accuracy contests per task at higher serving cost. For lemma/POS/dependency layers feeding classical features, spaCy remains the default answer in Python.
What to learn next
- Getting clean text out of PDFs — feeding the assembly line from the messiest common source.
- Named entity recognition — the NER station as a topic of its own.
- Sentence segmentation — the
senterstation and its rivals.