Sequence Labelling and Structure
Part-of-speech tagging
Part-of-speech tagging labels every word with its grammatical role in that exact sentence, which is the first structural clue almost every later NLP step depends on.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Part-of-speech tagging labels every word with the job it is doing in that sentence — noun, verb, adjective, and so on.
Think about a traffic police officer at a busy Chennai junction. A bus, a bike and an autorickshaw all arrive together. The officer does not care what colour any of them is. The officer only cares which lane each one belongs in, right now.
Part-of-speech tagging does the same thing to words. It looks at each word. It decides which "lane" the word belongs in — noun, verb, adjective, and a dozen other categories. The choice depends on the job that word is doing in this exact sentence.
Why it exists
The same word can do different jobs on different days. "Train" can be a noun, as in "I missed the train." It can also be a verb, as in "I train every morning." A computer reading raw text cannot tell which one it is.
Early language software tried a shortcut. Look up every word in a dictionary. Take the first meaning listed. This failed constantly, because dictionaries list every possible role a word can play, not the one it is playing right now.
Part-of-speech tagging fixes this by looking at context. It reads the words around a target word, not only the word by itself. Neighbours resolve a word's role. That single idea sits behind almost every later NLP technique you will meet, including attention.
How it works
A tagger reads a full sentence and assigns one label per word, using the surrounding words as evidence.
Sentence: I will train for the marathon
Tags: PRON AUX VERB ADP DET NOUN
^
"train" next to "will" looks like an action,
so it gets tagged as a VERB here.
Sentence: I will catch the train tomorrow
Tags: PRON AUX VERB DET NOUN NOUN
^
Same word "train", but here it follows
"the" — a determiner almost always
introduces a NOUN, not a verb.Nothing about the word "train" itself changed between the two sentences. Only its neighbours changed, and that was enough to flip the tag.
Where you have already seen it
- Grammar checkers. Word and Google Docs use part-of-speech tags to catch subject-verb mismatches like "the bus have arrived."
- Voice assistants. "Book a table" needs "book" tagged as a verb. "Read a book" needs it tagged as a noun.
- Search engines. "Runs" is a verb in one query and a noun in another, as in "running shoes." That distinction changes which results matter.
- Every later step in this section. Named entity recognition, dependency parsing and relation extraction all lean on part-of-speech tags, an early and cheap clue about sentence structure.
Remember this
- A part-of-speech tag describes the job a word is doing in one specific sentence, not a fixed property of the word.
- The same word can get different tags depending on what surrounds it.
- Part-of-speech tagging is usually the first structural layer other NLP tools are built on top of.
What to learn next
- BIO and BILOU tagging schemes — labelling whole entity spans, not single words.
- Named entity recognition — the task part-of-speech tagging usually feeds into.
- Dependency parsing — connecting tagged words into a full sentence structure.
Developer — Code and libraries.
spaCy ships a small, accurate, CPU-only part-of-speech tagger as part of its default pipeline. No GPU, no big download.
Setup
pip install spacy
python -m spacy download en_core_web_smen_core_web_sm is spaCy's small English pipeline, about 12 MB. It includes tokenization, part-of-speech tagging, and more, trained on web text.
Minimal runnable code
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Priya booked a fast train from Chennai to Bengaluru.")
for token in doc:
print(f"{token.text:10} {token.pos_:8} {token.tag_:6} {spacy.explain(token.tag_)}")Priya PROPN NNP noun, proper singular booked VERB VBD verb, past tense a DET DT determiner fast ADJ JJ adjective (English), other noun-modifier (Chinese) train NOUN NN noun, singular or mass from ADP IN conjunction, subordinating or preposition Chennai PROPN NNP noun, proper singular to ADP IN conjunction, subordinating or preposition Bengaluru PROPN NNP noun, proper singular . PUNCT . punctuation mark, sentence closer
Line by line
token.pos_ is the coarse tag — the "universal" set shared across languages: NOUN, VERB, ADJ, DET, ADP (a preposition or subordinating conjunction), PROPN (proper noun), PUNCT.
token.tag_ is the fine-grained tag, specific to English and borrowed from the Penn Treebank tag set — NNP for a proper singular noun, VBD for a past-tense verb. spacy.explain() turns any of these codes into a plain sentence, which saves a trip to a tag-set reference table.
fast was tagged ADJ. It modifies "train", the noun right after it. Move it elsewhere in the sentence and the tag can change.
Ambiguity, resolved by context
import spacy
nlp = spacy.load("en_core_web_sm")
for sent in ["I will train for the marathon tomorrow.", "I will catch the train tomorrow."]:
doc = nlp(sent)
for token in doc:
if token.text == "train":
print(f"{sent!r:45} -> 'train' tagged as {token.pos_} ({spacy.explain(token.pos_)})")'I will train for the marathon tomorrow.' -> 'train' tagged as VERB (verb) 'I will catch the train tomorrow.' -> 'train' tagged as NOUN (noun)
Same word, same spelling, opposite tags. The only thing that changed is what comes before it: "will train" reads as an action, "the train" reads as a thing.
Common mistakes
Assuming a word has one fixed part of speech. Many English words — "train", "book", "run", "light" — take different roles depending on the sentence. Never hardcode a word-to-tag lookup table.
Confusing pos_ and tag_. pos_ is coarse and consistent across languages, tag_ is fine-grained and English-specific. Mixing them up breaks code meant to run on multiple languages.
Running the tagger on lowercased text. spaCy's tagger was trained on normally-cased text. Lowercasing before tagging (a habit left over from older bag-of-words pipelines) removes a real signal — capitalisation is one clue the tagger uses to spot proper nouns — and measurably hurts accuracy.
Trusting the tagger blindly on messy text. Tweets, chat messages and transcribed speech confuse taggers trained on clean web and news text. Expect more errors on that kind of input, and consider it a reason to check output rather than a bug in your code.
Try it yourself
Feed the tagger a sentence with a different ambiguous word, such as "light": "Please light the lamp before it gets dark," versus "The bag was surprisingly light." Predict the tags before running it, then check.
What to learn next
- BIO and BILOU tagging schemes — the next layer of structure, built on top of tagged tokens.
- Dependency parsing — spaCy's parser uses these same tags as input features.
- spaCy pipelines — everything else
nlp()does in one call.
Researcher — Mathematics and papers.
The task, formally
Given a sequence of tokens x = (x_1, ..., x_n), assign a tag sequence y = (y_1, ..., y_n), y_i drawn from a fixed tag set T (36 tags in the Penn Treebank set; 17 in Universal POS).
The naive approach scores each y_i independently given local features of x_i. The better approach scores the whole sequence jointly:
y* = argmax over y of P(y | x)P(y | x)is the model's probability of tag sequenceygiven the observed tokensx.argmaxpicks the single highest-scoring tag sequence, not only the highest-scoring tag per position.
Scoring jointly matters because tags are not independent: an adjective is very likely to be followed by a noun, and a sequence-level model can use that.
From HMMs to neural taggers
Hidden Markov Models (used through the 1990s) model P(x, y) = P(y_1) * product over i=2..n of P(y_i | y_{i-1}) * product over i=1..n of P(x_i | y_i), decoded exactly with the Viterbi algorithm in O(n * |T|^2) time.
P(y_i | y_{i-1})is the transition probability between consecutive tags.P(x_i | y_i)is the emission probability of a word given its tag.
Maximum Entropy Markov Models and CRFs (2000s) replaced the generative emission term with a discriminative classifier over arbitrary overlapping features — capitalisation, suffix, surrounding words — without the independence assumptions an HMM needs. See Conditional random fields for the full mechanics.
Neural taggers (2015 onward) replace hand-built features with a learned encoder — a BiLSTM, or a transformer such as BERT — feeding a per-token classifier or a CRF layer on top. spaCy's en_core_web_sm tagger is a small CNN-based encoder, chosen for CPU speed over the small accuracy gain a transformer would give.
Accuracy and its ceiling
State-of-the-art English part-of-speech tagging exceeds 97% token accuracy on the Penn Treebank Wall Street Journal test set. This number is close to inter-annotator agreement — the rate at which two human linguists agree with each other — so most remaining errors are genuinely ambiguous cases, not model failures.
Accuracy drops sharply outside clean, edited English: 5 to 10 points lower on social media text (Gimpel et al., 2011), and considerably lower for morphologically rich, low-resource languages where far less annotated data exists.
Key references
- Marcus, M., Santorini, B. & Marcinkiewicz, M. (1993). Building a Large Annotated Corpus of English: The Penn Treebank. Computational Linguistics 19(2). Defines the 36-tag set still used by most English taggers.
- Toutanova, K., Klein, D., Manning, C. & Singer, Y. (2003). Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network. NAACL. The maximum-entropy tagger that set the standard for a decade.
- Nivre, J. et al. (2016). Universal Dependencies v1: A Multilingual Treebank Collection. LREC. Defines the 17-tag Universal POS set used across languages.
- Gimpel, K. et al. (2011). Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments. ACL.
Current state and open problems
Part-of-speech tagging is close to a solved problem for high-resource, edited text, and is now mostly consumed as a cheap intermediate feature inside larger pipelines rather than studied on its own. The open work has moved to two edges: robustness on noisy, informal and code-mixed text, and coverage for the thousands of languages with little or no annotated treebank data, where even the definition of "part of speech" borrowed from English grammar does not map cleanly.
What to learn next
- Conditional random fields — the sequence model that made joint tagging practical before neural taggers.
- BIO and BILOU tagging schemes — applying the same sequence-labelling idea to spans instead of single tokens.
- Attention — the mechanism modern taggers use in place of a fixed transition matrix.