Sequence Labelling and Structure

Part-of-speech tagging

Part-of-speech tagging labels every word with its grammatical role in that exact sentence, which is the first structural clue almost every later NLP step depends on.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Part-of-speech tagging labels every word with the job it is doing in that sentence — noun, verb, adjective, and so on.

Think about a traffic police officer at a busy Chennai junction. A bus, a bike and an autorickshaw all arrive together. The officer does not care what colour any of them is. The officer only cares which lane each one belongs in, right now.

Part-of-speech tagging does the same thing to words. It looks at each word. It decides which "lane" the word belongs in — noun, verb, adjective, and a dozen other categories. The choice depends on the job that word is doing in this exact sentence.

Why it exists

The same word can do different jobs on different days. "Train" can be a noun, as in "I missed the train." It can also be a verb, as in "I train every morning." A computer reading raw text cannot tell which one it is.

Early language software tried a shortcut. Look up every word in a dictionary. Take the first meaning listed. This failed constantly, because dictionaries list every possible role a word can play, not the one it is playing right now.

Part-of-speech tagging fixes this by looking at context. It reads the words around a target word, not only the word by itself. Neighbours resolve a word's role. That single idea sits behind almost every later NLP technique you will meet, including attention.

How it works

A tagger reads a full sentence and assigns one label per word, using the surrounding words as evidence.

  Sentence:   I     will   train      for    the   marathon
  Tags:       PRON  AUX    VERB       ADP    DET   NOUN
                            ^
                            "train" next to "will" looks like an action,
                            so it gets tagged as a VERB here.

  Sentence:   I     will   catch  the   train      tomorrow
  Tags:       PRON  AUX    VERB   DET   NOUN        NOUN
                                        ^
                                        Same word "train", but here it follows
                                        "the" — a determiner almost always
                                        introduces a NOUN, not a verb.

Nothing about the word "train" itself changed between the two sentences. Only its neighbours changed, and that was enough to flip the tag.

Where you have already seen it

  • Grammar checkers. Word and Google Docs use part-of-speech tags to catch subject-verb mismatches like "the bus have arrived."
  • Voice assistants. "Book a table" needs "book" tagged as a verb. "Read a book" needs it tagged as a noun.
  • Search engines. "Runs" is a verb in one query and a noun in another, as in "running shoes." That distinction changes which results matter.
  • Every later step in this section. Named entity recognition, dependency parsing and relation extraction all lean on part-of-speech tags, an early and cheap clue about sentence structure.

Remember this

  • A part-of-speech tag describes the job a word is doing in one specific sentence, not a fixed property of the word.
  • The same word can get different tags depending on what surrounds it.
  • Part-of-speech tagging is usually the first structural layer other NLP tools are built on top of.

What to learn next

Developer — Code and libraries.

spaCy ships a small, accurate, CPU-only part-of-speech tagger as part of its default pipeline. No GPU, no big download.

Setup

bash
pip install spacy
python -m spacy download en_core_web_sm

en_core_web_sm is spaCy's small English pipeline, about 12 MB. It includes tokenization, part-of-speech tagging, and more, trained on web text.

Minimal runnable code

pos_tag.py
import spacy

nlp = spacy.load("en_core_web_sm")
doc = nlp("Priya booked a fast train from Chennai to Bengaluru.")

for token in doc:
    print(f"{token.text:10} {token.pos_:8} {token.tag_:6} {spacy.explain(token.tag_)}")
Output
Priya      PROPN    NNP    noun, proper singular
booked     VERB     VBD    verb, past tense
a          DET      DT     determiner
fast       ADJ      JJ     adjective (English), other noun-modifier (Chinese)
train      NOUN     NN     noun, singular or mass
from       ADP      IN     conjunction, subordinating or preposition
Chennai    PROPN    NNP    noun, proper singular
to         ADP      IN     conjunction, subordinating or preposition
Bengaluru  PROPN    NNP    noun, proper singular
.          PUNCT    .      punctuation mark, sentence closer

Line by line

token.pos_ is the coarse tag — the "universal" set shared across languages: NOUN, VERB, ADJ, DET, ADP (a preposition or subordinating conjunction), PROPN (proper noun), PUNCT.

token.tag_ is the fine-grained tag, specific to English and borrowed from the Penn Treebank tag set — NNP for a proper singular noun, VBD for a past-tense verb. spacy.explain() turns any of these codes into a plain sentence, which saves a trip to a tag-set reference table.

fast was tagged ADJ. It modifies "train", the noun right after it. Move it elsewhere in the sentence and the tag can change.

Ambiguity, resolved by context

ambiguity.py
import spacy

nlp = spacy.load("en_core_web_sm")

for sent in ["I will train for the marathon tomorrow.", "I will catch the train tomorrow."]:
    doc = nlp(sent)
    for token in doc:
        if token.text == "train":
            print(f"{sent!r:45} -> 'train' tagged as {token.pos_} ({spacy.explain(token.pos_)})")
Output
'I will train for the marathon tomorrow.'     -> 'train' tagged as VERB (verb)
'I will catch the train tomorrow.'            -> 'train' tagged as NOUN (noun)

Same word, same spelling, opposite tags. The only thing that changed is what comes before it: "will train" reads as an action, "the train" reads as a thing.

Common mistakes

Assuming a word has one fixed part of speech. Many English words — "train", "book", "run", "light" — take different roles depending on the sentence. Never hardcode a word-to-tag lookup table.

Confusing pos_ and tag_. pos_ is coarse and consistent across languages, tag_ is fine-grained and English-specific. Mixing them up breaks code meant to run on multiple languages.

Running the tagger on lowercased text. spaCy's tagger was trained on normally-cased text. Lowercasing before tagging (a habit left over from older bag-of-words pipelines) removes a real signal — capitalisation is one clue the tagger uses to spot proper nouns — and measurably hurts accuracy.

Trusting the tagger blindly on messy text. Tweets, chat messages and transcribed speech confuse taggers trained on clean web and news text. Expect more errors on that kind of input, and consider it a reason to check output rather than a bug in your code.

Try it yourself

Feed the tagger a sentence with a different ambiguous word, such as "light": "Please light the lamp before it gets dark," versus "The bag was surprisingly light." Predict the tags before running it, then check.

What to learn next

Researcher — Mathematics and papers.

The task, formally

Given a sequence of tokens x = (x_1, ..., x_n), assign a tag sequence y = (y_1, ..., y_n), y_i drawn from a fixed tag set T (36 tags in the Penn Treebank set; 17 in Universal POS).

The naive approach scores each y_i independently given local features of x_i. The better approach scores the whole sequence jointly:

text
y* = argmax over y of P(y | x)
  • P(y | x) is the model's probability of tag sequence y given the observed tokens x.
  • argmax picks the single highest-scoring tag sequence, not only the highest-scoring tag per position.

Scoring jointly matters because tags are not independent: an adjective is very likely to be followed by a noun, and a sequence-level model can use that.

From HMMs to neural taggers

Hidden Markov Models (used through the 1990s) model P(x, y) = P(y_1) * product over i=2..n of P(y_i | y_{i-1}) * product over i=1..n of P(x_i | y_i), decoded exactly with the Viterbi algorithm in O(n * |T|^2) time.

  • P(y_i | y_{i-1}) is the transition probability between consecutive tags.
  • P(x_i | y_i) is the emission probability of a word given its tag.

Maximum Entropy Markov Models and CRFs (2000s) replaced the generative emission term with a discriminative classifier over arbitrary overlapping features — capitalisation, suffix, surrounding words — without the independence assumptions an HMM needs. See Conditional random fields for the full mechanics.

Neural taggers (2015 onward) replace hand-built features with a learned encoder — a BiLSTM, or a transformer such as BERT — feeding a per-token classifier or a CRF layer on top. spaCy's en_core_web_sm tagger is a small CNN-based encoder, chosen for CPU speed over the small accuracy gain a transformer would give.

Accuracy and its ceiling

State-of-the-art English part-of-speech tagging exceeds 97% token accuracy on the Penn Treebank Wall Street Journal test set. This number is close to inter-annotator agreement — the rate at which two human linguists agree with each other — so most remaining errors are genuinely ambiguous cases, not model failures.

Accuracy drops sharply outside clean, edited English: 5 to 10 points lower on social media text (Gimpel et al., 2011), and considerably lower for morphologically rich, low-resource languages where far less annotated data exists.

Key references

  • Marcus, M., Santorini, B. & Marcinkiewicz, M. (1993). Building a Large Annotated Corpus of English: The Penn Treebank. Computational Linguistics 19(2). Defines the 36-tag set still used by most English taggers.
  • Toutanova, K., Klein, D., Manning, C. & Singer, Y. (2003). Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network. NAACL. The maximum-entropy tagger that set the standard for a decade.
  • Nivre, J. et al. (2016). Universal Dependencies v1: A Multilingual Treebank Collection. LREC. Defines the 17-tag Universal POS set used across languages.
  • Gimpel, K. et al. (2011). Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments. ACL.

Current state and open problems

Part-of-speech tagging is close to a solved problem for high-resource, edited text, and is now mostly consumed as a cheap intermediate feature inside larger pipelines rather than studied on its own. The open work has moved to two edges: robustness on noisy, informal and code-mixed text, and coverage for the thousands of languages with little or no annotated treebank data, where even the definition of "part of speech" borrowed from English grammar does not map cleanly.

What to learn next

What to learn next

These follow on from what you just read.

  • Sequence Labelling and Structure

    BIO and BILOU tagging schemes

    BIO and BILOU are ways of writing down where an entity starts, continues and ends using one tag per token, which is what turns entity spans into something a per-token classifier can learn.

  • Sequence Labelling and Structure

    Conditional random fields

    A conditional random field tags a whole sequence at once instead of one word at a time, so it can rule out impossible tag combinations like a sentence ending mid-entity.

  • Sequence Labelling and Structure

    Training NER on your own entities

    Off-the-shelf named entity recognisers only know the entity types they were trained on, so extracting your own categories means labelling examples and training a small model from scratch.