Sentiment, Opinion and Text Mining
Classifying into a 300-label taxonomy
Hierarchical text classification picks a broad category first, then a specific label within it, instead of choosing directly from hundreds of flat options at once.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Hierarchical text classification picks a broad category first, then narrows down within it. It never chooses directly from hundreds of options at once.
Picture a library. You do not scan every single book on every shelf to find one novel. You go to "Fiction" first, then "Mystery," then find the specific book. Broad, then narrow, one step at a time.
Text classification with hundreds of possible labels works the same way. Picking one of 300 labels directly is hard. Picking one of 12 broad categories, then one of 20 within it, is far easier.
Why it exists
A support inbox might sort tickets into 300 fine-grained categories. Think "battery draining fast," "screen cracked," "wrong size delivered," and so on. Choosing directly from all 300 at once overwhelms most models.
Splitting the choice into stages solves this. First pick a broad category, like "electronics" or "clothing," from a short list. Then pick the specific label from only the handful that belong to that category.
Each individual decision gets easier. A model comparing 12 categories has far less room for confusion than one comparing 300 labels simultaneously.
How it works
Ticket: "The battery on my phone dies by lunchtime, even though it's brand new."
|
v
Stage 1: broad category
electronics / clothing / furniture
|
v
electronics (0.999)
|
v
Stage 2: fine label, within electronics only
phone screen cracked / battery draining fast / wifi keeps disconnecting
|
v
battery draining fast (0.926)Notice the second stage never even considers clothing or furniture labels. Narrowing the field first is what makes the fine-grained decision manageable.
Where you have already seen it
- E-commerce category trees. "Electronics > Phones > Accessories > Cases" is picked one level at a time. It is never chosen from a flat list of ten thousand products.
- Support ticket routing systems. Department first, then issue type, then specific sub-issue, exactly the pattern shown above.
- News site sections. "Sports > Cricket > IPL," a hierarchy readers navigate, and often the same hierarchy the article was tagged with.
- Medical and legal coding systems. Broad category, then increasingly specific codes, because a flat list of thousands of codes is unusable directly.
Remember this
- Hierarchical classification splits one hard decision into several easier ones.
- Each stage only has to choose among the options that survived the previous stage.
- It is a strategy, not a specific algorithm — many different techniques can implement it.
What to learn next
- Text classification — the underlying technique used at every stage.
- Zero-shot text classification — the specific method used below, in full depth.
- One-vs-rest and one-vs-one — another classic strategy for handling many possible labels.
Developer — Code and libraries.
Below, the same zero-shot model from earlier lessons runs twice per ticket, once per stage. A real product taxonomy would have many more categories and labels than this; the two-step idea stays identical either way.
Setup
pip install transformers torchTwo-stage classification
from transformers import pipeline
classifier = pipeline(
"zero-shot-classification",
model="MoritzLaurer/deberta-v3-xsmall-zeroshot-v1.1-all-33",
)
# A small stand-in for a real product taxonomy. A real support inbox might have
# 20 broad categories and 300 fine-grained ones; the two-step idea is the same.
taxonomy = {
"electronics": ["phone screen cracked", "battery draining fast", "wifi keeps disconnecting"],
"clothing": ["wrong size delivered", "fabric tore after one wash", "colour faded quickly"],
"furniture": ["missing assembly screws", "surface scratched on arrival", "wobbly legs"],
}
ticket = "The battery on my phone dies by lunchtime even though it's brand new."
# Stage 1: which broad category?
top = classifier(ticket, candidate_labels=list(taxonomy.keys()))
category = top["labels"][0]
print(f"Category: {category} ({top['scores'][0]:.3f})")
# Stage 2: which fine-grained label, inside that category only?
fine = classifier(ticket, candidate_labels=taxonomy[category])
print(f"Fine label: {fine['labels'][0]} ({fine['scores'][0]:.3f})")Category: electronics (0.999) Fine label: battery draining fast (0.926)
Line by line
taxonomy is a dictionary mapping each category to its fine-grained labels. This is the whole data structure the hierarchy needs. Real taxonomies are far bigger, but the shape stays exactly the same.
Stage 1 chooses among three broad categories. Only three options, so even a small model handles this confidently. The score, 0.999, reflects how unambiguous the choice was.
Stage 2 only ever sees the three labels inside electronics. Clothing and furniture labels were never candidates at this stage. That elimination is the entire benefit of splitting the decision into two.
The same model, classifier, runs both stages. Nothing about the model itself is hierarchical. The hierarchy lives entirely in how the code calls it, twice, with a narrowing set of candidates each time.
Common mistakes
Running all labels through one flat classification call. With 300 candidate labels in one call, even a strong model's confidence spreads thin, and accuracy drops. Splitting into stages consistently helps once label counts get large.
Letting a wrong Stage 1 pick doom the entire result. Say Stage 1 wrongly picks "clothing." Stage 2 can never recover the correct "electronics" label. It is not even a candidate anymore. Track Stage 1 confidence and flag low-confidence cases for review.
Building a taxonomy with overlapping categories. Say a real product could reasonably belong to two categories. Stage 1 becomes a coin flip, and that uncertainty propagates into Stage 2. Keep top-level categories as distinct as the actual data allows.
Forgetting that more stages means more compute. Two model calls per item instead of one doubles inference cost. For very large label sets this is still worth it, but it is a real trade-off to measure, not assume.
Try it yourself
Add a fourth top-level category, "other", with no fine-grained labels of its own. Route anything with low Stage 1 confidence there, instead of forcing a guess. Compare how that changes the handling of an ambiguous ticket.
What to learn next
- Zero-shot text classification — the classification method used at both stages here.
- Macro, micro and weighted averaging — how to fairly evaluate a classifier across many unevenly sized categories.
- Few-shot text classification with SetFit — training a dedicated classifier once labelled examples exist for each category.
Researcher — Mathematics and papers.
Formalising the hierarchy
Hierarchical text classification, HTC, predicts a path through a label taxonomy T. That taxonomy is typically a tree, or a directed acyclic graph, rather than a single flat, unstructured set. Two structural approaches dominate the literature.
Local classifiers. Train a separate classifier at each node, or at each level, of the hierarchy. This is exactly what the developer block implements: one classifier for the top level, one for each subtree beneath it. This is simple to implement, and it reuses the same base model. But errors at a parent node are unrecoverable by construction. A child model never sees candidates outside its assigned parent.
Global (flat) classifiers with structural regularisation. Train a single model over all leaf labels at once. Incorporate the hierarchy into the loss function itself. It penalises predictions that violate parent-child consistency, rather than using separate models per level. HiAGM (Zhou et al., 2020) represents this direction. It uses graph neural networks over the label hierarchy. Label embeddings then inform each other structurally, rather than being predicted independently.
Why the local approach's error propagation matters
In a strict local, top-down approach, one formula captures it:
P(correct leaf) = P(correct at level 1) * P(correct at level 2 | correct at level 1) * ...
``` Each level's accuracy multiplies. A taxonomy with many levels can have surprisingly low end-to-end accuracy, even when every individual level classifies well. This is the formal version of the developer block's "wrong Stage 1 pick dooms the result" observation.
One mitigation: keep top-of-hierarchy classifiers as accurate as possible, since their errors compound the most. Another: beam search across levels. Carry forward the top-`k` categories, not only the single best, to recover from a near-miss early on.
## Zero-shot HTC via NLI
The developer block's approach runs zero-shot classification at each level, using an NLI-based model. This sidesteps needing labelled training data at any level. The cost: the same calibration caveats already covered in [zero-shot text classification](/learn/opinion-mining/zero-shot-text-classification). This is a genuinely practical choice when a taxonomy exists, but labelled examples per fine-grained category do not.
## Evaluation
Standard flat metrics, precision, recall, F1, can be computed at the leaf level, ignoring hierarchy entirely. But this discards useful information. A prediction one level off is a smaller error than one in a completely unrelated branch. **Hierarchical F-measure** (Kiritchenko, Matwin & Famili, 2006) instead credits partial correctness along the taxonomy path. It distinguishes "wrong leaf, right parent" from "wrong everything."
## Key references
- Kiritchenko, S., Matwin, S. & Famili, F. (2006). *Functional Annotation of Genes Using Hierarchical Text Categorization.* — introduces hierarchical F-measure.
- Zhou, J. et al. (2020). *Hierarchy-Aware Global Model for Hierarchical Text Classification.* ACL.
- Mao, Y. et al. (2019). *Hierarchical Text Classification with Reinforced Label Assignment.* [arXiv:1908.10419](https://arxiv.org/abs/1908.10419)
## Current state and open problems
LLMs handle large label taxonomies well when the taxonomy, or a relevant subset, sits directly in the prompt. The hierarchical narrowing then happens implicitly, through in-context reasoning, rather than a separately coded staged pipeline.
The open problem is taxonomy drift. Real-world taxonomies change: categories get added, merged and renamed as a business evolves. Neither local nor global HTC models handle this gracefully by default. A taxonomy change generally forces retraining or re-specification, regardless of which architecture was chosen.
## What to learn next
- [Zero-shot text classification](/learn/opinion-mining/zero-shot-text-classification) — the underlying classification method used at each level here.
- [One-vs-rest and one-vs-one](/learn/hard-targets/one-vs-rest-and-one-vs-one) — a flat alternative strategy worth comparing against.
- [Few-shot text classification with SetFit](/learn/opinion-mining/few-shot-text-classification) — training dedicated per-level classifiers once labelled data exists.
:::