NER (named entity recognition)
In one sentence NER finds and labels the names in text — people, places, organisations, dates, amounts — turning prose into structured data.
Updated
Named entity recognition, or NER, is the task of locating the proper names in text and labelling what kind of thing each one is.
It is reading with a pack of highlighters: yellow for people, green for places, pink for organisations, blue for dates and amounts. Given "Ratan Tata founded Tata Sons in Mumbai in 1868", the highlighters mark PERSON, ORG, LOCATION, DATE — and suddenly a sentence has become database rows.
[Ratan Tata]PERSON founded [Tata Sons]ORG in [Mumbai]LOC in [1868]DATEThat conversion — prose in, structured fields out — is why NER quietly powers so much: search engines understanding that your query names a film, news systems tracking which companies appear together, hospitals extracting drug names and dosages from clinical notes, banks pulling parties and amounts from contracts, and de-identification systems finding personal names to redact.
Technically it is classification per token: each token gets a tag, with a B/I/O scheme (begin, inside, outside) so multi-word names like "State Bank of India" hold together. Fine-tuned BERT-class models are the standard engine; spaCy makes a strong pretrained pipeline three lines of Python. The genuinely hard parts: names never seen before (models generalise from context, imperfectly), ambiguity ("Amazon" — river, company?), and domain shift — a news-trained model stumbles on medical or legal text, which is why domain-specific fine-tuning persists. LLMs now do capable zero-shot NER with a structured-output prompt; dedicated models keep winning where volume is high and latency budgets are tight.
Where to go next
- Full lesson: What is NLP?
- Related terms: bert, token, classification, structured-output