Natural Language Processing

What is NLP?

Natural language processing is how a computer reads, understands and writes human language instead of only handling numbers.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. The part that is genuinely hard
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Natural language processing, or NLP, is teaching a computer to work with human language.

That means reading it, making sense of it, and writing it back.

Think about your phone keyboard. You type "I am on my" and it offers "way". Nobody sat down and wrote a rule for that exact sentence. The keyboard learned it by reading millions of messages that other people typed first.

That small guessing trick is NLP in its most everyday form. Everything else in this section is the same idea, grown much larger.

Why it exists

Computers were built for numbers. A bank balance, a distance, a date — these have one meaning and one shape. Language has neither.

Take the word "bank". River bank. Money bank. Bank a turn on a bike. Same four letters, three different worlds. You picked the right one instantly because of the words around it.

Now try to write that rule down for a computer. People tried for forty years. They wrote grammar rules by hand, thousands of them. Every rule fixed one sentence and broke two others.

Here is a famous example. Someone built an early translation system in the 1960s. It turned "the spirit is willing but the flesh is weak" into a sentence about vodka and rotten meat. The words were translated correctly. The meaning was gone.

The breakthrough was giving up on hand-written rules. Instead of telling the computer what language means, we show it enormous amounts of real language. Then we let it find the patterns on its own. That switch is what made translation, voice assistants and chatbots possible.

How it works

Every NLP system, from your keyboard to the biggest chatbot, runs the same four steps.

  "the chai is hot"
         |
         v
  [1] CHOP into pieces         ->  "the" "chai" "is" "hot"
         |
         v
  [2] TURN each piece          ->  every piece becomes a
      into a list of numbers       list of numbers
         |
         v
  [3] MODEL does its work      ->  finds patterns across
                                   the whole sentence
         |
         v
  [4] TURN numbers back        ->  "यह चाय गरम है"
      into words or a label        or "positive" or an answer

Step one is called tokenization, which means cutting text into small chunks. Step two is called embedding, which means giving each chunk a list of numbers that stands for its meaning. Both get their own lesson.

The important thing to hold on to is this. A model never sees letters. It sees numbers all the way through, and words only appear again at the very end.

Where you have already seen it

You use NLP many times a day, mostly without noticing.

  • Your spam folder. Nobody reads your mail to sort it. A model learned what spam wording looks like.
  • Google Translate. Point your camera at a signboard and the words change language.
  • YouTube auto-captions. Speech becomes text, then text gets punctuated and cleaned.
  • Search that forgives typos. You type "recipy for pav bhaji" and still get the right page.
  • Any chatbot. It reads your question and writes a fresh answer, word by word.

The part that is genuinely hard

Language carries things that are not in the words at all.

If a friend says "wow, great timing" when you arrive an hour late, every word is positive. The meaning is not. Sarcasm, politeness, hints and jokes live in shared context, not in the sentence.

Models today handle this far better than five years ago, and still get it wrong. This is not a bug waiting to be fixed next month. It is the honest edge of the field. Anyone who tells you language is a solved problem is selling something.

Remember this

  • NLP is any system that reads or writes human language, from a keyboard suggestion to a chatbot.
  • Words are turned into numbers first, and only turned back into words at the very end.
  • Hand-written grammar rules failed. Learning patterns from huge amounts of real text is what worked.

What to learn next

Developer — Code and libraries.

You do not need a large model to do real NLP. A word-counting model plus logistic regression solves a surprising number of production problems, and it runs on any laptop in under a second.

Build that first. It gives you a baseline number, and a baseline is the only thing that tells you whether a bigger model was worth it.

Setup

bash
pip install scikit-learn

That is the whole dependency list. No GPU, no downloads at runtime.

A working sentiment classifier

sentiment.py
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression

train_texts = [
    "the food was delicious and the staff was kind",
    "wonderful service, I will come back again",
    "loved the biryani, best in the city",
    "great experience, very tasty food",
    "the food was cold and tasteless",
    "terrible service, we waited an hour",
    "worst biryani I have ever eaten",
    "bad experience, never coming back",
]
train_labels = ["positive"] * 4 + ["negative"] * 4

vec = CountVectorizer()          # text -> a table of word counts
X = vec.fit_transform(train_texts)

clf = LogisticRegression(max_iter=1000)   # extra passes so the fit finishes cleanly
clf.fit(X, train_labels)

tests = ["the biryani was tasty", "we waited an hour for cold food"]
for text, guess in zip(tests, clf.predict(vec.transform(tests))):
    print(f"{guess:8} <- {text}")
Output
positive <- the biryani was tasty
negative <- we waited an hour for cold food

Eight training sentences. Two correct predictions on sentences the model never saw.

Line by line

CountVectorizer() builds a vocabulary from the training text, then turns each sentence into a row of counts. If the vocabulary has 40 words, every sentence becomes 40 numbers. Most of them are zero.

By default it lowercases everything and keeps only tokens of two or more word characters. So "I" is dropped and "service," becomes service. Worth knowing before you debug a missing word.

fit_transform on training data, transform on test data. This distinction is the single most common bug in beginner NLP code. fit learns the vocabulary. Calling fit again on your test set rebuilds the vocabulary and quietly leaks test information into your model.

LogisticRegression learns one weight per vocabulary word, per class. A word that shows up mostly in positive reviews gets a positive weight. Prediction adds up the weights of the words present.

That is why the first test sentence lands on positive. The word tasty appears only in positive training rows, so it carries a positive weight, and nothing in the sentence pulls the other way.

Look at what the model learned

inspect.py
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression

texts = ["the food was delicious and the staff was kind",
         "wonderful service, I will come back again",
         "loved the biryani, best in the city",
         "great experience, very tasty food",
         "the food was cold and tasteless",
         "terrible service, we waited an hour",
         "worst biryani I have ever eaten",
         "bad experience, never coming back"]
labels = ["positive"] * 4 + ["negative"] * 4

vec = CountVectorizer()
clf = LogisticRegression(max_iter=1000).fit(vec.fit_transform(texts), labels)

words = list(vec.get_feature_names_out())
weights = clf.coef_[0]                     # one weight per vocabulary word

print("classes:", clf.classes_)
print("vocabulary size:", len(words))
for w in ["delicious", "tasty", "wonderful", "cold", "worst", "terrible"]:
    lean = clf.classes_[1] if weights[words.index(w)] > 0 else clf.classes_[0]
    print(f"  {w:10} leans {lean}")
Output
classes: ['negative' 'positive']
vocabulary size: 36
  delicious  leans positive
  tasty      leans positive
  wonderful  leans positive
  cold       leans negative
  worst      leans negative
  terrible   leans negative

There is a reason those six words were chosen and not, say, biryani. Each of the six appears in one class only, so its direction is settled by the data and will not move.

biryani appears once on each side. Its weight sits close to zero, and its sign can flip between scikit-learn versions. Printing it would look like a result and would be noise.

That is worth sitting with. Eight sentences is far too little data to learn stable weights for common words. If you inspect weights on a small dataset and read meaning into the near-zero ones, you are reading the random seed.

Common mistakes

Calling fit_transform on your test set. It rebuilds the vocabulary from test data. Your reported accuracy goes up and your production accuracy does not. Use fit_transform on train, transform everywhere else.

Reporting accuracy on imbalanced data. If 95 percent of your emails are not spam, a model that answers "not spam" every time scores 95 percent. Look at precision and recall per class, never accuracy alone.

Stripping punctuation without thinking. "Great." and "Great!!!" carry different heat. For sentiment work, punctuation is signal, not noise.

Reaching for a transformer first. Fine-tune a large model, get 0.91, and you have no idea whether that is good. Run the ten-line baseline first. If the baseline gets 0.89, the transformer bought you very little for a hundred times the cost.

Try it yourself

Add four more training sentences in Hinglish — for example "khana bahut accha tha" and "service ekdum bekaar". Then test on a Hinglish sentence the model has not seen.

Watch it fail. Then work out why. The vocabulary was built from English words, so a Hinglish word the model never saw contributes nothing at all. This is the exact problem that tokenization and embeddings were invented to fix.

What to learn next

Researcher — Mathematics and papers.

The core formulation

Most of modern NLP reduces to one object: a probability distribution over token sequences.

text
P(w_1, ..., w_T) = product over t = 1..T of P(w_t | w_1, ..., w_{t-1})
  • w_t is the token at position t, drawn from a finite vocabulary V.
  • T is the sequence length in tokens.
  • The factorisation is the chain rule of probability, exact and assumption-free.

The modelling problem is estimating each conditional P(w_t | w_{<t}). Everything from a bigram table to a frontier LLM is a different parameterisation of that same conditional.

Training objective

A model with parameters theta is fitted by minimising token-level cross-entropy:

text
L(theta) = -(1/T) * sum over t = 1..T of log P_theta(w_t | w_{<t})
  • theta is the full parameter vector.
  • log is the natural logarithm, so L is measured in nats per token.
  • The expectation is taken over the training corpus.

Perplexity is the exponentiated form, PPL = exp(L), reported in the same units as vocabulary size. A perplexity of 20 means the model is as uncertain as if choosing uniformly among 20 tokens at each step.

Two cautions. Perplexity is only comparable across models that share a tokenizer, since L is per token and tokenizers differ in how many tokens a sentence costs. And perplexity correlates with downstream quality only loosely — it is a training diagnostic, not an evaluation.

The classical approach and why it broke

The n-gram model truncates the conditioning context:

text
P(w_t | w_{<t}) ~= P(w_t | w_{t-n+1}, ..., w_{t-1})
  • n is the order, typically 3 to 5.

Maximum-likelihood estimation of this table has a parameter count of O(|V|^n), where |V| is vocabulary size. With |V| = 50000 and n = 5, that is 3 * 10^23 cells against a corpus of maybe 10^11 tokens. Almost every cell is zero.

This is data sparsity. It forced an entire literature on smoothing: Good-Turing, Katz back-off, and Kneser-Ney (Kneser & Ney, 1995). Kneser-Ney remained the strongest count-based method for two decades. Even with perfect smoothing, an n-gram model cannot represent that "chai" and "coffee" behave alike, because each is a separate atomic symbol.

The shift to distributed representations

Bengio et al. (2003), A Neural Probabilistic Language Model (JMLR 3), replaced the count table with a learned map from tokens to dense vectors, followed by a neural network. The parameter count becomes O(|V| * d + d * h) rather than O(|V|^n), where d is embedding width and h is hidden width. Similar words share statistical strength through nearby vectors, which is what solves sparsity rather than smoothing around it.

The trajectory from there:

YearWorkContribution
2013Mikolov et al., word2vecCheap embeddings from a shallow objective, at corpus scale
2014Sutskever, Vinyals & Le, seq2seqEncoder-decoder for variable-length output
2014Bahdanau, Cho & BengioAttention removes the fixed-size encoder bottleneck
2017Vaswani et al., TransformerAttention alone, fully parallel over sequence positions
2018Devlin et al., BERTBidirectional masked pre-training, then fine-tune
2020Brown et al., GPT-3In-context learning at 175B parameters
2022Ouyang et al., InstructGPTPreference optimisation aligns output with instructions

Cost

For a decoder-only Transformer with L layers, model width d, and sequence length n:

ComponentCost per forward pass
Feed-forward blocksO(L * n * d^2)
Self-attention scoresO(L * n^2 * d)
KV cache memoryO(L * n * d) per sequence

The quadratic attention term dominates once n exceeds roughly d. For d = 4096, that crossover sits near 4k tokens, which is why long-context work concentrates on attention rather than on the feed-forward blocks. See attention for the mitigations.

Training compute is approximated by C ~= 6 * N * D FLOPs, where N is non-embedding parameter count and D is training tokens. The factor of 6 counts one forward and two backward multiply-accumulates per parameter per token.

Key references

  • Shannon, C. (1948). A Mathematical Theory of Communication. Bell System Technical Journal. Origin of entropy and of predicting text as a measurable task.
  • Bengio, Y., Ducharme, R., Vincent, P. & Jauvin, C. (2003). A Neural Probabilistic Language Model. JMLR 3.
  • Mikolov, T. et al. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781
  • Sutskever, I., Vinyals, O. & Le, Q. (2014). Sequence to Sequence Learning with Neural Networks. arXiv:1409.3215
  • Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762
  • Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. (2018). BERT. arXiv:1810.04805
  • Brown, T. et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165
  • Kaplan, J. et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361
  • Hoffmann, J. et al. (2022). Training Compute-Optimal Large Language Models. arXiv:2203.15556
  • Ouyang, L. et al. (2022). Training language models to follow instructions with human feedback. arXiv:2203.02155

Current state and open problems

Kaplan et al. (2020) established smooth power-law relationships between loss, parameters and data. Hoffmann et al. (2022) corrected the exponent balance. Models of the era turned out to be badly under-trained relative to their size. The practical rule that followed was roughly 20 training tokens per parameter.

Four problems remain genuinely unsolved, and it is worth being precise about why.

Evaluation. Benchmark contamination is pervasive, since test sets leak into web-scale pre-training corpora. Reported gains on static benchmarks are therefore weak evidence. Held-out, time-gated and adversarially constructed evaluations are the current partial answer.

Factuality. Cross-entropy training rewards fluent continuations, not true ones. Nothing in the objective distinguishes a correct fact from a plausible-sounding one. Retrieval grounding reduces the failure rate but does not remove the incentive.

Long-context quality. Extending the context window is now routine. Demonstrating that a model uses the middle of that window as reliably as the ends is not — see Liu et al. (2023), Lost in the Middle.

Language coverage. Tokenizers, corpora and benchmarks are heavily English-weighted. Petrov et al. (2023) showed the same content costs several times more tokens in many Indic and African scripts. That translates directly into higher latency and higher price for those users. This is a measurable inequity, not a rounding error.

What to learn next

What to learn next

These follow on from what you just read.

  • Natural Language Processing

    Tokenization

    Tokenization is cutting text into small pieces called tokens, because a model can only work with a fixed list of known pieces.

  • Natural Language Processing

    Embeddings

    An embedding is a list of numbers that stands for a word or a sentence, arranged so that things with similar meaning end up close together.

  • Natural Language Processing

    Word2Vec

    Word2Vec learns a number list for every word by playing a guessing game about which words sit near which, then throws the game away and keeps the numbers.