What is NLP?
Natural language processing is how a computer reads, understands and writes human language instead of only handling numbers.
- 14 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Natural language processing, or NLP, is teaching a computer to work with human language.
That means reading it, making sense of it, and writing it back.
Think about your phone keyboard. You type "I am on my" and it offers "way". Nobody sat down and wrote a rule for that exact sentence. The keyboard learned it by reading millions of messages that other people typed first.
That small guessing trick is NLP in its most everyday form. Everything else in this section is the same idea, grown much larger.
Why it exists
Computers were built for numbers. A bank balance, a distance, a date — these have one meaning and one shape. Language has neither.
Take the word "bank". River bank. Money bank. Bank a turn on a bike. Same four letters, three different worlds. You picked the right one instantly because of the words around it.
Now try to write that rule down for a computer. People tried for forty years. They wrote grammar rules by hand, thousands of them. Every rule fixed one sentence and broke two others.
Here is a famous example. Someone built an early translation system in the 1960s. It turned "the spirit is willing but the flesh is weak" into a sentence about vodka and rotten meat. The words were translated correctly. The meaning was gone.
The breakthrough was giving up on hand-written rules. Instead of telling the computer what language means, we show it enormous amounts of real language. Then we let it find the patterns on its own. That switch is what made translation, voice assistants and chatbots possible.
How it works
Every NLP system, from your keyboard to the biggest chatbot, runs the same four steps.
"the chai is hot"
|
v
[1] CHOP into pieces -> "the" "chai" "is" "hot"
|
v
[2] TURN each piece -> every piece becomes a
into a list of numbers list of numbers
|
v
[3] MODEL does its work -> finds patterns across
the whole sentence
|
v
[4] TURN numbers back -> "यह चाय गरम है"
into words or a label or "positive" or an answerStep one is called tokenization, which means cutting text into small chunks. Step two is called embedding, which means giving each chunk a list of numbers that stands for its meaning. Both get their own lesson.
The important thing to hold on to is this. A model never sees letters. It sees numbers all the way through, and words only appear again at the very end.
Where you have already seen it
You use NLP many times a day, mostly without noticing.
- Your spam folder. Nobody reads your mail to sort it. A model learned what spam wording looks like.
- Google Translate. Point your camera at a signboard and the words change language.
- YouTube auto-captions. Speech becomes text, then text gets punctuated and cleaned.
- Search that forgives typos. You type "recipy for pav bhaji" and still get the right page.
- Any chatbot. It reads your question and writes a fresh answer, word by word.
The part that is genuinely hard
Language carries things that are not in the words at all.
If a friend says "wow, great timing" when you arrive an hour late, every word is positive. The meaning is not. Sarcasm, politeness, hints and jokes live in shared context, not in the sentence.
Models today handle this far better than five years ago, and still get it wrong. This is not a bug waiting to be fixed next month. It is the honest edge of the field. Anyone who tells you language is a solved problem is selling something.
Remember this
- NLP is any system that reads or writes human language, from a keyboard suggestion to a chatbot.
- Words are turned into numbers first, and only turned back into words at the very end.
- Hand-written grammar rules failed. Learning patterns from huge amounts of real text is what worked.
What to learn next
- Tokenization — how a sentence gets chopped into pieces.
- Embeddings — how those pieces become numbers that carry meaning.
- What is machine learning? — the learning idea underneath all of this.
Developer — Code and libraries.
You do not need a large model to do real NLP. A word-counting model plus logistic regression solves a surprising number of production problems, and it runs on any laptop in under a second.
Build that first. It gives you a baseline number, and a baseline is the only thing that tells you whether a bigger model was worth it.
Setup
pip install scikit-learnThat is the whole dependency list. No GPU, no downloads at runtime.
A working sentiment classifier
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression
train_texts = [
"the food was delicious and the staff was kind",
"wonderful service, I will come back again",
"loved the biryani, best in the city",
"great experience, very tasty food",
"the food was cold and tasteless",
"terrible service, we waited an hour",
"worst biryani I have ever eaten",
"bad experience, never coming back",
]
train_labels = ["positive"] * 4 + ["negative"] * 4
vec = CountVectorizer() # text -> a table of word counts
X = vec.fit_transform(train_texts)
clf = LogisticRegression(max_iter=1000) # extra passes so the fit finishes cleanly
clf.fit(X, train_labels)
tests = ["the biryani was tasty", "we waited an hour for cold food"]
for text, guess in zip(tests, clf.predict(vec.transform(tests))):
print(f"{guess:8} <- {text}")positive <- the biryani was tasty negative <- we waited an hour for cold food
Eight training sentences. Two correct predictions on sentences the model never saw.
Line by line
CountVectorizer() builds a vocabulary from the training text, then turns each sentence into a row of counts. If the vocabulary has 40 words, every sentence becomes 40 numbers. Most of them are zero.
By default it lowercases everything and keeps only tokens of two or more word characters. So "I" is dropped and "service," becomes service. Worth knowing before you debug a missing word.
fit_transform on training data, transform on test data. This distinction is the single most common bug in beginner NLP code. fit learns the vocabulary. Calling fit again on your test set rebuilds the vocabulary and quietly leaks test information into your model.
LogisticRegression learns one weight per vocabulary word, per class. A word that shows up mostly in positive reviews gets a positive weight. Prediction adds up the weights of the words present.
That is why the first test sentence lands on positive. The word tasty appears only in positive training rows, so it carries a positive weight, and nothing in the sentence pulls the other way.
Look at what the model learned
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression
texts = ["the food was delicious and the staff was kind",
"wonderful service, I will come back again",
"loved the biryani, best in the city",
"great experience, very tasty food",
"the food was cold and tasteless",
"terrible service, we waited an hour",
"worst biryani I have ever eaten",
"bad experience, never coming back"]
labels = ["positive"] * 4 + ["negative"] * 4
vec = CountVectorizer()
clf = LogisticRegression(max_iter=1000).fit(vec.fit_transform(texts), labels)
words = list(vec.get_feature_names_out())
weights = clf.coef_[0] # one weight per vocabulary word
print("classes:", clf.classes_)
print("vocabulary size:", len(words))
for w in ["delicious", "tasty", "wonderful", "cold", "worst", "terrible"]:
lean = clf.classes_[1] if weights[words.index(w)] > 0 else clf.classes_[0]
print(f" {w:10} leans {lean}")classes: ['negative' 'positive'] vocabulary size: 36 delicious leans positive tasty leans positive wonderful leans positive cold leans negative worst leans negative terrible leans negative
There is a reason those six words were chosen and not, say, biryani. Each of the six appears in one class only, so its direction is settled by the data and will not move.
biryani appears once on each side. Its weight sits close to zero, and its sign can flip between scikit-learn versions. Printing it would look like a result and would be noise.
That is worth sitting with. Eight sentences is far too little data to learn stable weights for common words. If you inspect weights on a small dataset and read meaning into the near-zero ones, you are reading the random seed.
Common mistakes
Calling fit_transform on your test set. It rebuilds the vocabulary from test data. Your reported accuracy goes up and your production accuracy does not. Use fit_transform on train, transform everywhere else.
Reporting accuracy on imbalanced data. If 95 percent of your emails are not spam, a model that answers "not spam" every time scores 95 percent. Look at precision and recall per class, never accuracy alone.
Stripping punctuation without thinking. "Great." and "Great!!!" carry different heat. For sentiment work, punctuation is signal, not noise.
Reaching for a transformer first. Fine-tune a large model, get 0.91, and you have no idea whether that is good. Run the ten-line baseline first. If the baseline gets 0.89, the transformer bought you very little for a hundred times the cost.
Try it yourself
Add four more training sentences in Hinglish — for example "khana bahut accha tha" and "service ekdum bekaar". Then test on a Hinglish sentence the model has not seen.
Watch it fail. Then work out why. The vocabulary was built from English words, so a Hinglish word the model never saw contributes nothing at all. This is the exact problem that tokenization and embeddings were invented to fix.
What to learn next
- Tokenization — replace word counting with subword pieces.
- Text classification — the full pipeline, metrics and pitfalls.
- Embeddings — dense vectors that survive unseen words.
Researcher — Mathematics and papers.
The core formulation
Most of modern NLP reduces to one object: a probability distribution over token sequences.
P(w_1, ..., w_T) = product over t = 1..T of P(w_t | w_1, ..., w_{t-1})w_tis the token at positiont, drawn from a finite vocabularyV.Tis the sequence length in tokens.- The factorisation is the chain rule of probability, exact and assumption-free.
The modelling problem is estimating each conditional P(w_t | w_{<t}). Everything from a bigram table to a frontier LLM is a different parameterisation of that same conditional.
Training objective
A model with parameters theta is fitted by minimising token-level cross-entropy:
L(theta) = -(1/T) * sum over t = 1..T of log P_theta(w_t | w_{<t})thetais the full parameter vector.logis the natural logarithm, soLis measured in nats per token.- The expectation is taken over the training corpus.
Perplexity is the exponentiated form, PPL = exp(L), reported in the same units as vocabulary size. A perplexity of 20 means the model is as uncertain as if choosing uniformly among 20 tokens at each step.
Two cautions. Perplexity is only comparable across models that share a tokenizer, since L is per token and tokenizers differ in how many tokens a sentence costs. And perplexity correlates with downstream quality only loosely — it is a training diagnostic, not an evaluation.
The classical approach and why it broke
The n-gram model truncates the conditioning context:
P(w_t | w_{<t}) ~= P(w_t | w_{t-n+1}, ..., w_{t-1})nis the order, typically 3 to 5.
Maximum-likelihood estimation of this table has a parameter count of O(|V|^n), where |V| is vocabulary size. With |V| = 50000 and n = 5, that is 3 * 10^23 cells against a corpus of maybe 10^11 tokens. Almost every cell is zero.
This is data sparsity. It forced an entire literature on smoothing: Good-Turing, Katz back-off, and Kneser-Ney (Kneser & Ney, 1995). Kneser-Ney remained the strongest count-based method for two decades. Even with perfect smoothing, an n-gram model cannot represent that "chai" and "coffee" behave alike, because each is a separate atomic symbol.
The shift to distributed representations
Bengio et al. (2003), A Neural Probabilistic Language Model (JMLR 3), replaced the count table with a learned map from tokens to dense vectors, followed by a neural network. The parameter count becomes O(|V| * d + d * h) rather than O(|V|^n), where d is embedding width and h is hidden width. Similar words share statistical strength through nearby vectors, which is what solves sparsity rather than smoothing around it.
The trajectory from there:
| Year | Work | Contribution |
|---|---|---|
| 2013 | Mikolov et al., word2vec | Cheap embeddings from a shallow objective, at corpus scale |
| 2014 | Sutskever, Vinyals & Le, seq2seq | Encoder-decoder for variable-length output |
| 2014 | Bahdanau, Cho & Bengio | Attention removes the fixed-size encoder bottleneck |
| 2017 | Vaswani et al., Transformer | Attention alone, fully parallel over sequence positions |
| 2018 | Devlin et al., BERT | Bidirectional masked pre-training, then fine-tune |
| 2020 | Brown et al., GPT-3 | In-context learning at 175B parameters |
| 2022 | Ouyang et al., InstructGPT | Preference optimisation aligns output with instructions |
Cost
For a decoder-only Transformer with L layers, model width d, and sequence length n:
| Component | Cost per forward pass |
|---|---|
| Feed-forward blocks | O(L * n * d^2) |
| Self-attention scores | O(L * n^2 * d) |
| KV cache memory | O(L * n * d) per sequence |
The quadratic attention term dominates once n exceeds roughly d. For d = 4096, that crossover sits near 4k tokens, which is why long-context work concentrates on attention rather than on the feed-forward blocks. See attention for the mitigations.
Training compute is approximated by C ~= 6 * N * D FLOPs, where N is non-embedding parameter count and D is training tokens. The factor of 6 counts one forward and two backward multiply-accumulates per parameter per token.
Key references
- Shannon, C. (1948). A Mathematical Theory of Communication. Bell System Technical Journal. Origin of entropy and of predicting text as a measurable task.
- Bengio, Y., Ducharme, R., Vincent, P. & Jauvin, C. (2003). A Neural Probabilistic Language Model. JMLR 3.
- Mikolov, T. et al. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781
- Sutskever, I., Vinyals, O. & Le, Q. (2014). Sequence to Sequence Learning with Neural Networks. arXiv:1409.3215
- Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762
- Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. (2018). BERT. arXiv:1810.04805
- Brown, T. et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165
- Kaplan, J. et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361
- Hoffmann, J. et al. (2022). Training Compute-Optimal Large Language Models. arXiv:2203.15556
- Ouyang, L. et al. (2022). Training language models to follow instructions with human feedback. arXiv:2203.02155
Current state and open problems
Kaplan et al. (2020) established smooth power-law relationships between loss, parameters and data. Hoffmann et al. (2022) corrected the exponent balance. Models of the era turned out to be badly under-trained relative to their size. The practical rule that followed was roughly 20 training tokens per parameter.
Four problems remain genuinely unsolved, and it is worth being precise about why.
Evaluation. Benchmark contamination is pervasive, since test sets leak into web-scale pre-training corpora. Reported gains on static benchmarks are therefore weak evidence. Held-out, time-gated and adversarially constructed evaluations are the current partial answer.
Factuality. Cross-entropy training rewards fluent continuations, not true ones. Nothing in the objective distinguishes a correct fact from a plausible-sounding one. Retrieval grounding reduces the failure rate but does not remove the incentive.
Long-context quality. Extending the context window is now routine. Demonstrating that a model uses the middle of that window as reliably as the ends is not — see Liu et al. (2023), Lost in the Middle.
Language coverage. Tokenizers, corpora and benchmarks are heavily English-weighted. Petrov et al. (2023) showed the same content costs several times more tokens in many Indic and African scripts. That translates directly into higher latency and higher price for those users. This is a measurable inequity, not a rounding error.
What to learn next
- Attention — the mechanism the whole modern stack rests on.
- Embeddings — the distributional theory and its limits.
- How LLMs actually work — from this objective to a deployed system.