What is machine learning?
Machine learning is how a computer works out a rule by looking at examples, instead of being handed the rule by a programmer.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Machine learning is how a computer works out a rule by looking at examples, instead of being handed the rule.
Think about how a child learns to pick a ripe mango. Nobody gives the child a formula. Someone hands them mango after mango and says "ripe" or "not yet". After thirty mangoes, the child picks correctly — and cannot explain how.
That last part matters. The child learned a rule nobody wrote down. Machine learning does the same thing, with data instead of mangoes.
Why it exists
For most of computing history, a person wrote down every rule by hand. This works beautifully for some jobs. Calculating salary is a rule. Booking a train seat is a rule.
It falls apart the moment the world gets messy.
Take spam email. An engineer once wrote a rule to block any mail containing "free money". Spammers wrote "fr33 m0ney" instead. The engineer added another rule. The spammers changed again. This chase has no end.
Now try something harder. Write down the rule for recognising the digit 2, in words. Not a picture — words. Your 2 has a loop at the bottom. Your friend's 2 is a sharp zigzag. Your grandmother's 2 has a curl.
Nobody has ever written that rule successfully. And nobody needs to. You can collect seventy thousand handwritten 2s instead, and let the computer work out the pattern.
That is the whole idea. Stop writing rules. Start collecting examples.
How it works
A machine learning system has two phases. First it learns. Later it answers.
PHASE 1 — LEARNING (happens once, takes time)
Many examples Learning A trained
with answers → algorithm → model
(photos labelled looks for (the learned
cat or dog) patterns rule)
PHASE 2 — USING IT (happens millions of times, fast)
A new photo → Trained → "dog"
never seen before model (quite sure)Three things are needed, always:
- Data — the examples. Without examples there is nothing to learn from.
- A model — the shape of the rule the computer is allowed to learn.
- A score — a way to measure how wrong the current guess is.
The computer makes a guess, checks the score, adjusts, and repeats. Millions of times. That loop is the entire trick.
Where you have already seen it
You use machine learning many times a day without noticing.
- Google Photos. You type "beach" and it finds your beach photos. Nobody tagged them. A model learned what beaches look like.
- Your bank's fraud alert. A payment at 3 a.m. from another city gets blocked. A model learned what your normal spending looks like.
- YouTube's homepage. It learned which videos people like you watch next.
- Your phone keyboard. It suggests the next word from what millions of people typed.
- Google Translate. It learned from millions of sentences translated by humans.
An honest warning
Machine learning is not magic, and it is not always the right tool.
If a rule is clear and stable, write the rule. Do not train a model to check whether a number is even. That would be slower, less reliable, and harder to fix.
Machine learning also fails in ways that are hard to predict. A model trained on daytime photos may not work at night. A model trained on one city's customers may be wrong in another. It learns what is in the data — including the mistakes in the data.
Remember this
- Machine learning finds rules from examples, rather than being given rules by a programmer.
- It always needs three things — data, a model, and a score that measures being wrong.
- It produces likely answers, not guaranteed ones. That uncertainty never fully goes away.
What to learn next
- Supervised learning — learning from examples that come with the correct answers.
- Unsupervised learning — finding groups in data when no answers are given.
- Linear regression — the smallest real model, and the best one to learn first.
Developer — Code and libraries.
The clearest way to feel what machine learning is: write no rule, and watch the computer write one for you.
Setup
pip install scikit-learnThat is the only install. scikit-learn is the standard Python library for classical machine learning, meaning everything that is not deep neural networks. It runs on a CPU and needs no GPU.
Minimal runnable code
We will sort fruit. Each fruit has two features — features are the measured clues we give the model. Here they are weight and how rough the skin feels.
from sklearn.tree import DecisionTreeClassifier, export_text
# Each row is one fruit: [weight in grams, skin roughness score out of 10]
fruits = [[210, 2], [185, 3], [230, 1], [95, 7], [112, 8], [80, 6]]
# 1 means mango, 0 means lemon
labels = [1, 1, 1, 0, 0, 0]
model = DecisionTreeClassifier(random_state=0)
model.fit(fruits, labels)
# Two fruits the model has never seen before
new_fruits = [[200, 2], [100, 7]]
print("predictions:", model.predict(new_fruits))
# The rule the computer wrote by itself, printed in plain text
print(export_text(model, feature_names=["weight", "roughness"]))predictions: [1 0] |--- roughness <= 4.50 | |--- class: 1 |--- roughness > 4.50 | |--- class: 0
What actually happened
Read that output again. Nowhere in the code did anyone write if roughness < 4.5. The computer found the number 4.50 by itself, from six rows of data.
Line by line, the parts that are not obvious:
fit(fruits, labels)is the learning step. Every scikit-learn model uses this same method name. You will type.fit()for the rest of your career.random_state=0fixes the internal randomness. Decision trees break ties randomly. Setting this means you and I see identical output — leave it out and results can shift between runs.predict()takes a list of rows, not a single row. That is why the input is a list inside a list.export_text()prints the learned rule. Most models cannot do this. A decision tree can, which is why it is a good first model to look at.
Common mistakes
Passing a single row without the outer list. model.predict([200, 2]) raises a ValueError about a 2D array being expected. scikit-learn always wants rows-by-columns, even for one row. Write model.predict([[200, 2]]).
Calling predict() before fit(). You get NotFittedError. The model object exists, but it has learned nothing yet.
Passing text straight in. [["heavy", "smooth"]] fails. Classical models want numbers. Text has to be converted first, usually with OneHotEncoder or OrdinalEncoder.
Believing this model. Six rows is not a dataset. The tree found one rule that separates six fruits and ignored weight completely. Show it a rough-skinned mango and it fails. Small data produces confident, wrong models — a trap covered in overfitting and underfitting.
Try it yourself
Add one rough-skinned mango to the data: append [220, 8] to fruits and 1 to labels. Run it again and print the rule.
The tree will now have to use weight as well as roughness. Watch how the printed rule grows a second branch. That growth is the model responding to harder data — and it is the beginning of every real training story.
What to learn next
- Supervised learning — the setting this example belongs to, formalised.
- Linear regression — predict a number instead of a category.
- Model evaluation — how to check whether a model is any good.
Researcher — Mathematics and papers.
The formal setting
Statistical learning theory frames learning as risk minimisation under an unknown distribution.
Formulas here are written in plain text, since the site renders no maths typesetting library.
Let X be the input space (e.g. R^d)
Let Y be the output space ({0,1} for binary, R for regression)
Let D be an unknown joint distribution over X x Y
Draw a sample S = {(x_1, y_1), ..., (x_n, y_n)} i.i.d. from D
A hypothesis h : X -> Y, drawn from a hypothesis class H
A loss function L : Y x Y -> R>=0The quantity we care about is the true risk (also called expected risk or generalisation error):
R(h) = E_{(x,y) ~ D} [ L(h(x), y) ]R(h)— average loss of hypothesishover every example the world could produceE— expectation, the probability-weighted averageD— the data-generating distribution, permanently unknown to usL— the loss, how much a wrong answer costs
R(h) is not computable, because D is unknown. We substitute the empirical risk measured on our finite sample:
R_emp(h) = (1/n) * SUM_{i=1..n} L( h(x_i), y_i )n— number of training examples(x_i, y_i)— the i-th training pair
Choosing the hypothesis that minimises this is Empirical Risk Minimisation (ERM):
h_hat = argmin_{h in H} R_emp(h)The central question of the field is then the generalisation gap, R(h_hat) - R_emp(h_hat). Training measures the second term. Deployment pays the first.
Why generalisation is possible at all
Uniform convergence gives a distribution-free bound. Take a hypothesis class of VC dimension d. Draw a sample of size n. Then with probability at least 1 - delta, this holds for every h in H:
R(h) <= R_emp(h) + sqrt( ( d * (log(2n/d) + 1) + log(4/delta) ) / n )d— VC dimension, the largest set size thatHcan shatter (label in every possible way)n— sample sizedelta— the probability that the bound fails
Two readings follow. Richer hypothesis classes (larger d) need more data. And the gap shrinks on the order of sqrt(d/n).
This bound is honest but loose. For modern over-parameterised networks, d exceeds n by orders of magnitude. The bound goes vacuous while the models generalise well anyway. That contradiction is an open research area, not a settled matter. See overfitting and underfitting for the double-descent literature.
No Free Lunch
Wolpert (1996) proved that, averaged over all possible target functions, every learning algorithm has identical expected off-training-set error. There is no universally best learner.
The practical consequence is that all useful learning depends on inductive bias. That means assumptions built into H about which functions are plausible. Convolutions assume translation equivariance. Linear models assume additivity. Choosing a model is choosing an assumption, never avoiding one.
Cost
For a training set of n examples with d features:
| Method | Training cost | Prediction cost |
|---|---|---|
| Linear regression, normal equations | O(n d^2 + d^3) | O(d) |
| Linear model, gradient descent | O(n d) per epoch | O(d) |
| k-nearest neighbours | O(1) | O(n d) per query |
| Decision tree (CART) | O(n d log n) typical | O(depth) |
k-nearest neighbours illustrates that "training cost" alone is a misleading metric. It trains instantly and predicts slowly, which is the wrong trade for most production systems.
Key references
- Valiant, L. (1984). A Theory of the Learnable. CACM 27(11). Introduces the PAC (Probably Approximately Correct) framework.
- Vapnik, V. & Chervonenkis, A. (1971). On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities. Origin of VC dimension.
- Wolpert, D. (1996). The Lack of A Priori Distinctions Between Learning Algorithms. Neural Computation 8(7).
- Mitchell, T. (1997). Machine Learning. McGraw-Hill. Source of the standard task/experience/performance formulation of learning.
- Shalev-Shwartz, S. & Ben-David, S. (2014). Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press. Freely available from the authors, and the best single entry point to the theory above.
Current state
Classical statistical learning theory describes the small-n, small-d regime accurately. It describes deep learning poorly.
Zhang et al. (2017), Understanding Deep Learning Requires Rethinking Generalisation, made this concrete. Standard networks can fit random labels perfectly, and still generalise well on real labels. Uniform-convergence bounds cannot explain that.
Active alternatives include PAC-Bayes bounds, margin-based analyses, implicit regularisation of stochastic gradient descent, and the neural tangent kernel. None of these yet gives a complete account. Treat any confident claim that generalisation in deep learning is "solved" with suspicion.
What to learn next
- Overfitting and underfitting — the bias-variance decomposition and double descent.
- Model evaluation — estimators of
R(h)and their variance. - Optimization — how ERM is actually solved in practice.