What is a neural network?
A neural network is a stack of tiny decision-makers that learns patterns from solved examples instead of following rules you wrote by hand.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A neural network is a machine that learns patterns from examples, instead of following rules a person wrote for it.
Think about buying a ripe mango. You press it gently with your thumb. You look at the colour near the stem, and you hold it close and smell it. Nobody handed you a rulebook for this. You learned it by handling hundreds of mangoes and tasting the results.
A neural network learns the same way. You show it many examples with the answer attached. It slowly works out which clues matter.
Why this had to be invented
For a long time, people tried to write the rules by hand. To find a cat in a photo, a programmer would write instructions about whiskers, pointed ears and fur.
That approach kept breaking. A cat lying down looks nothing like a cat standing up. A black cat at night has no visible fur texture. Every new photo needed a new rule, and the rules started fighting each other.
The honest problem was this: people know how to recognise a cat, but they cannot explain how they do it. You cannot write down a rule you were never conscious of following.
Neural networks flip the job around. You stop writing rules. You collect examples, and the machine finds the rules on its own.
The one idea underneath it
The whole thing is built from one small part, repeated. That part is called a neuron — a tiny decision-maker that takes in several clues and produces one number.
Each clue arriving at a neuron passes through an importance knob. That knob is called a weight, and it decides how much that clue counts. A strong smell might get a big weight, and the sticker price might get almost none.
The neuron adds up all the clues after their knobs. Then it passes the total through a bend: a small rule that stops the answer being a straight line. Without that bend, stacking many neurons gives you nothing more than one neuron.
How it works
Neurons are arranged in layers: rows of neurons that all work at the same time. Each row hands its answers to the next row.
layer 1 layer 2
softness → ( neuron ) ─┐
├→ ( neuron ) → "ripe" (91% sure)
colour → ( neuron ) ─┤
│
smell → ( neuron ) ─┘Information walks left to right. Every arrow carries a number, and every arrow has its own importance knob.
The first layer notices small, dull things — an edge here, a patch of colour there. The next layer combines those into bigger things, like a shape. A later layer combines shapes into "this is a face".
When a network has many layers stacked like this, we call it deep learning. That is the only thing the word "deep" means here — a lot of layers, one after another.
This next part is the one that surprises everyone. Nobody chooses the knob settings. They start as meaningless numbers and get corrected, a little at a time, over thousands of examples. That correction process is the real magic, and it gets its own lesson.
Where you have already used one
You have used neural networks today, probably without noticing.
- Google Photos lets you search your own gallery for "beach" without you ever tagging a photo.
- UPI and bank apps flag a payment as suspicious within a second of you tapping pay.
- Google Translate turns a photo of a signboard into your language, live through the camera.
- YouTube decides which video to autoplay next.
None of these were built by someone writing down what a beach looks like.
What is honestly hard here
A neural network cannot explain itself. It can tell you a photo is a cat, with high confidence. It cannot tell you why, in words a person finds satisfying.
This is a real, unsolved limitation, not something we are hiding from you. It matters a great deal when these systems are used for loans, medical scans or hiring.
Remember this
- A neural network learns from examples with answers attached, not from rules you write.
- It is built from neurons — tiny decision-makers wired together in layers.
- Every connection has an importance knob, and learning means tuning those knobs.
What to learn next
- How neural networks learn — where the weights actually come from.
- Activation functions — why that bend between layers is not optional.
- Gradient descent — the downhill walk that does the tuning.
Developer — Code and libraries.
A neural network, stripped of all the vocabulary, is this repeated three times:
multiply by a weight matrix → add a bias → apply a non-linear functionThat is the entire forward pass. Let us build one with no framework at all, so nothing is hidden.
Setup
pip install numpyThat is the only dependency. This runs on any laptop in well under a second, and no GPU is involved.
A network you can read end to end
The weights below are written by hand rather than randomised. That is not how real training starts, but it means your output will match this page exactly, which matters when you are learning.
import numpy as np
# Three mangoes. Two clues each, scaled 0 to 1: how soft it feels, how yellow it looks.
mangoes = np.array([
[0.9, 0.8], # soft and yellow
[0.1, 0.2], # hard and green
[0.8, 0.1], # soft but still green
])
# Layer 1 weights: 2 clues -> 3 neurons. Fixed by hand so this file prints the same numbers for you.
W1 = np.array([[2.0, -1.5, 0.5],
[1.0, 2.0, -1.0]])
b1 = np.array([-0.5, -0.2, 0.1])
# Layer 2 weights: 3 neurons -> 1 final score.
W2 = np.array([[1.2], [0.8], [-0.7]])
b2 = np.array([-0.3])
def relu(x):
return np.maximum(0.0, x) # the bend: anything negative is flattened to zero
def sigmoid(x):
return 1.0 / (1.0 + np.exp(-x)) # squeezes any number into the range 0 to 1
hidden = relu(mangoes @ W1 + b1) # @ is matrix multiply: every clue reaches every neuron
score = sigmoid(hidden @ W2 + b2)
print("hidden layer:")
print(np.round(hidden, 3))
print("ripeness score (0 = unripe, 1 = ripe):")
print(np.round(score, 3))
print("parameters in this network:", W1.size + b1.size + W2.size + b2.size)hidden layer: [[2.1 0.05 0. ] [0. 0.05 0. ] [1.2 0. 0.4 ]] ripeness score (0 = unripe, 1 = ripe): [[0.906] [0.435] [0.703]] parameters in this network: 13
The soft, yellow mango scored 0.906. The hard, green one scored 0.435. The network has never been trained, so these numbers come purely from the weights typed above — but the structure that produces them is exactly the structure a trained network uses.
Line by line, the parts that are not obvious
mangoes @ W1 — one matrix multiply handles all three mangoes and all three neurons at once. Shape (3, 2) @ (2, 3) gives (3, 3): three mangoes, three neuron outputs each. Batching is not an optimisation you add later; it falls out of the shapes for free.
+ b1 — b1 has shape (3,) and the left side has shape (3, 3). NumPy broadcasts it, meaning it stretches the smaller array across every row. The bias is the neuron's baseline: it shifts the decision point, so a neuron can lean "yes" even when every input is zero.
relu — read the hidden layer output again. The third neuron produced 0.0 for two of the three mangoes. ReLU flattened a negative number to zero, so that neuron stayed silent for those inputs. Different inputs wake up different neurons, and that selectivity is where a network's power comes from.
sigmoid — the hidden layer emits any number at all. The final score needs to read as a confidence, so sigmoid squashes it into a value between zero and one.
Common mistakes
Getting the weight matrix shape backwards. If W1 is written as (3, 2) when it should be (2, 3), mangoes @ W1 raises:
ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 3 is different from 2)
The fix: a weight matrix is always (inputs_coming_in, neurons_going_out). Print .shape on both sides before the multiply — this one error costs beginners more hours than any other.
Leaving out the activation. Delete relu and the two layers collapse into a single layer, no matter how many you stack. There is a proof of this you can run yourself in activation-functions.
Feeding raw, unscaled numbers. Send in a price of 45000 alongside a softness of 0.8 and np.exp overflows. You get RuntimeWarning: overflow encountered in exp and a saturated output of 0.0 or 1.0 that never recovers. Scale your inputs to a similar range before they enter the network.
Reading an untrained network's output as meaningful. These scores come from hand-typed numbers. An untrained network is a random opinion machine; it becomes useful only after training.
Try it yourself
Change the third mango to [0.2, 0.9] — hard, but very yellow. Predict the score before you run it, then check. Then set the whole W2 column to zeros and see what every mango scores, and work out why they are all identical.
What to learn next
- How neural networks learn — where the weights actually come from.
- Activation functions — why that bend between layers is not optional.
- Gradient descent — the downhill walk that does the tuning.
Researcher — Mathematics and papers.
Formal definition
A feedforward neural network (a multilayer perceptron, MLP) is a parameterised function built by composing affine maps with pointwise non-linearities:
$$ h^{(0)} = x, \qquad h^{(l)} = \phi^{(l)}!\left(W^{(l)} h^{(l-1)} + b^{(l)}\right), \qquad f_\theta(x) = h^{(L)} $$
Where:
- $x \in \mathbb{R}^{d_0}$ — the input vector, with $d_0$ features.
- $L$ — the number of layers; $l \in {1, \dots, L}$ indexes them.
- $W^{(l)} \in \mathbb{R}^{d_l \times d_{l-1}}$ — the weight matrix of layer $l$.
- $b^{(l)} \in \mathbb{R}^{d_l}$ — the bias vector of layer $l$.
- $\phi^{(l)}$ — a non-linear function applied elementwise (ReLU, tanh, GELU).
- $h^{(l)} \in \mathbb{R}^{d_l}$ — the activations, or hidden representation, at layer $l$.
- $\theta = {W^{(l)}, b^{(l)}}_{l=1}^{L}$ — every trainable parameter, collected.
The non-linearity is what makes the composition non-trivial. If every $\phi^{(l)}$ were the identity, then $f_\theta(x) = W^{(L)} \cdots W^{(1)} x + \tilde{b}$, which is a single affine map of rank at most $\min_l d_l$. Depth would buy nothing but a rank constraint.
Cost
For a layer mapping $d_{l-1} \to d_l$ with batch size $B$:
- Parameters: $d_l d_{l-1} + d_l$.
- Forward FLOPs: $\approx 2 B d_l d_{l-1}$ (one multiply and one add per weight per sample).
- Backward FLOPs: $\approx 2 \times$ the forward cost, since gradients flow to both inputs and weights.
The rule of thumb for training a dense model is six FLOPs per parameter per token, decomposed as two forward and four backward.
Activation memory, not parameter memory, is usually the binding constraint during training: storing $h^{(l)}$ for the backward pass costs $O(B \sum_l d_l)$.
Universal approximation, stated honestly
Cybenko (1989), Approximation by superpositions of a sigmoidal function, and Hornik (1991), Approximation capabilities of multilayer feedforward networks, establish that a network with one hidden layer and a non-polynomial activation is dense in $C(K)$ for compact $K \subset \mathbb{R}^n$, under the sup norm.
Three caveats matter, and they are routinely dropped when this theorem is quoted:
- It is an existence result. It asserts that suitable weights exist; it says nothing about gradient descent finding them.
- The required width can be exponential in the input dimension. Existence at any cost is a weak guarantee.
- It says nothing about generalisation — approximating a function on the training set is not the same as approximating the data-generating distribution.
Depth genuinely helps, and there are separation theorems that prove it. Telgarsky (2016), Benefits of depth in neural networks, exhibits functions computable by a network of depth $k$ that require width exponential in $k$ at depth $O(1)$. Eldan and Shamir (2016) give a $3$-layer versus $2$-layer separation.
Parameter counting in practice
For an MLP with layer widths $[d_0, d_1, \dots, d_L]$, total parameters are $\sum_{l=1}^{L} (d_l d_{l-1} + d_l)$.
A concrete case: widths $[784, 256, 128, 10]$ gives $784 \cdot 256 + 256 = 200{,}960$, plus $256 \cdot 128 + 128 = 32{,}896$, plus $128 \cdot 10 + 10 = 1{,}290$ — a total of $235{,}146$ parameters. Note that the first layer holds 85% of them, which is why input dimensionality dominates small dense models.
The biological analogy is weak
The "neuron" name is historical. McCulloch and Pitts (1943) proposed a threshold logic unit as a model of neural activity, and Rosenblatt (1958) built the perceptron on that framing.
Modern artificial neurons diverge sharply from biology: real neurons communicate in discrete spikes with timing-dependent plasticity, there is no known biological mechanism implementing symmetric weight transport for backpropagation (the "weight transport problem", raised by Grossberg, 1987), and cortical connectivity is nothing like dense all-to-all layers. Treat the name as a label, not as evidence.
Where to read next
- Rumelhart, Hinton and Williams (1986), Learning representations by back-propagating errors — the paper that made depth trainable.
- Goodfellow, Bengio and Courville, Deep Learning (2016), chapter 6 — the standard reference treatment of feedforward networks.
- Zhang et al. (2017), Understanding deep learning requires rethinking generalization — networks fit random labels perfectly, so classical capacity arguments do not explain why they generalise.
What to learn next
- How neural networks learn — where the weights actually come from.
- Activation functions — why that bend between layers is not optional.
- Gradient descent — the downhill walk that does the tuning.