Vectors and matrices
A vector is an ordered list of measurements about one thing. A matrix is many of those lists stacked into a table. Almost everything inside an AI model is one of these two.
- 16 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A vector is an ordered list of measurements about one thing. A matrix is many of those lists stacked into a table.
Think about the card a tailor writes when you order a shirt. Chest, waist, sleeve, collar — the same measurements, always in the same order, one card per customer.
That card is a vector. The tailor's whole register, one row per customer, is a matrix. You have already used both; nobody told you their names.
Why you should care
A computer cannot see a mango, hear a song, or read a word. It has no eyes and no ears.
So every AI system starts by turning the thing in front of it into a list of measurements. A photo becomes a list. A sentence becomes a list. A song becomes a list. From that moment on, the model is working with lists and tables and nothing else.
Understanding these two objects means understanding the material every model is made of.
The second meaning
A vector carries one more idea: direction.
Tell a friend "the shop is that way, about ten minutes' walk". You have given two things at once — which way, and how far. That pair is also a vector. An arrow, not a point.
Both meanings are the same object seen from two sides. A list of measurements is an arrow, pointing into a space with one direction for each measurement. This is the sentence that takes a while to land. Read it twice; almost everyone has to.
How it works
One shirt order → chest, waist, sleeve, collar → a VECTOR (one row)
The whole register:
chest waist sleeve collar
Asha ... ... ... ...
Ravi ... ... ... ... → a MATRIX (many rows)
Meera ... ... ... ...Two useful things fall out of this arrangement.
Closeness. Two customers whose cards look alike need similar shirts. Comparing two vectors tells you how alike two things are. That is the whole engine behind "people who liked this also liked".
mango → sweet, fruit, small, grows on a tree
banana → sweet, fruit, small, grows on a tree → close together
laptop → metal, screen, heavy, made in a factory → far awayTransformation. A matrix is also a machine. Feed a list in one side and a different list comes out the other side.
a list going in → [ the matrix ] → a different list coming outA black-and-white photo filter is exactly this. Each dot in your photo has a red amount, a green amount and a blue amount. The filter turns those three into one grey amount, using the same fixed rule for every dot. That rule is a small matrix.
A real example you have seen
Face unlock on your phone. Your face is measured once and stored as a single list of numbers — not as a picture. When you look at the phone, it measures again. Then it checks whether the new list sits close to the stored one.
Music apps do the same with songs. Each song becomes a list. Songs whose lists sit near each other get recommended together. Nobody wrote a rule saying these two songs are similar; the lists put them near each other on their own.
An honest word
The word "matrix" frightens people, partly because of the film and partly because of how it is taught. Underneath, it is a table. Rows and columns, like a register.
There is one genuinely awkward part ahead. Combining two matrices is not a matter of multiplying the cells that line up. The rule is stranger, and nearly everyone finds it uncomfortable at first. That discomfort is a fair response to an unusual operation, not a sign that you are slow. The next lesson takes it apart slowly.
Remember this
- A vector is an ordered list of measurements about one thing, and also an arrow with direction and length.
- A matrix is a stack of those lists: a table. It can also act as a machine that turns one list into another.
- Every model turns pictures, words and sounds into vectors first, then works only with vectors and matrices.
What to learn next
- Linear algebra — what these objects can do together.
- Embeddings — vectors that carry the meaning of words.
- NumPy — how to hold vectors and matrices in Python.
Developer — Code and libraries.
Setup
pip install numpyEverything below is a tiny inline array. Nothing is downloaded, and it runs on any CPU in under a second.
Vectors, length, and similarity
import numpy as np
# Three made-up word vectors. Each row scores one word on four traits.
# traits: sweet fruit metal screen
vecs = np.array([
[0.9, 0.8, 0.1, 0.0], # mango
[0.8, 0.9, 0.0, 0.0], # banana
[0.0, 0.0, 0.9, 1.0], # laptop
])
words = ["mango", "banana", "laptop"]
print("shape:", vecs.shape)
print("length of the mango vector:", round(float(np.linalg.norm(vecs[0])), 3))
def cosine(a, b):
# 1.0 = pointing the same way, 0.0 = nothing in common.
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
print("mango vs banana:", round(cosine(vecs[0], vecs[1]), 3))
print("mango vs laptop:", round(cosine(vecs[0], vecs[2]), 3))
# All pairs at once, with no Python loop.
lengths = np.linalg.norm(vecs, axis=1, keepdims=True)
unit = vecs / lengths
print("all pairs:\n", np.round(unit @ unit.T, 3))shape: (3, 4) length of the mango vector: 1.208 mango vs banana: 0.99 mango vs laptop: 0.055 all pairs: [[1. 0.99 0.055] [0.99 1. 0. ] [0.055 0. 1. ]]
Real word vectors come from a trained model and have hundreds of dimensions instead of four. The mechanism is exactly the one above: normalise, then take a dot product. Every vector database on earth is a fast version of that last line.
Read the final grid as a table of every word against every word. The diagonal is 1.0 because a vector is identical to itself. The grid is symmetric, because similarity does not care which word you name first.
Matrices as layers, and as instructions
import numpy as np
# One neural-network layer is a matrix multiply plus a bias vector.
X = np.array([[1.0, 2.0, 3.0], # four examples, three features each
[0.0, 1.0, 0.0],
[2.0, 0.0, 1.0],
[1.0, 1.0, 1.0]])
W = np.array([[ 0.5, -1.0], # three features in, two outputs out
[ 1.5, 0.0],
[-0.5, 2.0]])
b = np.array([0.1, -0.1])
out = X @ W + b
print("X:", X.shape, " W:", W.shape, " out:", out.shape)
print(out)
# A matrix can also be an instruction: this one turns points a quarter turn.
quarter_turn = np.array([[0.0, -1.0],
[1.0, 0.0]])
point = np.array([3.0, 0.0])
print("point after one quarter turn :", quarter_turn @ point)
print("point after two quarter turns:", quarter_turn @ (quarter_turn @ point))
try:
W @ X
except ValueError as err:
print("shape error:", err)X: (4, 3) W: (3, 2) out: (4, 2) [[ 2.1 4.9] [ 1.6 -0.1] [ 0.6 -0.1] [ 1.6 0.9]] point after one quarter turn : [0. 3.] point after two quarter turns: [-3. 0.] shape error: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 4 is different from 2)
The rotation output is worth pausing on. The point starts three steps to the right. One quarter turn puts it three steps up. Another quarter turn puts it three steps to the left. The matrix stored none of that. It stored a rule, and the rule applied to whatever came in.
Line-by-line walkthrough
vecs.shape is (3, 4) — three rows, four columns. The convention across all of machine learning is rows are examples, columns are features. Once you internalise that, most shape errors solve themselves.
np.linalg.norm(v) is the length of the arrow, the straight-line distance from the origin to the tip. It is not the sum of the entries.
a @ b for two 1-D arrays is the dot product: multiply matching entries, then add them all up. It produces a single number, and that number is large when the two vectors point the same way.
Cosine similarity divides the dot product by both lengths, which removes magnitude and keeps only direction. That matters because a long document and a short one about the same topic should count as similar. Raw dot products would favour the long one.
keepdims=True keeps the result as shape (3, 1) instead of collapsing it to (3,). That shape is what lets vecs / lengths broadcast one length onto each row. Dropping it gives a wrong answer with no error message, which is the most dangerous kind.
X @ W requires the inner dimensions to match: (4, 3) @ (3, 2) works and yields (4, 2). The inner 3s meet and vanish; the outer numbers survive. Reversing the operands breaks, and the error message above shows NumPy checking exactly that.
out = X @ W + b processes four examples in one call. This is a batch, meaning many examples handled together, and it is why GPUs are worth using. The same work done in a Python loop over four separate examples gives identical numbers and wastes the hardware.
Common mistakes
1. Using * when you meant @.
import numpy as np
A = np.array([[1.0, 2.0], [3.0, 4.0]])
B = np.array([[1.0, 0.0], [0.0, 1.0]])
print("A * B ->\n", A * B)
print("A @ B ->\n", A @ B)A * B -> [[1. 0.] [0. 4.]] A @ B -> [[1. 2.] [3. 4.]]
B here is the identity matrix, so A @ B gives A back unchanged, which is the correct answer. A * B multiplies matching cells and quietly destroys half the values. The first form often runs without complaint because the shapes happen to line up, and it produces silently wrong numbers.
2. Cosine similarity without normalising. A raw dot product grows with vector length, so a longer vector looks more similar to everything. Divide by both norms, or normalise once up front as the example does.
3. Confusing (n,) with (n, 1).
A 1-D array of shape (3,) is neither a row nor a column. Combined with a (3, 3) matrix, broadcasting treats it as a row. That is silently wrong when you meant a column. Be explicit with v.reshape(-1, 1) or v[:, None].
4. Transposing as a reflex.
When a shape error appears, adding .T until it goes away will eventually make the code run. It will also compute something else. Print X.shape and W.shape, decide which dimension is meant to cancel, then act.
Try it yourself
Add a fourth word, "tablet", with the traits [0.0, 0.0, 0.7, 1.0], and rerun the all-pairs grid. Confirm that tablet lands closest to laptop and near zero against mango.
Then build a matrix that doubles the first coordinate and leaves the second alone. Apply it to [3.0, 0.0] and to [0.0, 3.0]. Work out on paper what it should give before running it. Getting the prediction right is a better signal of understanding than getting the code right.
What to learn next
- Linear algebra — rank, inverses, and decompositions.
- Embeddings — where real word vectors come from.
- Vector databases — similarity search at a scale of billions.
Researcher — Mathematics and papers.
Definitions
A vector space over the reals is a set closed under addition and scalar multiplication, satisfying the usual eight axioms. For our purposes it is R**n, and a vector is an ordered n-tuple of real numbers.
The standard inner product and the Euclidean norm it induces:
dot(u, v) = sum_{i=1..n} u[i] * v[i]
norm(v) = sqrt( dot(v, v) )
cos(u, v) = dot(u, v) / ( norm(u) * norm(v) )nis the dimension, the number of components.u[i]is thei-th component ofu.cos(u, v)lies in[-1, 1]and equals the cosine of the angle between the two vectors.
The p-norms generalise this:
norm_p(v) = ( sum_i abs(v[i])**p ) ** (1/p)p = 2 gives Euclidean length and is the norm behind weight decay. p = 1 gives the sum of absolute values and induces sparsity. This is why the lasso zeroes coefficients while ridge only shrinks them. p tending to infinity gives the maximum absolute component, which is the norm used in adversarial robustness work.
Matrix multiplication. Take A of shape (m, k) and B of shape (k, n). The product C = A @ B has shape (m, n), with
C[i][j] = sum_{p=1..k} A[i][p] * B[p][j]Read it as follows. Entry (i, j) of the output is the dot product of row i of A with column j of B.
The operation is associative and distributive, but not commutative. A @ B and B @ A usually differ, and often one of them is not even defined.
The more useful reading for machine learning: A @ B applies the linear map A to every column of B. A neural network layer is a linear map followed by a non-linearity. The whole forward pass is a composition of such maps.
Cost
Naive multiplication of two n x n matrices performs 2*n**3 - n**2 floating-point operations, so O(n**3). Strassen's 1969 algorithm reaches O(n**2.807) by trading multiplications for additions. It is occasionally used at large sizes, despite weaker numerical stability.
A long line of work has pushed the theoretical exponent omega below 2.372. The most recent is Alman, Duan, Vassilevska Williams, Xu, Xu and Zhou, 2024.
These are galactic algorithms. The hidden constants make them slower than the naive method at any size that fits in a data centre. Production BLAS kernels remain cubic and win on cache blocking, SIMD width, and tensor-core utilisation instead.
The practical performance model is arithmetic intensity. A matmul does O(n**3) work touching O(n**2) data, so intensity grows with n. Small matmuls are memory-bound; large ones are compute-bound. This is why frameworks fuse many small matmuls into one large batched call.
Rank and low-rank structure
The rank of a matrix is the number of linearly independent rows. Equivalently, the number of independent columns. Equivalently, the number of non-zero singular values.
A matrix of shape (m, n) with rank r carries roughly r * (m + n) independent numbers, not m * n.
The singular value decomposition factors any real matrix as
A = U @ S @ V.TUis(m, m)with orthonormal columns, the left singular vectors.Sis(m, n)diagonal with non-negative entriess[1] >= s[2] >= ... >= 0, the singular values.Vis(n, n)with orthonormal columns, the right singular vectors.V.Tis the transpose ofV.
Eckart and Young, 1936, extended to the spectral norm by Mirsky in 1960, proved a strong result. Truncate the SVD to the largest r singular values. What you get is the best possible rank-r approximation to A, under both the Frobenius and spectral norms.
This one theorem underwrites a great deal of practice:
- PCA is the SVD of a mean-centred data matrix. The top components are the directions of greatest variance.
- LoRA (Hu et al., 2021) freezes a pretrained weight matrix
Wand learns an updateB @ A. HereBis(d, r)andAis(r, d), withrfar smaller thand. Trainable parameters drop by orders of magnitude. The justification is empirical: fine-tuning updates turn out to be approximately low rank. - Embedding compression and model distillation both exploit the same redundancy.
Attention (Vaswani et al., 2017) is a stack of matrix products:
Attention(Q, K, V) = softmax( Q @ K.T / sqrt(d_k) ) @ VQ,K,Vare the query, key and value matrices, each(seq_len, d_k)for a single head.Q @ K.Tis(seq_len, seq_len): one similarity score per pair of positions. It is the same dot-product similarity as the cosine formula above, before normalisation.- Dividing by
sqrt(d_k)keeps the pre-softmax scores at a scale where the softmax gradient does not vanish.
That (seq_len, seq_len) intermediate is the quadratic memory cost of attention. FlashAttention (Dao et al., 2022) removes it by tiling the computation and never materialising the full matrix.
Numerical caveats
The condition number is kappa(A) = s_max / s_min, the largest singular value divided by the smallest. It bounds how much a relative input error can be amplified in the output. Solving a system with kappa near 1e8 in float32 leaves essentially no correct digits. Prefer np.linalg.solve or lstsq over forming np.linalg.inv(A) @ b, which is both slower and less stable.
Floating-point addition is not associative, so matrix multiplication is not either. Change the reduction order and you change the last bits of the result. A different thread count does it. So does a different tile size, or a different GPU. Over thousands of training steps, those bits diverge into visibly different models. This is the honest reason bit-exact reproducibility across hardware is hard, and why torch.use_deterministic_algorithms(True) costs speed.
Precision in practice: weights and activations are commonly bfloat16. That format keeps the float32 exponent range while giving up significand bits. So overflow is rare and rounding error is large.
Accumulation inside a matmul is nonetheless done in float32. Summing thousands of bfloat16 products in bfloat16 would lose the small terms entirely.
References
- Strang, G. Introduction to Linear Algebra, 6th ed., Wellesley-Cambridge Press, 2023. Also the MIT 18.06 lectures.
- Trefethen, L. N., Bau, D. Numerical Linear Algebra. SIAM, 1997. The book to read for conditioning and stability.
- Golub, G. H., Van Loan, C. F. Matrix Computations, 4th ed., Johns Hopkins University Press, 2013.
- Eckart, C., Young, G. "The approximation of one matrix by another of lower rank." Psychometrika 1(3), 211–218, 1936.
- Strassen, V. "Gaussian elimination is not optimal." Numerische Mathematik 13, 354–356, 1969.
- Alman, J., Duan, R., Vassilevska Williams, V., Xu, Y., Xu, Z., Zhou, R. "More Asymmetry Yields Faster Matrix Multiplication." arXiv:2404.16349, 2024.
- Vaswani, A., et al. "Attention Is All You Need." NeurIPS, 2017.
- Hu, E. J., Shen, Y., Wallis, P., et al. "LoRA: Low-Rank Adaptation of Large Language Models." arXiv:2106.09685, 2021.
- Dao, T., Fu, D. Y., Ermon, S., Rudra, A., Ré, C. "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness." NeurIPS, 2022.
What to learn next
- Linear algebra — eigenvalues, decompositions, and solving systems.
- Derivatives and gradients — differentiating through these products.
- Attention — the matrix products above, in their working context.