How images are stored in a computer
An image is a grid of numbers, one small set of numbers per dot, and every single thing in computer vision is arithmetic done on that grid.
- 21 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An image is a grid of numbers. Nothing else is stored.
Think of the giant screen at a cricket stadium, or the LED board outside a wedding hall. From your seat you see a face. Walk right up to it and the face disappears. You are standing in front of thousands of tiny lamps, in neat rows.
Each lamp is not a face. It is one small light, glowing at some brightness. The face only exists when you step back far enough for your eye to blend them.
A computer never gets to step back. It only ever holds the list of lamp brightnesses.
The word for one lamp
One dot of a picture is called a pixel, meaning one square of the grid. The word is short for "picture element".
A pixel does not know what it is part of. It holds a number, and that number says how bright it is. A low number is dark. A high number is bright.
Here is a picture that is five dots wide and five dots tall. This is what the computer actually stores:
0 0 255 0 0
0 0 255 0 0
255 255 255 255 255
0 0 255 0 0
0 0 255 0 0Squint at it. The 255s make a plus sign, because 0 means black and 255 means white.
You have read an image the way a computer does. That is the whole idea, and everything else in this section is built on it.
Why the numbers stop at 255
Every value is stored in one byte — the smallest useful lump of memory a computer works with. A byte can hold two hundred and fifty-six different values, counted from 0.
So brightness runs from 0 to 255. Not because eyes work that way, but because memory does.
This has a real consequence you will hit within your first week. There is nothing above 255. Try to make a bright pixel brighter and the value wraps around to a small number. A white patch turns black. The Developer section below shows this happening.
Colour is three numbers, not one
A black-and-white picture needs one number per dot. Colour needs three.
Those three are red, green and blue. Every colour you have ever seen on a screen is a mix of exactly those three lamps at different brightnesses.
pixel -> [ red 255, green 165, blue 0 ] an orange dot
pixel -> [ red 0, green 0, blue 0 ] a black dot
pixel -> [ red 255, green 255, blue 255] a white dotEach of those three lists is called a channel — one full grid of numbers, one per colour. A colour photo is three grids stacked on top of each other, lined up dot for dot.
How big does that get
A photo from a modern phone is about four thousand dots across and three thousand dots down. That is twelve million dots.
Three numbers per dot means thirty-six million numbers, for one photo. On a good phone camera, that is thirty-six megabytes before any squeezing.
Your gallery is not thirty-six megabytes per photo. The picture gets squeezed down for storage, and unpacked back into the full grid the moment anything looks at it. More on that squeezing below.
Rows first, then columns
This trips up nearly everyone, so read it slowly.
In maths class you were taught to say "across, then up" — the horizontal position first. Images do the opposite. You give the row first, then the column.
image[0][2] means row 0, column 2 -> the top edge, third dot along
image[2][0] means row 2, column 0 -> the left edge, third dot downAnd row 0 is the top row, not the bottom. Pictures are counted downward from the top left corner, the way you read a page.
Get this backwards and your program will run happily, produce a sideways image, and never tell you anything is wrong. Nearly every beginner loses an afternoon to it once.
Squeezing pictures down
You have seen the file endings. They are not interchangeable, and picking the wrong one damages your work.
PNG puts everything back exactly. Every number you saved comes out identical. It suits screenshots, diagrams, logos and anything with sharp edges or text.
JPEG throws information away on purpose. It keeps what your eye notices and discards what it does not. Photographs survive this beautifully. Text and sharp lines come back with faint smudges around them.
That word for throwing information away is lossy. The opposite, keeping everything, is lossless.
Here is the part that matters for your work. Save a photo as JPEG, open it, save it again, and again — the damage stacks up every time. Keep an untouched original of any picture you plan to feed to a model.
Where you have already seen all this
- A photo that looks blocky on a slow connection. You are watching a heavily squeezed version arrive first.
- A screenshot of text that looks fuzzy when shared. Someone saved it as JPEG.
- A photo editor's brightness slider. It is adding the same number to every pixel in the grid.
- A black-and-white filter. It is collapsing three numbers per dot down to one.
What is honestly confusing here
Two things catch out almost everyone, and neither is your fault.
The order of the colour channels is not agreed on. Most tools store red, green, blue. One very popular library, OpenCV, stores blue, green, red instead. Mix them up and your photo comes out looking like an underwater shot. The next lesson deals with this head on.
A picture on your screen is not the same as the picture in the file. Screens brighten dark values before showing them. So the numbers you print and the brightness you see do not line up. This is normal and it is explained properly in the Researcher section.
Remember this
- An image is a grid of numbers, one per dot, and a pixel is one dot.
- Brightness runs from 0 to 255, because one number is stored in one byte.
- Colour is three stacked grids: red, green and blue.
- Rows are counted first, from the top. Columns come second.
What to learn next
- OpenCV — the library that reads, writes and reshapes these grids for you.
- What is computer vision? — the wider picture this grid sits inside.
- Convolutional neural networks — what a model does once it has the grid.
Developer — Code and libraries.
Everything below runs on any laptop in under a second. There is no download, no dataset and no GPU. If images have felt abstract, five minutes here fixes that permanently.
Setup
pip install numpy pillowNumPy holds the grid. Pillow is only needed for the last example, about file formats.
An image you can read with your own eyes
import numpy as np
# A 5x5 grayscale image, typed out by hand. 0 is black, 255 is white.
img = np.array([
[ 0, 0, 255, 0, 0],
[ 0, 0, 255, 0, 0],
[255, 255, 255, 255, 255],
[ 0, 0, 255, 0, 0],
[ 0, 0, 255, 0, 0],
], dtype=np.uint8)
print("what the computer stores:")
print(img)
print()
print("the same numbers, drawn:")
for row in img:
print("".join("#" if value > 127 else "." for value in row))
print()
print("shape (rows, columns):", img.shape)
print("value at row 0, column 2:", img[0, 2])
print("value at row 0, column 0:", img[0, 0])
print("how many numbers in total:", img.size)what the computer stores: [[ 0 0 255 0 0] [ 0 0 255 0 0] [255 255 255 255 255] [ 0 0 255 0 0] [ 0 0 255 0 0]] the same numbers, drawn: ..#.. ..#.. ##### ..#.. ..#.. shape (rows, columns): (5, 5) value at row 0, column 2: 255 value at row 0, column 0: 0 how many numbers in total: 25
The array and the drawing are the same object. The second one is only a nicer costume.
dtype=np.uint8 is the important part. uint8 means an unsigned 8-bit integer: a whole number from 0 to 255, with no negatives allowed. Every image loading library on earth hands you uint8 by default.
Colour: the same grid, three deep
import numpy as np
# Two rows, three columns, three colour channels per pixel.
img = np.zeros((2, 3, 3), dtype=np.uint8)
img[0, 0] = [255, 0, 0] # pure red
img[0, 1] = [ 0, 255, 0] # pure green
img[0, 2] = [ 0, 0, 255] # pure blue
img[1, 0] = [255, 255, 0] # red and green together look yellow
img[1, 1] = [255, 255, 255] # all three full on is white
img[1, 2] = [ 40, 40, 40] # equal and low is dark grey
print("shape (height, width, channels):", img.shape)
print("bytes on disk, uncompressed:", img.nbytes)
print()
print("the pixel at row 0, column 0:", img[0, 0])
print(" its red :", img[0, 0, 0])
print(" its green:", img[0, 0, 1])
print(" its blue :", img[0, 0, 2])
print()
print("the red channel on its own:")
print(img[:, :, 0])
print("the blue channel on its own:")
print(img[:, :, 2])shape (height, width, channels): (2, 3, 3) bytes on disk, uncompressed: 18 the pixel at row 0, column 0: [255 0 0] its red : 255 its green: 0 its blue : 0 the red channel on its own: [[255 0 0] [255 255 40]] the blue channel on its own: [[ 0 0 255] [ 0 255 40]]
Read that shape out loud: height, width, channels. Not width first. NumPy indexes the slowest-changing axis first, and for an image that is the row.
img[:, :, 0] means "every row, every column, channel 0". You get back a plain 2-by-3 grid — the red channel as a standalone grayscale image.
The overflow that eats your afternoon
Brightening an image looks like adding a number to every pixel. Do it carelessly and white turns black.
import numpy as np
dark = np.array([250, 250, 250], dtype=np.uint8)
# Brightening by adding, the wrong way.
print("wrong :", dark + np.uint8(10))
# Brightening with room to work in, then coming back to uint8.
bright = np.clip(dark.astype(np.int16) + 10, 0, 255).astype(np.uint8)
print("right :", bright)wrong : [4 4 4] right : [255 255 255]
Two hundred and fifty plus ten gave four. There is no warning, no exception, no red text. uint8 has 256 slots and counting past the last one wraps back to the first.
The rule, worth a sticky note: widen, calculate, clip, narrow. Cast to a wider type, do the arithmetic, clip to the 0 to 255 range, then cast back. np.clip is what stops values escaping the range before you convert.
Turning colour into grey the right way
A naive grayscale conversion averages the three channels. That looks wrong to people, and there is a reason.
import numpy as np
colours = {
"pure red ": np.array([255, 0, 0]),
"pure green": np.array([ 0, 255, 0]),
"pure blue ": np.array([ 0, 0, 255]),
}
WEIGHTS = np.array([0.299, 0.587, 0.114]) # how much the eye trusts each channel
print("colour plain average eye-weighted")
for name, rgb in colours.items():
plain = rgb.mean()
weighted = (rgb * WEIGHTS).sum()
print(f"{name} {plain:6.1f} {weighted:6.1f}")colour plain average eye-weighted pure red 85.0 76.2 pure green 85.0 149.7 pure blue 85.0 29.1
The plain average calls all three colours equally bright. Your eye disagrees strongly. Pure green looks far brighter than pure blue, and the weighted version reflects that.
Those three weights are not arbitrary. They come from a broadcast television standard, and every mainstream library uses them or a close relative. The Researcher section names the standard and its successor.
PNG against JPEG, measured
import io
import numpy as np
from PIL import Image
# A 64x64 picture of the letter T: white strokes on black, sharp edges.
letter = np.zeros((64, 64), dtype=np.uint8)
letter[10:14, 8:40] = 255 # the top bar
letter[10:40, 20:24] = 255 # the stem
def roundtrip(array, fmt, **kwargs):
"""Save to memory, read back, and report the file size."""
buffer = io.BytesIO()
Image.fromarray(array).save(buffer, format=fmt, **kwargs)
size = buffer.tell()
buffer.seek(0)
return size, np.array(Image.open(buffer))
png_size, png_back = roundtrip(letter, "PNG")
jpg_size, jpg_back = roundtrip(letter, "JPEG", quality=80)
print("raw pixels :", letter.nbytes, "bytes")
print("PNG file :", png_size, "bytes")
print("JPEG file :", jpg_size, "bytes")
print("PNG pixel damage :", int(np.abs(png_back.astype(int) - letter).max()))
print("JPEG pixel damage :", int(np.abs(jpg_back.astype(int) - letter).max()))
print()
print("row 14, columns 16 to 27, before JPEG:", letter[14, 16:28])
print("row 14, columns 16 to 27, after JPEG :", jpg_back[14, 16:28])raw pixels : 4096 bytes PNG file : 117 bytes JPEG file : 485 bytes PNG pixel damage : 0 JPEG pixel damage : 17 row 14, columns 16 to 27, before JPEG: [ 0 0 0 0 255 255 255 255 0 0 0 0] row 14, columns 16 to 27, after JPEG : [ 8 0 0 14 238 255 255 244 0 0 0 0]
The byte counts depend on which Pillow and libjpeg build you have, so yours may differ by a few bytes. The pattern will not.
Look at the last two lines. Pixels that were pure black came back as 8 and 14. Pixels that were pure white came back as 238 and 244. That faint halo around a sharp edge is called ringing, and it is JPEG's signature damage.
Now look at the file sizes. On this image JPEG is four times larger than PNG and it corrupted the pixels. JPEG is built for photographs, where brightness changes gradually. Feed it hard edges and it loses on both counts.
Line by line, the parts that are not obvious
np.zeros((2, 3, 3), dtype=np.uint8) creates a black image. Zero is black in every channel, so a fresh array of zeros is a valid, very boring picture.
letter[10:14, 8:40] = 255 is a slice assignment: rows 10 up to but not including 14, columns 8 up to but not including 40. Assigning a single number to a slice fills the whole rectangle. This is how you draw without any drawing library.
io.BytesIO() is a file that lives in memory. Using it means the format comparison touches no disk and leaves nothing behind, which makes the example safe to run anywhere.
np.abs(back.astype(int) - original) must cast to a signed type first. Subtracting two uint8 arrays wraps exactly like the overflow example above, and you would measure nonsense.
Common mistakes
Assuming the channel order. OpenCV hands you blue, green, red. Pillow, matplotlib, PyTorch and almost everything else use red, green, blue. Convert explicitly rather than hoping. This is the single most common first-week bug in computer vision.
Doing arithmetic in uint8. Averaging two images, adding brightness, subtracting a background — all of these overflow silently. Cast to float, work, clip, cast back.
Forgetting to divide by 255. Pretrained models expect inputs scaled to a small range, usually 0 to 1 and then normalised further. Feed raw 0-to-255 values and accuracy collapses without a single error message.
Confusing shape with size. img.shape gives (height, width, channels). img.size gives the count of individual numbers, which is all three multiplied together. Printing the wrong one sends you hunting for a bug that is not there.
Re-saving JPEGs. Every save loses a little more. If a pipeline reads a JPEG, crops it and writes a JPEG, the damage compounds over runs. Decode once, work in arrays, write PNG for anything intermediate.
Try it yourself
Change the drawing threshold in one_image.py from 127 to 254 and re-run. The plus sign survives, because the strokes are exactly 255.
Now do something more revealing. Build a 16-by-16 image where the value equals the column number multiplied by sixteen, so it fades left to right. Print it with three characters instead of two: a space, a dash and a hash. Watch a smooth gradient appear out of the numbers.
Then save that gradient as JPEG at quality=80 and measure the damage the same way. It will be far smaller than for the letter T. Now you have measured, on your own machine, exactly which kinds of pictures JPEG is built for.
What to learn next
- OpenCV — the library that reads, writes and reshapes these grids for you.
- What is computer vision? — the wider picture this grid sits inside.
- Convolutional neural networks — what a model does once it has the grid.
Researcher — Mathematics and papers.
What a stored image actually is
A digital image is a sampled, quantised measurement of light, indexed by a discrete grid:
I : {0, ..., H-1} x {0, ..., W-1} x {0, ..., C-1} -> {0, ..., 2^b - 1}H,W— height and width in pixels.C— channel count: 1 for grayscale, 3 for RGB, 4 with an alpha channel.b— bit depth per channel, 8 in consumer formats, 10 to 16 in raw capture and medical imaging.
Two irreversible steps sit between the scene and this array. Sampling integrates irradiance over a finite sensor area, which band-limits and aliases. Quantisation maps a continuous intensity to one of 2^b levels, which sets the noise floor.
Everything downstream inherits both. A model that appears to learn "shape" may be learning sensor-specific quantisation and demosaicing artefacts, which is one mechanism behind failure under camera change.
Memory and layout
Uncompressed size in bytes:
bytes = H * W * C * ceil(b / 8)A 4000 x 3000 RGB image at 8 bits is 4000 * 3000 * 3 = 36,000,000 bytes, about 34.3 MiB. Batching 32 of these as float32 costs 32 * 36e6 * 4 = 4.6 GB, which is why input pipelines resize before batching, not after.
Layout matters as much as size. Two conventions dominate:
| Layout | Order | Used by |
|---|---|---|
| NHWC | batch, height, width, channel | TensorFlow, most image I/O, JPEG decoders |
| NCHW | batch, channel, height, width | PyTorch, cuDNN, TensorRT |
torchvision.transforms.ToTensor performs the NHWC to NCHW permutation and the divide by 255 in one step. A permutation is a stride change, not a copy, so tensor.is_contiguous() can be False afterwards. Some kernels then trigger a silent copy. Channels-last memory format on modern hardware is frequently faster for convolutions despite PyTorch's NCHW logical layout — tensor.to(memory_format=torch.channels_last) exposes it.
Colour encoding, and the mistake almost everyone makes
Stored 8-bit sRGB values are not proportional to light intensity. The sRGB standard (IEC 61966-2-1:1999) applies a transfer function close to a power law with exponent 1/2.2:
encode: V' = 1.055 * V^(1/2.4) - 0.055 for V > 0.0031308
V' = 12.92 * V otherwise
decode: V = ((V' + 0.055) / 1.055)^2.4 for V' > 0.04045
V = V' / 12.92 otherwiseV— linear light intensity, normalised to the 0 to 1 range.V'— the gamma-encoded value actually stored, also 0 to 1 before scaling to 0 to 255.
The consequence is concrete. Averaging two sRGB values does not average their light. Resizing, blurring and alpha compositing are all linear operations being applied to non-linear data, which is why naive downsampling of high-contrast content darkens it. Correct practice is decode to linear, operate, re-encode. Most deep learning pipelines skip this entirely and are internally consistent enough to get away with it, which is worth knowing but not worth imitating in colour-critical work.
Luma weights
The familiar grayscale weights are luma, denoted Y', computed on gamma-encoded values:
Rec. 601 (SD): Y' = 0.299 R' + 0.587 G' + 0.114 B'
Rec. 709 (HD): Y' = 0.2126 R' + 0.7152 G' + 0.0722 B'OpenCV's COLOR_RGB2GRAY and Pillow's "L" conversion both use Rec. 601. Video pipelines and modern displays assume Rec. 709. The difference is small on natural images and large on saturated colours, so a red-on-green test chart will not survive the swap.
True luminance, Y, is the Rec. 709 weighted sum of linearised channels. The prime marks in Y', R', G', B' denote gamma encoding, and dropping them is the notational sin that hides this whole distinction.
Chroma subsampling
Human vision resolves luminance far better than colour. JPEG and every video codec exploit this by converting RGB to Y'CbCr and storing the two chroma planes at reduced resolution.
| Notation | Chroma resolution | Bytes per pixel at 8 bits |
|---|---|---|
| 4:4:4 | full | 3 |
| 4:2:2 | half horizontally | 2 |
| 4:2:0 | half in both directions | 1.5 |
4:2:0 is the default for JPEG at typical quality settings and for essentially all consumer video. It halves the data. So any task depending on precise colour boundaries is working on interpolated data before your code ever sees it. Chroma keying, fine segmentation of thin coloured structures, and colour-based defect detection are all in that category. Force 4:4:4 at capture if colour edges carry your signal.
JPEG, in one paragraph of mechanism
JPEG (ITU-T T.81, 1992) converts to Y'CbCr and subsamples chroma. Each plane is split into 8 x 8 blocks and a 2-D discrete cosine transform is applied per block. The coefficients are divided by a quantisation table, rounded to integers, and entropy-coded. The quality setting scales the quantisation table.
Two properties follow directly and both matter for vision work. Rounding the quantised coefficients is the only lossy step. It is where ringing near sharp edges comes from: a step edge needs high-frequency coefficients, and those are quantised hardest. And because blocks are coded independently, errors do not cross the 8 x 8 boundary, which produces the familiar blocking at low quality.
Re-encoding is not idempotent in general. It is approximately idempotent if the image is re-saved at the same quality with identical sampling and block alignment, which is the basis of JPEG forensics: misaligned re-compression leaves detectable periodic artefacts.
Preprocessing contracts
A pretrained checkpoint carries an implicit contract about its input. Breaking any clause degrades accuracy silently.
x = (I / 255 - mean) / stdI— theuint8array as decoded.mean,std— per-channel constants baked in at pre-training.
The ImageNet convention is mean = [0.485, 0.456, 0.406] and std = [0.229, 0.224, 0.225], in RGB order. CLIP uses different constants. Inception-family models use [0.5, 0.5, 0.5] for both. The contract also fixes the resize interpolation, the crop rule and the channel order.
torchvision's weight enums expose the exact transform via weights.transforms(). Use it rather than reconstructing the pipeline from memory — a bilinear-versus-bicubic mismatch alone costs measurable accuracy, as documented in the torchvision "How to use pre-trained models" notes.
Bit depth and high dynamic range
8 bits per channel gives 256 levels, which is enough for a well-exposed display-referred image and not enough for anything else. Two failure modes recur:
Banding. Smooth gradients quantised to 256 levels show visible steps, especially after gamma adjustment stretches part of the range. Dithering hides it; more bits solve it.
Clipping. Scene dynamic range routinely exceeds 1000:1. An 8-bit display-referred encoding discards highlight and shadow detail at capture, and no post-processing recovers it. Raw sensor formats retain 12 to 14 bits precisely to defer that decision.
For scientific and medical imaging, 16-bit is standard and the 0-to-255 assumption is wrong. DICOM pixel data commonly carries 12 significant bits with a window/level transform applied only for display; feeding windowed 8-bit renders to a model discards information that the raw stored values still hold.
References
- IEC 61966-2-1:1999. Multimedia systems and equipment — Colour measurement and management — Part 2-1: Default RGB colour space — sRGB.
- ITU-R BT.601-7 (2011) and ITU-R BT.709-6 (2015) — the luma coefficient definitions.
- ITU-T Recommendation T.81 (1992). Digital compression and coding of continuous-tone still images. The JPEG specification.
- Wallace, G. K. (1992). The JPEG Still Picture Compression Standard. IEEE Trans. Consumer Electronics 38(1).
- Poynton, C. (2012). Digital Video and HD: Algorithms and Interfaces, 2nd ed. The standard reference for gamma, luma and subsampling, and the source of the
Y'versusYdiscipline. - Reinhard, E. et al. (2010). High Dynamic Range Imaging, 2nd ed.
- Ehrlich, M. & Davis, L. (2019). Deep Residual Learning in the JPEG Transform Domain. arXiv:1812.11690 — networks operating on DCT coefficients directly, skipping decode.
Open issues worth knowing
Compression as an unmodelled covariate. Training data scraped from the web arrives at mixed JPEG qualities. Test data from a controlled capture rig does not. Models pick up compression statistics as features, and this contributes to the accuracy drop under distribution shift measured by Recht et al. (2019) and by ImageNet-C's JPEG corruption axis.
Colour management is absent from most pipelines. ICC profiles attached to source images are routinely dropped at decode. An image tagged Adobe RGB, decoded as sRGB, has systematically shifted colours, and nothing in the pipeline reports it.
Decode is a real cost. At high throughput, JPEG decode dominates the input pipeline and starves the accelerator. Hardware decode (NVIDIA nvJPEG, DALI) or pre-decoding to a raw shard format is standard practice at scale. It is the first thing to check when utilisation is low and nobody can find the bottleneck.
What to learn next
- OpenCV — the library that reads, writes and reshapes these grids for you.
- What is computer vision? — the wider picture this grid sits inside.
- Convolutional neural networks — what a model does once it has the grid.