Weather and climate data formats
Weather and climate data is stored as a giant grid of numbers laid over the map, stacked by height and repeated over time — and two file formats, NetCDF and GRIB, hold almost all of it.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Weather and climate data is usually stored as a giant grid of numbers laid over the map, changing over time.
Picture an ice cube tray. Every little compartment holds one value, arranged neatly in rows and columns. Now imagine stacking dozens of these trays on top of each other, one tray for every height in the sky. Imagine a brand new stack arriving every single hour, forever. That stack of trays is close to what a weather dataset actually looks like inside a computer.
Why it exists
Real measurements are messy. Weather balloons, satellites, ships, and ground stations report from wherever they happen to be. They measure at whatever moment they happen to — never neatly lined up. Before any model can use this data, scientists interpolate it onto a regular grid — evenly spaced boxes covering the map. Evenly spaced heights. Evenly spaced times.
Two file formats hold almost all of this gridded data:
- NetCDF (Network Common Data Form) — the standard in climate science. It is self-describing: the file itself states its units, its grid spacing, and what each number means.
- GRIB (GRIdded Binary) — the standard most national weather agencies use for operational forecasts, more compact, built for fast transmission.
How it works
Grid: 0.25° x 0.25° boxes covering the globe (about 28km per box near the equator)
Levels: many heights in the atmosphere -- near the ground, up to the stratosphere
Time: a fresh full grid every 1, 3, or 6 hours
|
v
stored as one array: [time, level, lat, lon] = temperature
(or humidity, wind speed, pressure, rainfall...)A single day of global data, at every height, every few hours, adds up to billions of individual numbers. That is why compact formats and efficient storage matter as much as the numbers themselves.
Where you have already seen it
The temperature and rain forecast on your phone's weather app almost never comes from a person looking at the sky. It comes from a NetCDF or GRIB file, produced by an agency like the India Meteorological Department, the US NOAA, or Europe's ECMWF. The nearest grid point's value gets interpolated to your exact location.
Remember this
- Real-world measurements are messy; weather and climate data gets interpolated onto a clean grid before most models ever touch it.
- NetCDF and GRIB are the two formats that carry almost all of this data, in research and in operational forecasting alike.
- A gridded dataset is, at its core, a large, labelled array — time, height, latitude, and longitude, each with a number attached.
What to learn next
- Machine-learned weather forecasting — the models trained directly on data stored this way.
- What is time series data? — the general pattern this data follows at any single grid point.
- Downscaling climate projections — turning this coarse global grid into something useful for one specific town.
Developer — Code and libraries.
Setup
pip install numpy xarray netcdf4xarray is the standard Python library for working with labelled, multi-dimensional data like this. It reads and writes NetCDF files directly, keeping track of coordinates and units for you.
Minimal runnable code
Building a tiny real NetCDF file, then reading it back the way you would read a genuine weather dataset.
import numpy as np
import xarray as xr
rng = np.random.default_rng(0)
lats = np.array([28.5, 28.75, 29.0])
lons = np.array([76.5, 76.75, 77.0, 77.25])
times = np.array(["2026-06-01T00", "2026-06-01T06", "2026-06-01T12"], dtype="datetime64[ns]")
temperature = 30 + rng.normal(0, 2, size=(len(times), len(lats), len(lons)))
ds = xr.Dataset(
{"temperature": (["time", "lat", "lon"], temperature, {"units": "degC"})},
coords={"time": times, "lat": lats, "lon": lons},
)
ds.to_netcdf("weather_demo.nc") # a real NetCDF file, the format weather agencies actually use
loaded = xr.open_dataset("weather_demo.nc")
# Delhi is roughly at 28.6N, 77.2E -- find the nearest grid cell to it.
point = loaded["temperature"].sel(lat=28.6, lon=77.2, method="nearest")
print(f"file variables: {list(loaded.data_vars)}")
print(f"grid shape (time, lat, lon): {loaded['temperature'].shape}")
print(f"units: {loaded['temperature'].attrs['units']}")
print(f"temperature near Delhi, over time: {point.values.round(1)}")file variables: ['temperature'] grid shape (time, lat, lon): (3, 3, 4) units: degC temperature near Delhi, over time: [30.2 28.5 28.2]
What actually happened
ds.to_netcdf() wrote an actual, standards-compliant NetCDF file to disk — the same kind of file a national weather agency would publish, though far smaller. xr.open_dataset() read it back with every coordinate and unit intact. NetCDF stores that metadata inside the file itself, not in a separate document you have to keep track of.
.sel(lat=28.6, lon=77.2, method="nearest")is the everyday operation in this field: "give me the grid cell closest to this real-world point." Real coordinates almost never land exactly on a grid line.- The
{"units": "degC"}dictionary attached to the temperature variable travels with the data permanently. Losing track of units — Celsius versus Kelvin, especially — is a genuinely common, and genuinely costly, real-world mistake in this field.
Common mistakes
Assuming a grid point is a single physical location. A 0.25° grid box is about 28km wide near the equator. A value at one point represents an average condition across that entire box, not a precise reading at one address.
Loading an entire global, multi-year file into memory at once. Real climate datasets can be hundreds of gigabytes. xarray supports lazy loading and chunked access (chunks= when opening a file) specifically so you never need to.
Losing track of the time zone. Weather and climate files are almost always stored in UTC. Forgetting to convert before comparing against a local weather report is a common source of "off by a few hours" bugs.
Try it yourself
Add a second variable to the dataset — humidity, with its own random values and a "%" units attribute — inside the same xr.Dataset(...) call. Save, reload, and print both variables' units. A real weather file holds many variables side by side, sharing the same grid.
What to learn next
- Machine-learned weather forecasting — models trained directly on data stored exactly this way.
- What is time series data? — the shape this data takes once you follow a single grid point through time.
- Forecasting solar and wind output — a direct downstream use of this kind of gridded weather data.
Researcher — Mathematics and papers.
Data assimilation
Raw observations are not dropped directly onto the nearest grid cell. Data assimilation combines sparse, irregular observations with a short-range physics-based forecast (the "background"), to produce the best consistent gridded state. It typically uses 4D-Var or ensemble Kalman filter methods — a large-scale relative of the Kalman filter covered in sensor fusion and the Kalman filter. ECMWF's ERA5 reanalysis, one of the most widely used research datasets in this field, is the output of exactly this process, run globally at 0.25° resolution, hourly, back to 1940.
File format internals
NetCDF (built on HDF5 since NetCDF-4) stores data in self-describing binary form: dimensions, variables, and attributes are all queryable from the file without external documentation, and supports internal compression and chunking for efficient partial reads. GRIB stores each field as an independent, compactly encoded message, optimised for the operational need to transmit and archive massive volumes of forecast output quickly. It is less convenient for exploratory analysis, but remains the default output format for most national forecasting centres.
Common libraries and conventions
xarray, used above, layers labelled, NumPy-compatible arrays over NetCDF, GRIB (via cfgrib), and Zarr. Zarr in particular has become the preferred format for large-scale cloud-native climate data, since it is designed for parallel, chunked read and write directly against object storage (S3-compatible buckets), unlike NetCDF's traditional single-file model. The CF (Climate and Forecast) metadata conventions standardise variable naming, units, and coordinate reference systems across the field, so files from different agencies remain interoperable.
Scale
A single ERA5 variable, globally, hourly, at 0.25° resolution, for one year, is on the order of tens of gigabytes; the full multi-variable, multi-level archive back to 1940 is measured in petabytes. This scale is the direct motivation for chunked, cloud-native formats and for lazy-loading libraries — no single machine loads the full archive into memory.
Key references
- Hersbach, H. et al. (2020). The ERA5 Global Reanalysis. Quarterly Journal of the Royal Meteorological Society.
- Eaton, B. et al. NetCDF Climate and Forecast (CF) Metadata Conventions.
- WMO. Manual on Codes: GRIB2.
What to learn next
- Machine-learned weather forecasting — the models trained on archives like ERA5.
- Downscaling climate projections — refining this coarse grid to local scale.
- What is time series data? — the statistical structure underlying a single grid point's history.