Estimating how long and how much a run will cost
Before you start a training run, you can work out roughly how long it will take and what it will cost, the same way a contractor quotes a job before picking up a tool.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
You can estimate the time and money a training run will need before you start it. Measure a small piece of the work first.
The analogy you have already lived
You have asked a painter to quote a job before they touch a wall. They do not paint the whole house to find out the price.
They paint one square metre, time themselves, and look at the paint used. Then they multiply that by the size of your house. A three-day job does not become a surprise on day five, because someone measured before starting.
Training a model works the same way. You do not need to run the whole job to know roughly what it will cost. You need to run a small, honest slice of it first.
Why it exists
Training a large model can take days and run on rented, expensive machines. Some teams start a run and walk away. They come back to a bill they did not expect, or a job still running two days late.
Both problems have the same fix: measure before you commit. A rough number known in advance beats an exact number discovered too late.
The two things you are estimating
Time. How many hours or days until this run finishes.
Money. What that time will cost, on whatever machine you rented to do it.
They are connected but not identical. A faster machine can finish sooner and still cost more per hour. A cheaper machine can take so much longer that the total bill is worse.
How it works
run a SMALL slice of the real job
(a few minutes, on the real hardware)
|
v
see how much work got done in that time
|
v
scale that up to the FULL job you actually want
|
v
"this will take about 14 hours, and cost about $35"The whole idea rests on one assumption: the small slice must look like the real job. Same kind of data, same size of model, same machine. A slice that cuts corners gives you a confident, wrong answer.
A real example you have seen
Cloud photo-editing apps that show "estimated processing time: 4 minutes" before you press start. They are not guessing. They timed a piece of the job on your file and scaled it up.
Video export tools do the same thing. That "time remaining" number updates as it goes. The estimate gets more accurate once more real work has happened.
The honest part
Every estimate here is a guess with a margin of error, not a promise. Real training runs slow down for reasons a two-minute test cannot see.
A machine other jobs are also using. Data that has to travel from a slow disk. A temperature limit that quietly throttles the hardware after an hour.
Treat your estimate as "probably in this range." Keep checking the real run, rather than trusting the first number forever.
Remember this
- Time a small, honest slice of the real job before committing to the whole thing.
- You are estimating two separate numbers: how long, and how much it costs.
- Treat the result as a range, and keep checking the real run against it.
What to learn next
- Caching model predictions — the cheapest run is the one you skip entirely.
- Unit economics of an AI feature — turning a cost estimate into a per-user number that matters to a business.
- Model serving — what happens to a trained model once its cost has been justified.
Developer — Code and libraries.
Setup
pip install torchMeasure a small slice, then scale it up
import time
import torch
import torch.nn as nn
torch.manual_seed(0)
# Stand-in for your real model -- the technique is what matters here, not
# this exact architecture.
model = nn.Sequential(nn.Linear(768, 2048), nn.ReLU(), nn.Linear(2048, 768))
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)
batch = torch.randn(32, 768)
target = torch.randn(32, 768)
loss_fn = nn.MSELoss()
def one_step():
optimizer.zero_grad()
loss = loss_fn(model(batch), target)
loss.backward()
optimizer.step()
# Warm-up steps are thrown away. The first few calls pay for memory
# allocation and kernel selection, which happens once, not every step.
for _ in range(5):
one_step()
MEASURE_STEPS = 50
start = time.perf_counter()
for _ in range(MEASURE_STEPS):
one_step()
elapsed = time.perf_counter() - start
seconds_per_step = elapsed / MEASURE_STEPS
print(f"measured: {elapsed:.3f}s for {MEASURE_STEPS} steps")
print(f"seconds per step: {seconds_per_step:.5f}")
# ---- Everything below is arithmetic on the measurement above, not a new
# ---- measurement. It assumes the real run behaves like this small slice.
TOTAL_STEPS = 100_000
estimated_hours = (seconds_per_step * TOTAL_STEPS) / 3600
print(f"estimated time for {TOTAL_STEPS:,} steps: {estimated_hours:.2f} hours")
ILLUSTRATIVE_PRICE_PER_HOUR = 2.50 # an example on-demand rate, not a live quote
estimated_cost = estimated_hours * ILLUSTRATIVE_PRICE_PER_HOUR
print(f"estimated cost at ${ILLUSTRATIVE_PRICE_PER_HOUR}/hour: ${estimated_cost:.2f}")measured: 0.349s for 50 steps seconds per step: 0.00697 estimated time for 100,000 steps: 0.19 hours estimated cost at $2.5/hour: $0.48
Every number above is a real measurement from this exact machine, on one run. Your seconds per step will differ, sometimes by a lot. It depends on your CPU or GPU, what else is running, even how warm the machine already is. Re-run it on your own hardware rather than trusting this output.
Line-by-line walkthrough
The warm-up loop. The first call to a new layer size does extra work behind the scenes. Timing it would make your estimate too pessimistic, so it is thrown away deliberately.
MEASURE_STEPS = 50. Enough steps that one unlucky slow step does not dominate the average, few enough that the measurement itself stays cheap.
The price constant. It is written in capitals and commented as illustrative on purpose. Real prices change by provider, region and instance type, sometimes by the hour. Never let a hard-coded number like this go stale in a script someone else will trust.
Common mistakes
Measuring at a small batch size, then assuming the full batch size behaves the same. Larger batches often run more than proportionally faster on a GPU, up to a point. Then they hit a memory wall the small batch never showed you. Measure at the batch size you will actually train with.
Forgetting data loading. This script generates its batch in memory. A real job reads files from disk or a network instead. That can dominate the time if your storage is slow. Time a real batch of your real data, not a fake one.
Extrapolating from too few steps. Five steps can be timed by luck — a system hiccup, or a lucky cache hit. Fifty to a few hundred gives a steadier average.
Ignoring the parts that are not training. Loading the dataset, saving checkpoints, and running evaluation all take real time. They are easy to leave out of the estimate entirely.
Try it yourself
Change the batch size from 32 to 128 and re-run. Compare how much seconds per step changed against how much the batch size changed. They rarely move together in a simple, predictable way.
What to learn next
- Caching model predictions — the cheapest run is the one you skip entirely.
- Unit economics of an AI feature — turning a cost estimate into a per-user number that matters to a business.
- Model serving — what happens to a trained model once its cost has been justified.
Researcher — Mathematics and papers.
The compute-first estimate
For transformer training, the total floating-point operations needed is well approximated by
$$C \approx 6ND$$
where $C$ is total training compute in FLOPs, $N$ is the number of trainable parameters, and $D$ is the number of training tokens seen. The factor of 6 comes from roughly 2 FLOPs per parameter per token for the forward pass and roughly 4 for the backward pass (Kaplan et al., 2020).
Time and cost follow directly:
$$T = \frac{C}{r \cdot U}, \qquad \text{Cost} = T \cdot p$$
where $r$ is the accelerator's peak FLOPs per second (from its datasheet), $U$ is model FLOPs utilisation (MFU) — the fraction of that peak you actually achieve, typically 0.3–0.6 for a well-tuned transformer training job — and $p$ is the price per unit time of the hardware.
MFU is the number that separates a confident estimate from a fantasy. Peak hardware FLOPs are a ceiling nobody reaches in practice; memory bandwidth, communication between devices and Python-side overhead all eat into it. The PaLM paper (Chowdhery et al., 2022) reports MFU explicitly for this reason, and treats it as an engineering metric worth optimising in its own right.
Choosing how much to train
Kaplan et al. (2020) showed loss follows a power law in $N$, $D$ and $C$, which implied — under their fitted exponents — that scaling parameters mattered more than scaling data. Hoffmann et al. (2022), the Chinchilla paper, refit those laws with a wider sweep. They reached the opposite practical conclusion: most large models of that era were significantly undertrained relative to their size. For a fixed compute budget $C$, loss is minimised near
$$N_{\text{opt}} \propto C^{0.5}, \qquad D_{\text{opt}} \propto C^{0.5}$$
i.e. parameters and tokens should scale together, roughly one-to-one, not parameters alone. This single result reshaped how compute budgets get allocated between model size and dataset size across the field.
Where the estimate breaks down
- Communication overhead across multiple accelerators is not captured by $C = 6ND$ at all — it depends on topology, interconnect bandwidth, and parallelism strategy, and can dominate at scale.
- Data pipeline stalls show up as low $U$ with no change to $C$, and are invisible to a compute-only estimate.
- Checkpointing and evaluation add wall-clock time without adding to $C$.
A compute-first estimate bounds the minimum achievable time. Real time is always that minimum divided by an achieved utilisation you have to measure, not assume.
Papers
- Kaplan et al., Scaling Laws for Neural Language Models, 2020 — arxiv.org/abs/2001.08361
- Hoffmann et al., Training Compute-Optimal Large Language Models (Chinchilla), 2022 — arxiv.org/abs/2203.15556
- Chowdhery et al., PaLM: Scaling Language Modeling with Pathways, 2022 — arxiv.org/abs/2204.02311
What to learn next
- Caching model predictions — the cheapest run is the one you skip entirely.
- Unit economics of an AI feature — turning a cost estimate into a per-user number that matters to a business.
- Model serving — what happens to a trained model once its cost has been justified.