Self-hosting vs API: the break-even point
Paying per API call is cheap at low volume and expensive at high volume. Owning the hardware is the opposite. There is a real crossover point between them.
- 9 min read
- 3 reading levels
- Published
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Paying an API per use is cheap at low volume and expensive at high volume. Owning the hardware yourself is the reverse. There is a real crossover point between them.
The analogy you have already lived
If you do laundry once a month, you go to a laundromat and pay per load. Buying a washing machine for that would be a waste of money. Most months it would sit unused, after a large amount paid upfront for very little use.
If you do laundry every single day, the laundromat's per-load price adds up fast. Buying a machine starts to look cheap by comparison, spread across all that use. Somewhere between "once a month" and "every day" is a point where the better choice flips.
Running a model through a paid API versus running it yourself works the same way.
Why it exists
An API charges per use — a small amount, but repeated every single time. Self-hosting flips the shape entirely. It is a large, mostly fixed cost: renting or buying hardware, keeping it running. That cost barely changes whether you use it once, or a million times a month.
Neither option is universally cheaper. Which one wins depends entirely on how much you actually use it.
The two cost shapes
API cost grows with usage. Ten times the calls means roughly ten times the bill. At low volume this is very cheap — you pay for exactly what you use, and nothing else.
Self-hosted cost is mostly fixed. A rented machine costs about the same whether it answers ten questions or ten million this month. At low volume this is a bad deal — a lot of money for very little use. At high volume it can be far cheaper. The fixed cost gets spread across so much use that the cost per question becomes tiny.
How it works
cost
^
| API cost
| ,-'
| ,-'
| ,-'
| self-hosted cost ---------------------------- (mostly flat)
| ,-------------------X
| / ^
|/ |
+-----------------------------------------------> usage
BREAK-EVEN POINTBelow the crossing point, the API is cheaper. Above it, self-hosting is cheaper. The whole decision comes down to honestly estimating where your usage sits relative to that point.
A real example you have seen
Small apps and side projects almost always call a paid AI API. Their usage sits far below any sensible break-even point, so owning hardware would be pure waste. Large companies serving the same kind of feature to millions of users often run their own models instead. At their volume, the fixed cost of owning hardware spreads thin enough to easily beat the per-call price.
The honest part
Self-hosting is not only a cost decision. It also means someone has to keep that machine running, patched, and monitored. That is real ongoing work an API handles for you invisibly, as part of its price. A break-even calculation that only counts hardware cost and ignores the people-time to run it is not the whole picture.
Remember this
- API cost scales with usage. Self-hosted cost is mostly fixed.
- There is a real crossover volume where the cheaper option flips.
- Self-hosting's fixed cost is not only hardware — it includes the ongoing work of running it.
What to learn next
- The one-box production stack — what a minimal self-hosted setup can look like once you cross the break-even point.
- Choosing a GPU for inference — sizing the hardware side of a self-hosting decision.
- Docker for ML — packaging a self-hosted model so it runs reliably wherever you put it.
Developer — Code and libraries.
Setup
No installs needed — this uses only Python's standard library.
Finding the crossover point
def api_cost(calls_per_month, price_per_1k_calls):
return (calls_per_month / 1000) * price_per_1k_calls
def self_hosted_cost(monthly_gpu_rental, fixed_engineering_overhead):
return monthly_gpu_rental + fixed_engineering_overhead
# Illustrative example figures -- check real quotes before using numbers like
# these for an actual decision.
API_PRICE_PER_1K_CALLS = 3.00
MONTHLY_GPU_RENTAL = 1800.0 # one always-on GPU instance, on-demand rate
FIXED_ENGINEERING_OVERHEAD = 400.0 # ops time to keep a self-hosted service healthy
hosted_total = self_hosted_cost(MONTHLY_GPU_RENTAL, FIXED_ENGINEERING_OVERHEAD)
# Find the crossover call volume where API cost first exceeds self-hosting cost.
volume, step = 0, 10_000
while api_cost(volume, API_PRICE_PER_1K_CALLS) < hosted_total:
volume += step
print(f"self-hosted fixed cost: ${hosted_total:,.2f}/month")
print(f"API price: ${API_PRICE_PER_1K_CALLS:.2f} per 1,000 calls\n")
for calls in [50_000, 200_000, volume, 1_000_000]:
a = api_cost(calls, API_PRICE_PER_1K_CALLS)
cheaper = "API" if a < hosted_total else "self-host"
print(f"{calls:>9,} calls/month: API ${a:>9,.2f} self-host ${hosted_total:>9,.2f} cheaper: {cheaper}")
print(f"\nbreak-even is around {volume:,} calls/month")self-hosted fixed cost: $2,200.00/month API price: $3.00 per 1,000 calls 50,000 calls/month: API $ 150.00 self-host $ 2,200.00 cheaper: API 200,000 calls/month: API $ 600.00 self-host $ 2,200.00 cheaper: API 740,000 calls/month: API $ 2,220.00 self-host $ 2,200.00 cheaper: self-host 1,000,000 calls/month: API $ 3,000.00 self-host $ 2,200.00 cheaper: self-host break-even is around 740,000 calls/month
Every number here is deterministic arithmetic on the constants at the top. Real and exact, but entirely dependent on those constants being accurate for your situation. The crossover point — about 740,000 calls a month, in this made-up example — is the number to focus on. Not the specific dollar figures.
Line-by-line walkthrough
The while loop searches for the crossover. This is a simple, honest way to find where two lines cross, without needing to solve an equation by hand. Start at zero, and step up until the API line overtakes the flat self-hosted line.
step = 10_000. Controls how precisely the crossover point is found. A coarser step finds it faster but less precisely. A finer step costs more loop iterations, for a more exact answer. For a real decision, precision to the nearest ten thousand calls is more than enough.
FIXED_ENGINEERING_OVERHEAD. Easy to leave out, and doing so makes self-hosting look artificially cheap. The honest part of the beginner section above is encoded directly in this line. It is a real, ongoing cost, not a rounding error.
Common mistakes
Using list price instead of your actual price. Reserved instances, committed-use discounts, and enterprise API contracts can change either side of this calculation substantially. Use the price you will actually pay, not a public sticker price.
Ignoring that "usage" is not constant. Traffic that spikes on some days and drops on others makes a fixed self-hosted machine sit idle sometimes, and overloaded other times. An API absorbs that variance automatically; self-hosting needs its own answer for it — see autoscaling for one.
Comparing today's volume to today's break-even, and forgetting growth. If usage is growing quickly, the right question is not "which is cheaper right now." It is "which is cheaper for where we will be in six months." A decision made on this month's numbers alone can be stale by the time it is implemented.
Forgetting that self-hosting has a floor of capability. An API gives you access to models you could not necessarily run yourself. Not at any cost, depending on their size and your hardware. This break-even calculation only makes sense between two options that can both actually do the job.
Try it yourself
Change MONTHLY_GPU_RENTAL to represent buying the hardware outright, amortised over 24 months, instead of renting monthly. Recompute the break-even point, and compare it to the rented version.
What to learn next
- The one-box production stack — what a minimal self-hosted setup can look like once you cross the break-even point.
- Choosing a GPU for inference — sizing the hardware side of a self-hosting decision.
- Docker for ML — packaging a self-hosted model so it runs reliably wherever you put it.
Researcher — Mathematics and papers.
The general break-even model
Let $v$ be monthly call volume, $p_{\text{api}}$ the API price per call, $F$ the self-hosted fixed cost (hardware amortisation or rental, plus operational overhead), and $c_{\text{marginal}}$ the small marginal cost of self-hosting per call (electricity, negligible compute overhead beyond the fixed capacity already paid for). Total cost under each strategy is
$$C_{\text{api}}(v) = p_{\text{api}} \cdot v, \qquad C_{\text{host}}(v) = F + c_{\text{marginal}} \cdot v$$
The break-even volume $v^$ solves $C_{\text{api}}(v^) = C_{\text{host}}(v^*)$:
$$v^* = \frac{F}{p_{\text{api}} - c_{\text{marginal}}}$$
Because $c_{\text{marginal}} \ll p_{\text{api}}$ almost always holds — self-hosted marginal cost is typically a small fraction of API price — this simplifies closely to $v^* \approx F / p_{\text{api}}$: break-even volume scales roughly linearly with fixed cost and inversely with API price.
Utilisation as the hidden variable
$F$ itself is not fixed in reality — it depends on how efficiently the self-hosted capacity is used. A GPU rented for the whole month but serving traffic only during business hours delivers a much higher effective per-call cost than the naive $F/v$ suggests, because $v$ was computed against calendar time, not against the fraction of capacity actually consumed. Reading GPU utilisation honestly and autoscaling both bear directly on making $F$ smaller in practice, which shifts $v^*$ down and makes self-hosting viable at lower reported volume than a naive calculation implies.
Option value and the cost of flexibility
A purely cost-based break-even calculation omits a real economic factor. API usage carries option value — the ability to change providers, models, or scale down to zero with no penalty. Self-hosting forfeits that, in exchange for its lower marginal cost at scale. Formally, this is analogous to the make-or-buy decision in operations research. Committing to fixed capacity trades flexibility for a lower expected cost, at a known, stable demand level. It is a worse trade whenever demand is volatile, or the technology is likely to change quickly — a genuine risk, with model quality and pricing both moving fast.
Hybrid strategies dominate pure strategies
In practice, the cost-minimal strategy is rarely a pure choice. Route baseline, predictable traffic to self-hosted capacity sized near its break-even utilisation. Route burst traffic above that baseline to an API. This captures most of self-hosting's savings, without paying for capacity that sits idle during low-traffic periods — the same idea as cloud "reserved plus on-demand" capacity planning, applied to model serving specifically.
What to learn next
- The one-box production stack — what a minimal self-hosted setup can look like once you cross the break-even point.
- Choosing a GPU for inference — sizing the hardware side of a self-hosting decision.
- Docker for ML — packaging a self-hosted model so it runs reliably wherever you put it.