Caching and Cost Control

Unit economics of an AI feature

Unit economics means working out whether a single use of your AI feature makes money or loses it, once every real cost is counted.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. What actually gets counted
  5. How it works
  6. A real example you have seen
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Unit economics means working out whether one single use of your feature makes money or loses it. Every real cost has to be counted.

The analogy you have already lived

A tiffin service charges a fixed amount per tiffin. To know if that price actually makes sense, someone has to add it all up. The vegetables, the rice, the gas for cooking, the container, the delivery person's time. Per tiffin, not per month.

If all of that adds up to more than the price charged, every single tiffin sold loses money. Selling more of them makes the business worse, not better. Getting this number right, per unit, matters more than getting the total revenue number right.

Why it exists

An AI feature can seem to be doing great — lots of users, lots of usage — and still quietly lose money on every single use. Growth hides a bad unit economics problem for a while, and then makes it much bigger.

The whole point of measuring cost per use is to catch this early. Before the business gets built on top of a number that does not actually work.

What actually gets counted

Every real cost that a single use causes has to be added up, not just the obvious one.

  • The model cost. A paid API call, or a share of the electricity and hardware for a model you run yourself.
  • The infrastructure cost — servers, storage, monitoring — divided fairly across every use.
  • Anything saved by caching or other tricks from earlier lessons, which lowers this number.

Against that, compare what one use is actually worth. What a customer pays for it, or what value it creates if there is no direct price.

How it works

   cost to run the model, for ONE use
 + a fair share of infrastructure cost, for ONE use
 -----------------------------------------------------
 = total cost, for ONE use

   compare that to what ONE use is worth

   worth more than it costs?  -> the feature is a healthy unit
   costs more than it's worth? -> more usage makes things worse, not better

A real example you have seen

Free apps that suddenly start charging, or start limiting free usage, after growing very fast. Often the reason is exactly this. The free version's unit economics never worked, and it only became visible once enough people were using it for the total bill to matter.

The honest part

A feature can lose money per use on purpose, for a while. It acts as a loss leader to attract customers, while the team works on the cost side. That is a real, valid business decision. The failure is not losing money per use. The failure is not knowing you are, and being surprised by the bill later.

Remember this

  • Unit economics asks: does one single use of this feature make money or lose it?
  • Every real cost has to be counted, including infrastructure, not only the obvious model cost.
  • Losing money per use can be a deliberate choice. Not knowing you are is the actual danger.

What to learn next

Developer — Code and libraries.

Setup

No installs needed — this uses only Python's standard library.

A small unit economics calculator

unit_economics.py
from dataclasses import dataclass

@dataclass
class FeatureCosts:
    llm_cost_per_call: float       # $ per uncached model call
    infra_cost_per_month: float    # $ fixed cost: servers, monitoring, etc.
    calls_per_month: int
    cache_hit_rate: float          # 0.0 to 1.0, fraction served from cache
    price_per_use: float           # $ the customer pays, or the internal value assigned

def unit_economics(c: FeatureCosts):
    uncached_calls = c.calls_per_month * (1 - c.cache_hit_rate)
    model_cost = uncached_calls * c.llm_cost_per_call
    total_cost = model_cost + c.infra_cost_per_month
    revenue = c.calls_per_month * c.price_per_use
    margin = revenue - total_cost
    return {
        "cost_per_call": round(total_cost / c.calls_per_month, 4),
        "total_cost": round(total_cost, 2),
        "revenue": round(revenue, 2),
        "margin": round(margin, 2),
        "margin_pct": round(margin / revenue * 100, 1),
    }

before_caching = FeatureCosts(
    llm_cost_per_call=0.01, infra_cost_per_month=500,
    calls_per_month=200_000, cache_hit_rate=0.0, price_per_use=0.02,
)
after_caching = FeatureCosts(
    llm_cost_per_call=0.01, infra_cost_per_month=500,
    calls_per_month=200_000, cache_hit_rate=0.6, price_per_use=0.02,
)

for label, config in [("no caching", before_caching), ("60% cache hit rate", after_caching)]:
    print(f"--- {label} ---")
    for k, v in unit_economics(config).items():
        print(f"  {k:14}: {v}")
Output
--- no caching ---
  cost_per_call : 0.0125
  total_cost    : 2500.0
  revenue       : 4000.0
  margin        : 1500.0
  margin_pct    : 37.5
--- 60% cache hit rate ---
  cost_per_call : 0.0065
  total_cost    : 1300.0
  revenue       : 4000.0
  margin        : 2700.0
  margin_pct    : 67.5

This output is exact, deterministic arithmetic. Plug in the same numbers and you get exactly these figures, on any machine. What varies between businesses is which numbers to plug in, not the arithmetic itself.

Line-by-line walkthrough

uncached_calls = calls_per_month * (1 - cache_hit_rate). Only calls that actually reach the model cost money at llm_cost_per_call. This is the direct, calculable link between semantic caching or prompt prefix caching and the bottom line. A 60% hit rate here nearly halves total cost, and almost doubles margin.

infra_cost_per_month is fixed, not per call. Servers and monitoring cost roughly the same whether 10,000 or 200,000 calls happen. Dividing a fixed cost across a variable number of calls means this number improves automatically as usage grows. That is the opposite of the model cost, which grows with usage.

margin_pct. The percentage is often more useful for decisions than the absolute margin. It is comparable across features of very different sizes.

Common mistakes

Forgetting infrastructure cost entirely. Counting only the per-call model price makes every feature look healthier than it is. Servers, logging, monitoring and storage cost money whether or not a single model call happens.

Averaging across very different usage patterns. A feature might be profitable for light users and lose money on heavy ones. A single blended number can hide a real problem inside an average that looks fine.

Treating the cache hit rate as a constant. It is a measured number that changes as your traffic changes. Watch it — do not assume it once and forget it. See semantic caching for how a threshold shift alone can move it.

Ignoring cost that scales with something other than call count. Storage cost for a growing set of cached embeddings, for instance, grows with how much content exists. Not how many requests happen. That is a different variable your formula needs to account for separately.

Try it yourself

Add a storage_cost_per_month field that scales with a new documents_indexed number, independent of calls_per_month, and fold it into total_cost. Recompute the margin for a feature with a lot of content but comparatively few calls.

What to learn next

Researcher — Mathematics and papers.

Contribution margin and the make-or-buy decision

The per-unit calculation above is a contribution margin: revenue per unit minus variable cost per unit, deliberately excluding fixed costs that do not change with volume. Contribution margin, not gross margin, is the correct number for deciding whether to accept one more unit of usage. A feature can have a low or even negative overall margin at current volume, while still having a positive contribution margin — meaning growth genuinely helps rather than hurts.

This distinction feeds directly into self-hosting vs API: the break-even point. Self-hosting typically converts a variable cost, per-call API pricing, into a larger fixed cost — owned or reserved hardware — plus a much smaller variable cost: electricity, amortised depreciation. The crossover volume where self-hosting's lower marginal cost overtakes its higher fixed cost is exactly a contribution-margin calculation applied to two different cost structures.

Cost allocation for shared infrastructure

When infrastructure serves multiple features, infra_cost_per_month in the demo needs an allocation rule, and the rule chosen changes every downstream decision. Common approaches, in increasing order of accuracy and implementation cost:

  • Even split across features — simplest, systematically wrong whenever features differ meaningfully in resource intensity.
  • Usage-proportional — allocate by each feature's share of total calls. Better, but it still ignores that some calls are far more expensive than others, like a long RAG context versus a short classification call.
  • Resource-proportional — allocate by measured GPU-seconds, memory, or tokens actually consumed per feature. The most accurate approach, and the only one that correctly rewards a feature for adopting the caching techniques earlier in this section.

Marginal versus average cost at scale

Average cost per call, as computed in the demo, masks a distinction that matters at scale. Marginal cost — the cost of the next call — is often lower than average cost once fixed infrastructure is already paid for. It can fall further still with volume discounts from API providers or cloud vendors. Capacity planning and pricing decisions should generally be made against marginal cost, not average cost, since average cost systematically overstates the true cost of growth once fixed costs are sunk.

What to learn next