Caching and Cost Control

Cutting LLM API costs

Most of an LLM API bill comes from a small number of controllable choices — how long the prompt is, which model answers, and how often you ask at all.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The levers, in plain words
  5. How it works
  6. A real example you have seen
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Most of an LLM API bill comes from a few controllable choices. How long your prompts are, which model answers, and how often you ask at all.

The analogy you have already lived

A grocery bill goes up because of a hundred small choices. A bigger pack than needed. The branded item over the ordinary one. Food thrown away after it spoiled. Cutting the bill rarely means eating less. It means shopping smarter.

An LLM API bill works the same way. Cutting it does not usually mean answering fewer questions. It means asking better questions, in fewer words, of the right model. It means not asking twice for something you already know.

Why it exists

LLM APIs typically charge by the token. A token is a small chunk of text, roughly a word or part of a word. A longer prompt costs more. A longer answer costs more. A more capable model usually costs more per token than a smaller one.

None of that is hidden. It is printed on every provider's pricing page. Most teams still overspend, because these small choices are made once, early, and never revisited as usage grows.

The levers, in plain words

Shorter prompts. Every instruction, every example, every unnecessary sentence in your prompt is paid for, every single time it is sent.

The right-sized model. A simple task — classifying a message as "spam" or "not spam" — rarely needs the most expensive model available. Save the expensive model for genuinely hard questions.

Not asking twice. Every earlier lesson in this section directly reduces how often the model gets called at all: caching answers, catching paraphrases, reusing a shared prompt prefix.

Shorter answers, where a short answer is enough. Asking for a 200-word explanation when 20 words would answer the question pays for 180 words nobody reads.

How it works

   your prompt, as written
              |
   trim it: remove instructions and examples that don't change the answer
              |
   send it to the SMALLEST model that can actually do this task well
              |
   check the cache first (see the earlier lessons in this section)
              |
   ask for only as much output as you actually need

A real example you have seen

Customer support tools often use a small, fast model to sort incoming messages into categories. They call in a bigger, more expensive model only for the messages that are genuinely complicated. Nobody notices the difference in quality. The bill for sorting messages barely registers, next to what it would cost to run every message through the expensive model.

The honest part

Cutting cost this way has a floor. A genuinely hard question needs a genuinely capable model, and a genuinely long document needs a genuinely long prompt. The goal is removing waste, not removing capability the task actually needs. Cutting past that floor shows up as worse answers, not a smaller bill worth having.

Remember this

  • Most API cost comes from prompt length, model choice, and how often you ask.
  • Cutting cost usually means removing waste, not removing capability the task needs.
  • The techniques from earlier in this section are cost levers too, not only speed ones: caching, semantic matching, prefix reuse.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install tiktoken

Measuring the effect of prompt length and model tier

cost_levers.py
import tiktoken

enc = tiktoken.get_encoding("cl100k_base")

VERBOSE_PROMPT = """You are a highly capable and knowledgeable assistant. Please read the
following customer support ticket very carefully and thoroughly, considering
all possible angles and interpretations, and then provide a comprehensive,
detailed, and thorough response that addresses every possible aspect of the
customer's concern in a friendly and professional manner.

Ticket: My order #4821 hasn't arrived yet. It's been 10 days."""

TIGHT_PROMPT = """Answer the support ticket in 2 sentences.

Ticket: Order #4821 hasn't arrived in 10 days."""

# Illustrative example rates, not live prices -- check your provider's
# current pricing page before using numbers like these for real decisions.
PRICE_PER_1K_INPUT_TOKENS = {"large-tier": 0.0030, "small-tier": 0.00015}

def cost_for(prompt, tier, calls_per_day):
    tokens = len(enc.encode(prompt))
    daily_cost = (tokens * calls_per_day / 1000) * PRICE_PER_1K_INPUT_TOKENS[tier]
    return tokens, daily_cost

CALLS_PER_DAY = 50_000
for label, prompt, tier in [
    ("verbose prompt, large-tier model", VERBOSE_PROMPT, "large-tier"),
    ("tight prompt, large-tier model",   TIGHT_PROMPT,   "large-tier"),
    ("tight prompt, small-tier model",   TIGHT_PROMPT,   "small-tier"),
]:
    tokens, daily_cost = cost_for(prompt, tier, CALLS_PER_DAY)
    print(f"{label:35} {tokens:4} tokens/call   ${daily_cost:8.2f}/day")
Output
verbose prompt, large-tier model      79 tokens/call   $   11.85/day
tight prompt, large-tier model        23 tokens/call   $    3.45/day
tight prompt, small-tier model        23 tokens/call   $    0.17/day

The token counts are exact and reproducible — tiktoken deterministically counts the same text the same way every time. The dollar figures use made-up example prices, and exist only to show the shape of the saving. Substitute your real provider's current per-token price to get a real number.

Line-by-line walkthrough

tiktoken.get_encoding("cl100k_base"). The exact tokenizer varies by model and provider — this one matches several OpenAI models. Different models can tokenize the same text into a different number of tokens. Always check which tokenizer matches the model you are actually billed for.

The two prompts encode to 79 versus 23 tokens. Nothing about what the model is asked to do changed in a way that matters. Both prompts ask it to address the same ticket. The verbose version spends tokens on politeness and repetition that the model does not need to do its job well.

Combining both levers multiplies the saving. Tightening the prompt alone cut the bill by about 70% here. Moving to a smaller model on top of that cut it by another 95%. The two levers are independent and stack.

Common mistakes

Trimming a prompt until the model's answers get worse. Every instruction removed should be one that genuinely was not changing the output. Test before and after on a real set of examples, not just on cost.

Assuming a smaller model is always worse. For narrow, well-defined tasks — classification, extraction, short factual answers — a small model can match a large one closely. Often at a fraction of the cost. This is worth testing directly on your own task, not assumed either way.

Optimising input tokens and ignoring output tokens. Output is usually priced too, sometimes at a higher rate than input. An open-ended prompt that invites a long answer can cost more in output than the entire input side combined.

Cutting cost without checking quality regularly. A model provider can change pricing, or a prompt tweak can quietly change behaviour. Recheck both cost and quality on a schedule, not only once when the system was first built.

Try it yourself

Add a third prompt variant that keeps the instructions but removes the greeting and sign-off politeness language only. Measure how many tokens that alone saves, separate from shortening the instructions.

What to learn next

Researcher — Mathematics and papers.

Where API cost is actually spent

For a request-response LLM call, total cost decomposes as

$$\text{Cost} = n_{\text{in}} \cdot p_{\text{in}} + n_{\text{out}} \cdot p_{\text{out}}$$

where $n_{\text{in}}$ and $n_{\text{out}}$ are input and output token counts and $p_{\text{in}}$, $p_{\text{out}}$ their respective per-token prices. Providers commonly price $p_{\text{out}} > p_{\text{in}}$, sometimes by 3–5x, because generation is autoregressive and inherently more expensive per token than the parallel processing of a prompt — see prefill and decode for why the two phases have such different cost profiles. This asymmetry means constraining output length (explicit instructions, max_tokens, stop sequences) is frequently the single highest-leverage lever available, and is undervalued relative to prompt trimming in most cost-cutting discussions.

Model routing as a cost-quality frontier

Choosing model tier per request is a special case of the general cost-quality Pareto frontier. For a fixed budget, the routing policy that maximises aggregate quality sends each request to the cheapest model expected to answer it correctly, escalating only when that model's confidence is low. This is formalised and extended in model cascades. The cost curve produced by good routing is provably better than any single fixed model choice, whenever task difficulty varies across requests — which it does in essentially every real workload.

Prompt compression as an active research area

Beyond manual prompt trimming, learned prompt compression methods (LLMLingua, Jiang et al., 2023) train a small model to identify and remove tokens from a prompt with minimal impact on downstream task performance. Reported compression ratios reach 2–5x on some benchmarks, with small measured quality loss. This automates what the demo does by hand, at the cost of an additional model in the pipeline and its own (much smaller) compute cost.

Batching and rate-limit economics

Many providers price batch or asynchronous APIs, where results are not needed within seconds, at a discount relative to synchronous calls — commonly reported around 50%. This lets the provider schedule the work against spare capacity. Any workload without a hard real-time latency requirement (see offline batch scoring) should default to the batch tier unless there is a specific reason not to.

Reading

What to learn next