Cutting LLM API costs
Most of an LLM API bill comes from a small number of controllable choices — how long the prompt is, which model answers, and how often you ask at all.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Most of an LLM API bill comes from a few controllable choices. How long your prompts are, which model answers, and how often you ask at all.
The analogy you have already lived
A grocery bill goes up because of a hundred small choices. A bigger pack than needed. The branded item over the ordinary one. Food thrown away after it spoiled. Cutting the bill rarely means eating less. It means shopping smarter.
An LLM API bill works the same way. Cutting it does not usually mean answering fewer questions. It means asking better questions, in fewer words, of the right model. It means not asking twice for something you already know.
Why it exists
LLM APIs typically charge by the token. A token is a small chunk of text, roughly a word or part of a word. A longer prompt costs more. A longer answer costs more. A more capable model usually costs more per token than a smaller one.
None of that is hidden. It is printed on every provider's pricing page. Most teams still overspend, because these small choices are made once, early, and never revisited as usage grows.
The levers, in plain words
Shorter prompts. Every instruction, every example, every unnecessary sentence in your prompt is paid for, every single time it is sent.
The right-sized model. A simple task — classifying a message as "spam" or "not spam" — rarely needs the most expensive model available. Save the expensive model for genuinely hard questions.
Not asking twice. Every earlier lesson in this section directly reduces how often the model gets called at all: caching answers, catching paraphrases, reusing a shared prompt prefix.
Shorter answers, where a short answer is enough. Asking for a 200-word explanation when 20 words would answer the question pays for 180 words nobody reads.
How it works
your prompt, as written
|
trim it: remove instructions and examples that don't change the answer
|
send it to the SMALLEST model that can actually do this task well
|
check the cache first (see the earlier lessons in this section)
|
ask for only as much output as you actually needA real example you have seen
Customer support tools often use a small, fast model to sort incoming messages into categories. They call in a bigger, more expensive model only for the messages that are genuinely complicated. Nobody notices the difference in quality. The bill for sorting messages barely registers, next to what it would cost to run every message through the expensive model.
The honest part
Cutting cost this way has a floor. A genuinely hard question needs a genuinely capable model, and a genuinely long document needs a genuinely long prompt. The goal is removing waste, not removing capability the task actually needs. Cutting past that floor shows up as worse answers, not a smaller bill worth having.
Remember this
- Most API cost comes from prompt length, model choice, and how often you ask.
- Cutting cost usually means removing waste, not removing capability the task needs.
- The techniques from earlier in this section are cost levers too, not only speed ones: caching, semantic matching, prefix reuse.
What to learn next
- Model cascades — automating the "right-sized model" choice per request instead of picking one model for everything.
- Semantic caching — the biggest lever of all: not calling the model twice for the same question.
- Self-hosting vs API: the break-even point — when cutting the per-call price stops being enough, and owning the model becomes cheaper.
Developer — Code and libraries.
Setup
pip install tiktokenMeasuring the effect of prompt length and model tier
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
VERBOSE_PROMPT = """You are a highly capable and knowledgeable assistant. Please read the
following customer support ticket very carefully and thoroughly, considering
all possible angles and interpretations, and then provide a comprehensive,
detailed, and thorough response that addresses every possible aspect of the
customer's concern in a friendly and professional manner.
Ticket: My order #4821 hasn't arrived yet. It's been 10 days."""
TIGHT_PROMPT = """Answer the support ticket in 2 sentences.
Ticket: Order #4821 hasn't arrived in 10 days."""
# Illustrative example rates, not live prices -- check your provider's
# current pricing page before using numbers like these for real decisions.
PRICE_PER_1K_INPUT_TOKENS = {"large-tier": 0.0030, "small-tier": 0.00015}
def cost_for(prompt, tier, calls_per_day):
tokens = len(enc.encode(prompt))
daily_cost = (tokens * calls_per_day / 1000) * PRICE_PER_1K_INPUT_TOKENS[tier]
return tokens, daily_cost
CALLS_PER_DAY = 50_000
for label, prompt, tier in [
("verbose prompt, large-tier model", VERBOSE_PROMPT, "large-tier"),
("tight prompt, large-tier model", TIGHT_PROMPT, "large-tier"),
("tight prompt, small-tier model", TIGHT_PROMPT, "small-tier"),
]:
tokens, daily_cost = cost_for(prompt, tier, CALLS_PER_DAY)
print(f"{label:35} {tokens:4} tokens/call ${daily_cost:8.2f}/day")verbose prompt, large-tier model 79 tokens/call $ 11.85/day tight prompt, large-tier model 23 tokens/call $ 3.45/day tight prompt, small-tier model 23 tokens/call $ 0.17/day
The token counts are exact and reproducible — tiktoken deterministically counts the same text the same way every time. The dollar figures use made-up example prices, and exist only to show the shape of the saving. Substitute your real provider's current per-token price to get a real number.
Line-by-line walkthrough
tiktoken.get_encoding("cl100k_base"). The exact tokenizer varies by model and provider — this one matches several OpenAI models. Different models can tokenize the same text into a different number of tokens. Always check which tokenizer matches the model you are actually billed for.
The two prompts encode to 79 versus 23 tokens. Nothing about what the model is asked to do changed in a way that matters. Both prompts ask it to address the same ticket. The verbose version spends tokens on politeness and repetition that the model does not need to do its job well.
Combining both levers multiplies the saving. Tightening the prompt alone cut the bill by about 70% here. Moving to a smaller model on top of that cut it by another 95%. The two levers are independent and stack.
Common mistakes
Trimming a prompt until the model's answers get worse. Every instruction removed should be one that genuinely was not changing the output. Test before and after on a real set of examples, not just on cost.
Assuming a smaller model is always worse. For narrow, well-defined tasks — classification, extraction, short factual answers — a small model can match a large one closely. Often at a fraction of the cost. This is worth testing directly on your own task, not assumed either way.
Optimising input tokens and ignoring output tokens. Output is usually priced too, sometimes at a higher rate than input. An open-ended prompt that invites a long answer can cost more in output than the entire input side combined.
Cutting cost without checking quality regularly. A model provider can change pricing, or a prompt tweak can quietly change behaviour. Recheck both cost and quality on a schedule, not only once when the system was first built.
Try it yourself
Add a third prompt variant that keeps the instructions but removes the greeting and sign-off politeness language only. Measure how many tokens that alone saves, separate from shortening the instructions.
What to learn next
- Model cascades — automating the "right-sized model" choice per request instead of picking one model for everything.
- Semantic caching — the biggest lever of all: not calling the model twice for the same question.
- Self-hosting vs API: the break-even point — when cutting the per-call price stops being enough, and owning the model becomes cheaper.
Researcher — Mathematics and papers.
Where API cost is actually spent
For a request-response LLM call, total cost decomposes as
$$\text{Cost} = n_{\text{in}} \cdot p_{\text{in}} + n_{\text{out}} \cdot p_{\text{out}}$$
where $n_{\text{in}}$ and $n_{\text{out}}$ are input and output token counts and $p_{\text{in}}$, $p_{\text{out}}$ their respective per-token prices. Providers commonly price $p_{\text{out}} > p_{\text{in}}$, sometimes by 3–5x, because generation is autoregressive and inherently more expensive per token than the parallel processing of a prompt — see prefill and decode for why the two phases have such different cost profiles. This asymmetry means constraining output length (explicit instructions, max_tokens, stop sequences) is frequently the single highest-leverage lever available, and is undervalued relative to prompt trimming in most cost-cutting discussions.
Model routing as a cost-quality frontier
Choosing model tier per request is a special case of the general cost-quality Pareto frontier. For a fixed budget, the routing policy that maximises aggregate quality sends each request to the cheapest model expected to answer it correctly, escalating only when that model's confidence is low. This is formalised and extended in model cascades. The cost curve produced by good routing is provably better than any single fixed model choice, whenever task difficulty varies across requests — which it does in essentially every real workload.
Prompt compression as an active research area
Beyond manual prompt trimming, learned prompt compression methods (LLMLingua, Jiang et al., 2023) train a small model to identify and remove tokens from a prompt with minimal impact on downstream task performance. Reported compression ratios reach 2–5x on some benchmarks, with small measured quality loss. This automates what the demo does by hand, at the cost of an additional model in the pipeline and its own (much smaller) compute cost.
Batching and rate-limit economics
Many providers price batch or asynchronous APIs, where results are not needed within seconds, at a discount relative to synchronous calls — commonly reported around 50%. This lets the provider schedule the work against spare capacity. Any workload without a hard real-time latency requirement (see offline batch scoring) should default to the batch tier unless there is a specific reason not to.
Reading
- Jiang et al., LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, EMNLP 2023 — arxiv.org/abs/2310.05736
What to learn next
- Model cascades — automating the "right-sized model" choice per request instead of picking one model for everything.
- Semantic caching — the biggest lever of all: not calling the model twice for the same question.
- Self-hosting vs API: the break-even point — when cutting the per-call price stops being enough, and owning the model becomes cheaper.