Prompt an LLM or fine-tune a small model?
Prompting a large LLM needs no training but costs more per document. Fine-tuning a small model costs more upfront but gets cheaper at scale. The right choice depends on volume.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Prompting a large LLM needs no training and starts working today. Fine-tuning a small model costs more upfront but gets cheaper the more you use it.
Picture two ways to get a job done. Hire a freelance consultant for each task, paying every time. Or train a permanent employee once, paying a salary regardless of how much work comes in.
Both get the job done. Which one is smarter depends entirely on how much work there is. That is the real decision behind "prompt an LLM or fine-tune a small model."
Why it exists
A large LLM can summarise almost anything well, the moment you ask. Describe any style you want in the prompt. No training data needed, no waiting.
That convenience has a running cost. Every single document you summarise costs money, every time, forever. At low volume, this is by far the cheaper option overall.
A fine-tuned small model needs upfront work: gathering examples, training, testing. But once it is trained, running it costs far less per document. At high volume, that upfront cost pays for itself many times over.
How it works
Low volume: prompt an LLM
(no setup, higher cost per document, fine at small scale)
High volume: fine-tune a small model
(setup cost once, much lower cost per document after)
Somewhere in between: a break-even point
where both options cost the sameBelow that break-even point, prompting wins. Above it, fine-tuning wins. The exact point depends on your actual costs, which change over time as pricing changes.
Where you have already seen it
- A startup prototyping with a hosted LLM API, then later training their own model as they scale. The exact break-even trade-off in action.
- Large companies running their own smaller, fine-tuned models for high-volume internal tasks. Millions of documents a day makes the upfront cost worth it.
- A one-person side project prompting an LLM directly, nothing more. At that volume, fine-tuning would never pay for itself.
- News organisations fine-tuning house-style summarisers, while using a general LLM for one-off special requests. Both approaches, used where each makes sense.
Remember this
- Prompting an LLM has no setup cost, but a cost on every single document.
- Fine-tuning has a real upfront cost, but a much lower cost per document after that.
- The right choice depends on volume, not on which approach is "better" in general.
What to learn next
- Fine-tuning your own summariser — the fine-tuning half of this trade-off, in full.
- What is a large language model? — background on the prompting half of this trade-off.
- Extractive summarisation — the very first lesson in this section, if you are starting the whole path again.
Developer — Code and libraries.
Below is a small, illustrative model of the trade-off. The numbers are made up for demonstration. Real pricing changes constantly and varies by provider. Treat this as a way to think about the shape of the trade-off, not a live quote.
Setup
Nothing to install. This example uses only the standard library.
Finding the break-even point
# Illustrative figures only, not live pricing from any provider.
# Swap in your own numbers before trusting this for a real budget decision.
llm_cost_per_1k_docs = 8.00 # prompting a hosted LLM, per 1,000 documents
finetune_setup_cost = 40.00 # one-time: labelling plus a small GPU rental
finetuned_cost_per_1k_docs = 0.30 # running your own fine-tuned model, per 1,000
def total_cost(num_docs, use_finetuned):
docs_in_thousands = num_docs / 1000
if use_finetuned:
return finetune_setup_cost + docs_in_thousands * finetuned_cost_per_1k_docs
return docs_in_thousands * llm_cost_per_1k_docs
for docs in [500, 2_000, 10_000, 50_000]:
prompt_cost = total_cost(docs, use_finetuned=False)
finetune_cost = total_cost(docs, use_finetuned=True)
cheaper = "fine-tuned" if finetune_cost < prompt_cost else "prompted LLM"
print(f"{docs:>7,} docs prompt=${prompt_cost:7.2f} "
f"fine-tune=${finetune_cost:7.2f} cheaper: {cheaper}")500 docs prompt=$ 4.00 fine-tune=$ 40.15 cheaper: prompted LLM 2,000 docs prompt=$ 16.00 fine-tune=$ 40.60 cheaper: prompted LLM 10,000 docs prompt=$ 80.00 fine-tune=$ 43.00 cheaper: fine-tuned 50,000 docs prompt=$ 400.00 fine-tune=$ 55.00 cheaper: fine-tuned
Line by line
Every number here is a placeholder. The real ones depend on the LLM provider, the model size you fine-tune, and your own GPU pricing. What matters is the shape of the result, not these specific numbers.
finetune_setup_cost never changes with document count. It is paid once, regardless of whether you summarise 10 documents or 10 million.
The crossover happened between 2,000 and 10,000 documents. Below it, prompting was cheaper. Above it, fine-tuning was cheaper. Every real decision has a crossover point like this one, somewhere.
Cost is not the only variable. This model ignores quality, latency and maintenance effort entirely. A full decision needs all of them, not cost alone.
Common mistakes
Comparing cost while ignoring quality differences. A fine-tuned small model can match or beat a general LLM on a narrow task it was trained for. It can also lose badly outside that narrow task. Compare quality on your actual documents, not assumptions.
Forgetting maintenance cost. A fine-tuned model needs retraining as your data or requirements shift. A prompted LLM updates automatically when the provider improves it. Maintenance is a real, ongoing cost on both sides, only a different kind.
Ignoring latency requirements. A small fine-tuned model running locally can respond faster than a network call to a hosted LLM. For a real-time product feature, that difference can matter more than raw cost.
Assuming the decision is permanent. Many teams start by prompting an LLM, then fine-tune a smaller model later, once volume and requirements are clear. Starting simple and switching later is a completely reasonable strategy.
Try it yourself
Change finetune_setup_cost to 200.00, representing a more realistic labelling effort, and rerun. Watch how far the break-even point shifts to the right.
What to learn next
- Fine-tuning — the general technique behind the fine-tuning side of this comparison.
- fp32, bf16, fp8 and int4 — cutting the running cost of a fine-tuned model further.
- Dynamic batching — another major lever on real-world cost at high volume.
Researcher — Mathematics and papers.
The trade-off, formally
Total cost of ownership over n documents, for the LLM-prompting strategy:
C_llm(n) = n * p_llmp_llmis marginal cost per document: input and output tokens, multiplied by API price per token.
For the fine-tuning strategy:
C_ft(n) = S + n * p_ftSis fixed setup cost: data collection, labelling, compute for training, engineering time.p_ftis marginal serving cost per document, typically far belowp_llmfor a smaller model.
The break-even document count n* solves C_llm(n*) = C_ft(n*):
n* = S / (p_llm - p_ft)Below n*, prompting dominates on cost. Above it, fine-tuning dominates. This is a standard build-versus-buy calculation, structurally identical across many domains, not specific to summarisation.
Quality is not captured by this model at all
Cost crossover says nothing about output quality. A fine-tuned small model can exceed a much larger general LLM on a narrow, well-represented task. It has seen thousands of in-domain examples the general model never specifically trained on. Conversely, it can fail badly outside that narrow distribution, where a general LLM's broader training gives it an advantage.
Kaplan et al. (2020) and later scaling-law work show larger models are more sample-efficient in-context. A large LLM's few-shot prompt performance can rival a fine-tuned model's, on some tasks. No training data collection needed at all. This narrows, but does not eliminate, the quality gap that used to favour fine-tuning by default.
Variable costs beyond the token price
Latency. A fine-tuned model served locally or on dedicated infrastructure has no network round-trip to an external API. For latency-sensitive applications, this can matter more than the raw per-document cost difference.
Data control. Fine-tuning on sensitive documents can be done entirely within an organisation's own infrastructure. Prompting a hosted LLM sends that data to a third party. That is a real constraint in regulated industries, regardless of relative cost.
Maintenance drift. A fine-tuned model's quality is frozen at training time and can drift as real-world input distribution shifts. A hosted LLM updates behind an unchanged API, for better or worse, without the maintainer needing to retrain anything.
The hybrid pattern
Production systems increasingly combine both. A cheap, fine-tuned model handles the high-volume, well-understood majority of traffic. Harder or unusual cases route to a larger, prompted LLM instead. A lightweight classifier, or a confidence threshold on the small model's own output, decides which is which. This captures most of fine-tuning's cost benefit, while keeping a large LLM's general-case robustness as a fallback.
Key references
- Kaplan, J. et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361
- Brown, T. et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165 — the original in-context learning result underpinning the "prompt it directly" option.
- Hu, E. et al. (2021). LoRA. arXiv:2106.09685 — lowers
Sin the cost model above, by cutting fine-tuning's setup cost.
Current state and open problems
Falling hosted-LLM prices and improving small open-weight models are both moving n* in the developer block's model. They pull in opposite directions, at the same time. Cheaper API calls push n* higher, favouring prompting for longer. Better, cheaper-to-train small models push n* lower, favouring fine-tuning sooner. Which effect currently dominates depends on the specific task and month, not on any fixed rule.
The open problem: most teams choose from habit, or unfamiliarity with fine-tuning. An actual cost-quality analysis, like this lesson's, is rare. A small amount of real measurement, on real documents, at real volume, usually beats intuition here.
What to learn next
- Fine-tuning — full detail on the setup-cost side of this trade-off.
- fp32, bf16, fp8 and int4 — cutting
p_ftfurther once you have chosen to fine-tune. - Extractive summarisation — back to the start of the section, for a full second pass.