Observability for LLM Applications
Token and cost telemetry
Token and cost telemetry means recording exactly how much of a model's usage limit and money every single request consumed, the same way a meter records every unit of electricity used.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Token and cost telemetry means recording exactly how much a request cost, in both usage and money, every single time.
A token is a small chunk of text, often a word or part of a word. A language model reads and writes text one token at a time, and providers charge by how many go in and come out.
The analogy you have already lived
An electricity meter outside a house ticks up with every unit used. Nobody guesses the bill at the end of the month. The meter counted every single unit, all along.
An LLM app without token telemetry is a house with no meter. The bill still arrives. Nobody can explain, room by room, where the units went.
Why it exists
Every call to a language model has a real, per-token price. A short question costs little. A long document pasted into a prompt, sent a hundred thousand times a day, adds up fast.
Without recording tokens and cost per request, a monthly bill is a single, unexplained number. With it, the exact request, prompt, or feature driving the cost is one query away.
How it works
request ----> [ count tokens IN ] ----> call the model
|
v
[ count tokens OUT ] <----+
|
v
cost: (tokens in x price in)
+ (tokens out x price out)
|
v
recorded alongside the traceBoth counts, and the cost they produce, get logged next to the request's own trace span, from the previous lesson.
A real example you have seen
A prepaid mobile SIM shows exact data used, down to the megabyte, after every session. That is the same idea: usage is metered precisely, not estimated, so a surprising bill always has an explanation attached.
The honest part
Counting tokens correctly is trickier than it sounds. Different providers, and even different versions of the same provider's models, split text into tokens slightly differently.
A token count from one tokenizer is an estimate for a different model, not an exact figure. Use the specific tokenizer for the specific model you are calling, whenever it is available.
Remember this
- A token is the small chunk of text a model is billed by.
- Cost telemetry means recording tokens and price per request, not only a monthly total.
- Different models tokenize text differently — a count from one is an estimate for another.
What to learn next
- Versioning prompts — a prompt change is often the reason cost or token counts moved.
- Tracing an LLM application — where this telemetry gets attached, per request.
- Replaying production traffic — reusing logged requests like these to test a cheaper model safely.
Developer — Code and libraries.
Setup
pip install tiktokentiktoken downloads its encoding file the first time get_encoding runs, so that first run needs an internet connection. It is cached locally after that.
Counting tokens and cost per call
tiktoken is OpenAI's own tokenizer library, and a reasonable stand-in for exploring the idea even against other providers, since most modern tokenizers behave similarly on English text.
import tiktoken
encoder = tiktoken.get_encoding("cl100k_base")
# Price per 1,000 tokens, in a made-up currency unit. Real prices change
# often and depend on the exact model -- always read these from the
# provider's current pricing page, never hardcode them long-term.
PRICE_PER_1K = {"input": 0.15, "output": 0.60}
calls = [
{
"prompt": "Summarise this support ticket in one sentence: "
"customer says the app crashes when uploading a photo over 10MB.",
"completion": "The app crashes when a user uploads a photo larger than 10MB.",
},
{
"prompt": "Translate to Hindi: Your order has been shipped.",
"completion": "aapka order bhej diya gaya hai.",
},
]
total_cost = 0.0
total_input_tokens = 0
total_output_tokens = 0
for i, call in enumerate(calls, start=1):
input_tokens = len(encoder.encode(call["prompt"]))
output_tokens = len(encoder.encode(call["completion"]))
cost = (input_tokens / 1000) * PRICE_PER_1K["input"] + \
(output_tokens / 1000) * PRICE_PER_1K["output"]
total_input_tokens += input_tokens
total_output_tokens += output_tokens
total_cost += cost
print(f"call {i}: input={input_tokens} tokens output={output_tokens} tokens cost=${cost:.5f}")
print()
print(f"total input tokens: {total_input_tokens}")
print(f"total output tokens: {total_output_tokens}")
print(f"total cost: ${total_cost:.5f}")call 1: input=24 tokens output=15 tokens cost=$0.01260 call 2: input=10 tokens output=13 tokens cost=$0.00930 total input tokens: 34 total output tokens: 28 total cost: $0.02190
The token counts and cost figures above are exact and reproducible — tiktoken is deterministic for a fixed piece of text and a fixed encoding. The PRICE_PER_1K values are made up for this lesson; substitute your real provider's current published prices.
Line-by-line walkthrough
encoder.encode(text) returns a list of integer token IDs. Its length is the token count — the code never needs to look at what the tokens actually are.
The cost formula treats input and output tokens separately, at different prices, because most providers price them differently. Output tokens are frequently priced several times higher than input tokens.
Nothing here calls a real model. The prompt and completion strings are written by hand, standing in for a real request-response pair, so the token counts are real even though no API was called.
Common mistakes
Using one tokenizer for every provider's model. A token count from cl100k_base is close for many models and exact for none but OpenAI's own. When a provider ships its own tokenizer, use it.
Recording only total cost, with no split by feature or endpoint. A total that cannot be broken down by which part of the product caused it is a number you cannot act on.
Forgetting that a failed or retried call still cost tokens. A request that errors out after the model already generated a response still burned real tokens. Log cost regardless of whether the call ultimately succeeded.
Never checking whether a long, unnecessary prompt is driving the bill. Logging cost per call makes it visible when one part of the system is quietly sending a much larger prompt than it needs to.
Try it yourself
Add a third call whose prompt pastes in a long block of repeated text, five times longer than the others. Rerun, and check how much of the total cost that one call ends up responsible for.
What to learn next
- Versioning prompts — a prompt change is often the reason cost or token counts moved.
- Tracing an LLM application — where this telemetry gets attached, per request.
- Replaying production traffic — reusing logged requests like these to test a cheaper model safely.
Researcher — Mathematics and papers.
Byte-pair encoding, briefly
cl100k_base and similar tokenizers are trained with byte-pair encoding (BPE): starting from individual bytes, the most frequent adjacent pair in a training corpus is repeatedly merged into a new symbol, until a target vocabulary size is reached (Sennrich et al., 2016, originally for neural machine translation; adapted for GPT-family models). This is why token counts do not map cleanly onto words — common words become one token, and rare words split into several sub-word pieces.
Why cross-model token counts do not transfer
Different model families train their own BPE (or unigram-language-model) vocabularies on different corpora, so the same string tokenizes to a different length under each. As a rough calibration for English prose, cl100k_base averages roughly 1.3 tokens per word — useful for back-of-envelope estimates, unreliable for billing-accurate figures on a model whose actual tokenizer is unknown or unavailable locally.
Building telemetry that scales
At production volume, per-request cost logging is best implemented as a structured event emitted alongside the trace span from the previous lesson — gen_ai.usage.input_tokens and gen_ai.usage.output_tokens are the relevant OpenTelemetry GenAI semantic-convention attribute names, keeping cost data queryable in the same system as latency and error data, rather than in a separate spreadsheet.
$$C = \sum_{i=1}^{n} \left( t^{\text{in}}_i \cdot p^{\text{in}} + t^{\text{out}}_i \cdot p^{\text{out}} \right)$$
Where $C$ is total cost over $n$ requests, $t^{\text{in}}_i$ and $t^{\text{out}}_i$ are the input and output token counts for request $i$, and $p^{\text{in}}$, $p^{\text{out}}$ are the per-token prices. Aggregating this by a dimension — endpoint, customer, feature flag — turns it from a single number into a cost-attribution report, the LLM-era equivalent of cloud cost tagging.
Cost-reduction levers this telemetry reveals
Once cost is attributed per request, the standard levers become visible as concrete, measurable opportunities rather than guesses: prompt caching for repeated context, shorter system prompts, routing simple requests to a smaller model, and truncating retrieved context to only what is actually used downstream.
Papers
- Sennrich, Haddow and Birch, Neural Machine Translation of Rare Words with Subword Units, ACL 2016 — arxiv.org/abs/1508.07909, the BPE paper.
- Kudo and Richardson, SentencePiece: A Simple and Language Independent Subword Tokenizer, EMNLP 2018 — a widely used alternative tokenization approach, relevant since not every model family uses BPE.
What to learn next
- Versioning prompts — a prompt change is often the reason cost or token counts moved.
- Tracing an LLM application — where this telemetry gets attached, per request.
- Replaying production traffic — reusing logged requests like these to test a cheaper model safely.