Observability for LLM Applications

Observability for agent runs

An agent run takes a different path through tools and decisions almost every time, so watching it needs every earlier technique in this section wired into a full record of that path.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An agent run takes a different path through decisions and tools each time. Watching one means recording the full path, not only the final answer.

An agent is a system that decides its own next step, instead of following one fixed sequence.

The analogy you have already lived

A relay race has four runners passing one baton. If the team loses, nobody only looks at the final time.

They review each leg separately: who was slow out of the blocks, where the handoff fumbled, which runner actually lost the race. The final time alone cannot show any of that.

An agent's run is that relay. A wrong final answer could have gone wrong at any one of several handoffs along the way.

Why it exists

A simple LLM call has one step: prompt in, answer out. Tracing it, from earlier in this section, is straightforward.

An agent can call a tool, read the result, decide to call another tool, and repeat. The number of times, and the order, differs on every run. Watching only the final answer hides every decision that led there.

How it works

   agent_run
     think (step 0)  -->  decides: use the calculator
     tool_call: calculator("240 * 0.15")  -->  36.0
     think (step 1)  -->  decides: search instead
     tool_call: search("who invented the calculator")  -->  a fact
     think (step 2)  -->  decides: enough information, answer now
     final_answer

Every decision and every tool call becomes its own recorded step, nested inside the run. It is the same span-and-trace shape from earlier in this section, applied to a path that branches differently each time.

A real example you have seen

A travel-booking assistant that searches flights, checks a calendar, and books a ticket can fail at any one of those steps. Without a full record of the path, "why did it book the wrong date?" cannot be answered. It can only be guessed at.

The honest part

Agents can genuinely get stuck: calling the same tool with the same input, over and over, never reaching an answer. This is common enough to plan for, not a rare edge case.

A step-count limit that forces a stop is not a workaround. It is a required safety net for a system that decides its own number of steps.

Remember this

  • An agent decides its own next step, so its path differs on every run.
  • Every decision and tool call needs its own recorded step, not only the final answer.
  • A step-count limit is a required safety net, because agents can genuinely get stuck.

What to learn next

Developer — Code and libraries.

Setup

No install needed. This script only uses Python's standard library.

Tracing a healthy run, and a stuck one

Reusing the span tracer from tracing an LLM application, trace two runs of a small agent: one that reaches an answer, and one whose policy has a bug and never does.

agent_trace.py
import time
import uuid
from contextlib import contextmanager

TRACE = []
_stack = []
_t0 = time.perf_counter()


@contextmanager
def span(name, **attributes):
    span_id = str(uuid.uuid4())[:8]
    start = time.perf_counter()
    depth = len(_stack)
    _stack.append(span_id)
    try:
        yield
    finally:
        _stack.pop()
        duration_ms = (time.perf_counter() - start) * 1000
        TRACE.append({"name": name, "start_ms": (start - _t0) * 1000,
                       "duration_ms": duration_ms, "depth": depth, **attributes})


def print_trace():
    TRACE.sort(key=lambda r: r["start_ms"])
    for row in TRACE:
        attrs = {k: v for k, v in row.items()
                 if k not in ("name", "start_ms", "duration_ms", "depth")}
        indent = "  " * row["depth"]
        label = f"{indent}{row['name']}"
        print(f"{label:<32}{attrs}")


def calculator_tool(expression):
    return eval(expression, {"__builtins__": {}})


def search_tool(query):
    facts = {"who invented the calculator": "Blaise Pascal built an early mechanical calculator in 1642."}
    return facts.get(query.lower(), "no result found")


def fake_agent_policy(step, scratchpad):
    """
    Stands in for an LLM deciding the next action given the scratchpad so
    far. A real agent asks a model "what should I do next?" and parses its
    reply; this fixed, deterministic script plays that role so the whole
    example runs with no API key.
    """
    plan = [
        ("calculator", "240 * 0.15"),
        ("search", "who invented the calculator"),
        ("final_answer", None),
    ]
    return plan[step]


def run_agent(max_steps=5):
    scratchpad = []
    with span("agent_run", max_steps=max_steps):
        for step in range(max_steps):
            with span("think", step=step):
                tool, arg = fake_agent_policy(step, scratchpad)

            if tool == "final_answer":
                with span("final_answer"):
                    answer = (f"15% of 240 is {scratchpad[0]}. "
                              f"The calculator was invented by: {scratchpad[1]}")
                return answer, step + 1

            with span("tool_call", tool=tool, arg=str(arg)):
                if tool == "calculator":
                    result = calculator_tool(arg)
                elif tool == "search":
                    result = search_tool(arg)
                scratchpad.append(result)

    return None, max_steps  # ran out of steps without finishing


answer, steps_used = run_agent(max_steps=5)
print(f"answer: {answer}")
print(f"steps used: {steps_used}\n")
print_trace()


def buggy_agent_policy(step, scratchpad):
    """A policy with a bug: it keeps re-searching and never reaches
    final_answer, standing in for a real agent stuck re-planning forever."""
    return ("search", "who invented the calculator")


def run_buggy_agent(max_steps=4):
    TRACE.clear()
    scratchpad = []
    with span("agent_run", max_steps=max_steps):
        for step in range(max_steps):
            with span("think", step=step):
                tool, arg = buggy_agent_policy(step, scratchpad)
            with span("tool_call", tool=tool, arg=str(arg)):
                scratchpad.append(search_tool(arg))
    return None, max_steps


print("\n--- a second run: a policy with a bug ---\n")
answer2, steps_used2 = run_buggy_agent(max_steps=4)
hit_cap = steps_used2 >= 4
print(f"answer: {answer2}")
print(f"steps used: {steps_used2}   hit the step cap without finishing: {hit_cap}\n")
print_trace()
Output
answer: 15% of 240 is 36.0. The calculator was invented by: Blaise Pascal built an early mechanical calculator in 1642.
steps used: 3

agent_run                       {'max_steps': 5}
  think                         {'step': 0}
  tool_call                     {'tool': 'calculator', 'arg': '240 * 0.15'}
  think                         {'step': 1}
  tool_call                     {'tool': 'search', 'arg': 'who invented the calculator'}
  think                         {'step': 2}
  final_answer                  {}

--- a second run: a policy with a bug ---

answer: None
steps used: 4   hit the step cap without finishing: True

agent_run                       {'max_steps': 4}
  think                         {'step': 0}
  tool_call                     {'tool': 'search', 'arg': 'who invented the calculator'}
  think                         {'step': 1}
  tool_call                     {'tool': 'search', 'arg': 'who invented the calculator'}
  think                         {'step': 2}
  tool_call                     {'tool': 'search', 'arg': 'who invented the calculator'}
  think                         {'step': 3}
  tool_call                     {'tool': 'search', 'arg': 'who invented the calculator'}

The content and step counts above are exact and reproducible, since every tool and policy here is deterministic. Notice the second trace: four identical tool_call: search steps in a row is precisely the pattern an observability dashboard should flag automatically.

Line-by-line walkthrough

fake_agent_policy and buggy_agent_policy both stand in for a real call to a model asking "what next?". Everything else in run_agent — the loop, the spans, the step cap — is real agent-orchestration logic, unchanged by whether the policy is real or fake.

The max_steps guard is what turns "stuck forever" into "stuck for at most max_steps, then stopped and flagged". Removing it would let the buggy policy loop without end.

The repeated tool_call: search rows in the second trace are the actual observable symptom of the bug — not a crash, not an error, only the same action taken again and again.

Common mistakes

Tracing only the final answer, not each tool call. As shown above, the second run's answer is only None. Without the full trace, there is no way to see why — only that it failed.

Setting no step limit at all. A real agent calling a real, billed tool in a loop is not only a bug. It is an uncapped bill, accumulating for as long as nobody notices.

Treating "hit the step cap" as a rare edge case. It should be tracked as its own rate, over time, the way an error rate is tracked. A rising rate of capped-out runs is an early warning worth alerting on.

Not logging which tool was called with which arguments. "The agent used a tool" is not debuggable. "The agent called search with arg='who invented the calculator', three times in a row" is.

Try it yourself

Add a check after the loop in run_agent that looks at the last three tool_call spans and flags a warning if all three used the identical tool and argument. Run it against both traces above, and confirm it only fires on the second one.

What to learn next

Researcher — Mathematics and papers.

Why agent observability is a strict superset of single-call tracing

A single LLM call has one span. An agent run has an unbounded, run-dependent number of spans, with control flow (branching, looping, early termination) that only exists at run time. This means agent observability tooling needs to represent a tree of arbitrary, dynamic shape per request, rather than a fixed pipeline of known steps — the OpenTelemetry GenAI conventions extend their span model with gen_ai.agent.* attributes and nested tool-call spans specifically to represent this.

Loop detection formally

The repeated-tool-call pattern in the developer example is a special case of a more general problem: detecting that an agent's trajectory has entered a cycle in its own state space. A simple, cheap detector: hash each (tool, arguments, and optionally result) tuple, and flag when the same hash appears more than $k$ times within a sliding window of the last $w$ steps — exactly what the "try it yourself" exercise above implements in miniature. More sophisticated approaches embed the full scratchpad state and flag high cosine similarity between the current state and a recent prior state, catching near-duplicate loops that vary superficially (slightly reworded queries) but are functionally identical.

Attribution in multi-agent systems

When multiple agents coordinate (a planner agent, a specialist agent, a critic agent), a failure's root cause can live in a different agent than the one that produced the visibly wrong output — a critic that approved a bad plan is the deeper cause, even though the specialist that executed it produced the visible error. Tracing needs to preserve the full multi-agent call graph, with each agent's spans attributed to their originating agent identity, for this kind of root-cause attribution to be possible after the fact at all.

Cost and latency compound differently in agents

For a single call, cost and latency are roughly one draw from the fixed distributions covered in token and cost telemetry. For an agent, both compound across an unknown number of steps, so the relevant monitored quantities shift from per-call averages to per-run totals and their tail: $P_{99}$ of total steps, total cost, and total latency per completed run, since a small fraction of runs looping near the step cap can dominate total cost even while barely moving a per-call average.

Papers

  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, ICLR 2023 — arxiv.org/abs/2210.03629, the reasoning-then-acting loop structure the developer example is a minimal instance of.
  • Wang et al., A Survey on Large Language Model based Autonomous Agents, 2024 — arxiv.org/abs/2308.11432, covers failure modes including loops and their detection more broadly.

What to learn next

What to learn next

These follow on from what you just read.

  • Incident Response for ML Systems

    On-call for ML systems

    On-call means someone is reachable and ready to act when a model-serving system breaks, on a rotating schedule, so a 2 a.m. failure never depends on luck.

  • Incident Response for ML Systems

    Runbooks for model failures

    A runbook is a written, step-by-step set of checks for one specific kind of failure, so the person who gets paged does not have to think from zero.

  • Incident Response for ML Systems

    Triaging an ML incident

    Triage is deciding how bad an incident is and what to do first, in the first few minutes, before anyone understands the full cause.