Skip to content
Huggehub Global Digital StudioParis --:--London --:--New York --:--Now accepting new projects →AI / Commerce / Technology / GrowthEurope • United Kingdom • United States

Engineering · 7 min read ·

LLM Observability: The Metrics to Track for AI Agents

Traditional monitoring tells you the server is up. LLM observability tells you whether the agent actually did the job — and what it cost.

By Hatim El Badaoui

A central black sphere connected to twelve glass and metal nodes

An AI agent can return HTTP 200 on every request and still fail half its tasks. It can pick the wrong tool, invent an order number or quietly cost ten times more than last week. LLM observability is the practice of tracing what models and agents actually do, so you can measure quality, cost and reliability — not just uptime.

Start with traces, not dashboards

The unit of LLM observability is the trace: one run of an agent, broken into spans for each step.

  • Input — the user request and the prompt actually sent.
  • Model call — model name, parameters, tokens in and out, latency.
  • Tool calls — which tool, which arguments, success or error.
  • Decision — what the agent chose to do next and why.
  • Result — the final output and whether it was accepted.

With complete traces you can answer almost any question later. Without them, every incident becomes guesswork.

The LLM observability metrics that matter

MetricWhat it tells you
Task success rateShare of runs that achieved the goal — the metric the business cares about.
Latency (p50 / p95)Time per run and time to first token; p95 shows the experience of unlucky users.
Cost per runTokens × price across all model calls, plus paid tool calls.
Tool failure rateErrors, timeouts and invalid arguments per tool — often the real cause of "the AI is bad".
Hallucination / groundednessHow often outputs contain claims not supported by the provided data.
Escalation rateHow often the agent hands over to a human — too low can be as worrying as too high.
Steps per runRising step counts signal loops or confused planning.

Break every metric down per agent, per tool and per model version. An average across everything hides the one tool that fails 40% of the time.

Measuring quality: evals

Success rate needs a definition of success. Three practical options, from cheapest to most reliable:

  1. Rules — output is valid JSON, cites a source, stays under a length, uses an allowed tool.
  2. LLM-as-judge — a second model grades outputs against a rubric; calibrate it against human labels.
  3. Human review — a weekly sample of traces scored by the people who own the process.

Keep a fixed set of test cases and run them on every prompt or model change. A regression you catch in CI costs nothing; one you catch in production costs customers.

Open-source LLM observability

Hosted platforms (Datadog LLM Observability, LangSmith, Langfuse Cloud, Arize) are worth it at scale. But the core is simple enough to own. Our agent-observability library collects per-run traces — prompt, model, tool, decision, result — and computes success rate, average latency, average cost, tool-failure rate and hallucination rate, with per-agent breakdowns. It is pure Python standard library, so it runs inside any agent, including serverless ones.

Whatever you use, align span names and attributes with the OpenTelemetry GenAI semantic conventions, so you can switch vendors later without re-instrumenting.

Alerts worth setting

  • Task success rate drops more than a set margin day over day.
  • Cost per run doubles against its seven-day average.
  • Any tool's failure rate crosses a threshold.
  • Median steps per run increase — a sign of loops.

Observability is also security

The same traces show denied actions, unusual tool sequences and prompt-injection attempts. Pair observability with the controls in AI agent guardrails and you have both the brakes and the dashboard.

Need agents you can measure? Our AI agent development work ships with tracing and evals from the first version.

Frequently asked questions

What is LLM observability?

Tracing and measuring what language models and agents do — prompts, model calls, tool calls, decisions, results, cost and quality — so you can debug, evaluate and improve them.

How is it different from normal application monitoring?

Monitoring checks that services are up and fast. LLM observability checks whether the AI did the right thing, how much it cost and why it failed when it did.

How do you measure hallucinations?

Compare outputs with the source data the agent was given, using rules, an LLM judge calibrated on human labels, or human review of sampled traces.

Let's build what's next

Let's build what's next.

Have a project, product or ambitious idea? Tell us where you want to go.