AgentOps Series  PART 2 OF 6  — 

AgentOps Part 2: You Can't Fix What You Can't See — Tracing Agent Runs End to End

Tracing agent runs means capturing not just what your system did, but what it decided and why. Here's what a complete agent trace actually contains — and where to start.

The Three Pillars Don't Change. The Target Does.

Traditional observability rests on three pillars: metrics, logs, and traces. Those pillars still apply to agent systems. What changes is the target.

In a standard web service, a trace captures what the system did: which services were called, in what order, at what latency. That's sufficient when the system's behavior is deterministic and its logic is in the code.

Agent systems like LangGraph and CrewAI operate non-linearly — through nodes, conditions, and dynamic state management. Each decision point is a potential failure surface. A trace that only captures execution sequence misses the most important signal: what did the agent decide at each branch, and why?

Without that, you can see that something failed. You cannot understand it.

What a Complete Agent Trace Actually Contains

A trace sufficient for meaningful debugging in production agent systems has five distinct layers:

LLM calls. Every prompt sent and completion received — including token counts, latency, cost, and the model version. When a reasoning failure occurs, the prompt is almost always part of the evidence.

Tool invocations. Every external call — arguments passed, results returned, failures encountered. An agent that receives an empty or ambiguous tool result and continues confidently is a hallucination waiting to happen. The trace needs to capture both sides of that exchange.

Reasoning steps. The intermediate planning and decision-making between observable actions. This is the hardest layer to capture because most frameworks don't preserve it automatically. But it's the most diagnostic signal when things go wrong.

Memory and state snapshots. What did the agent know at each point in the run? What was in working memory when it made the call that started the failure chain? Without this, post-hoc analysis is mostly inference.

Inter-agent boundaries. In multi-agent systems, every handoff is a fidelity risk. Context that was accurate in one agent's frame may be stale, truncated, or misinterpreted by the next. Those boundaries need to be first-class instrumentation targets.

Session Replay Is the Actual Goal

Anomaly detection and alerting are table stakes. The real objective of agent observability is session replay: a complete, ordered record that lets you reconstruct an agent run step by step and identify exactly where divergence occurred.

This isn't just a debugging convenience. It's the prerequisite for everything else in the AgentOps stack — anomaly detection, root cause analysis, feedback loops. None of those work if you can't replay what actually happened.

Tools like LangSmith, AgentOps, and Langfuse provide session replay capability natively. If you're building on a framework that doesn't, that's the gap to close first.

Where to Start

Instrument in priority order: tool calls first, LLM calls second, state snapshots third, inter-agent boundaries last. Tool calls are where the most consequential failures happen and where the evidence is most recoverable. Inter-agent boundaries require coordination across your full system — save that complexity for after the simpler instrumentation is in place and producing signal.

The Real Point

Observability is the prerequisite for every other capability in the AgentOps stack. You cannot detect anomalies in behavior you haven't captured. You cannot do root cause analysis on a run you can't replay. You cannot close a feedback loop against failures you can't characterize.

The teams that will distinguish themselves in production agentic AI won't necessarily have the best models. They'll have the best visibility into what their models are actually doing — and the infrastructure to act on that visibility systematically.