In Parts 1 through 3, we built up the foundation — AgentOps as a discipline, observability as the prerequisite, anomaly detection as the first line of defense. All of that work serves one goal: being able to answer the question that actually matters when something goes wrong in production.
Why did the agent do that?
Not what happened. Not when. Why — specifically, what the agent believed, what it knew, and what it was reasoning about at the exact moment it made the wrong call. That's root cause analysis in agentic systems. And it's a fundamentally different problem than anything MLOps or LLMOps had to solve.
Why Traditional RCA Breaks Down
RCA is straightforward in traditional software. A service fails, you pull the logs, trace the request, find the line that threw the exception. Deterministic systems leave deterministic evidence. The failure is reproducible. The cause is isolatable.
Agents don't work that way.
Recent studies put failure rates in multi-agent systems as high as 86.7% (Cemri et al., MAST, 2025 — published at NeurIPS 2025) — errors propagating through intricate dependencies in ways that are genuinely hard to unwind. The failures aren't point failures. They're trajectory failures. An orchestrator marks a constraint at step 14. A downstream agent overlooks it at step 16. The failure is distributed across agents and steps, not localized to a single error event. You can't find the root cause by looking at the line that failed. You have to reconstruct the entire reasoning chain leading up to it.
That's a different class of problem.
Three Layers. Most Teams Are Only Working One.
RCA in agentic systems operates at three distinct layers.
Execution layer. The trace — tool calls, LLM completions, state transitions, sequence. This is where most RCA efforts live. It's necessary. It's not sufficient. Knowing the agent called a tool at step seven that returned empty doesn't tell you why the agent continued confidently from that point rather than stopping.
Reasoning layer. The intermediate planning and decision-making between observable steps. Agent execution histories contain rich information — not just what happened, but why agents made decisions and how they reasoned through them. The problem is most frameworks don't preserve this. The reasoning that led to a bad tool selection lives in intermediate completions that get discarded after the next step executes. By the time the failure surfaces, the evidence is already gone.
Memory and belief layer. What did the agent know when it made the call that started the failure chain? What was in its working memory? What beliefs had it formed from prior steps that turned out to be wrong? A fundamental limitation runs through most agent architectures: agents have amnesia. Most LLMs are stateless, and agents lack systematic mechanisms to record their execution experience in a way that's actually inspectable after the fact. You can't do meaningful RCA on a system whose memory state you can't examine.
Most teams are working the execution layer only. The reasoning layer gets captured inconsistently. The memory and belief layer is almost never instrumented at all.
Causal Attribution — Still Unsolved
The research frontier here is causal attribution — not just where a failure occurred in the trajectory, but which agent, which decision, and which state transition was causally responsible for what happened downstream.
Direct LLM-based approaches struggle to capture fine-grained causal clues within lengthy execution contexts. Spectrum-based methods hit prohibitive token costs through repeated trajectory replays. Fine-tuned models introduce training overhead and don't generalize across agent designs. None of these are production-viable today.
The most interesting recent direction is Project Ariadne — a structural causal model approach that uses counterfactual interventions to audit whether reasoning traces in LLM agents are faithful generative drivers or post-hoc rationalizations. That distinction matters for RCA: if the reasoning trace you're analyzing wasn't actually driving the agent's decisions — if it was generated after the fact to justify an action the model took for other reasons — then you're analyzing a rationalization, not the actual cause. Your entire RCA is built on fabricated evidence.
That's a practical problem, not an academic one. Anyone doing serious RCA on LLM-based agents should be thinking about it.
What's Actually Working in Production
Teams doing RCA well aren't waiting for the research to catch up. The approach is layered.
Structured trajectory logging. Every intermediate output gets preserved — planning steps, intermediate reasoning, rejected tool options. If it influenced the next decision, it needs to be in the record. This is an instrumentation discipline problem before it's a tooling problem. You have to commit to capturing it before a failure occurs.
Constraint validation at decision points. Instrument your agent to record which constraints were active at each decision point. When a failure occurs, you can immediately check whether the constraint was present, whether the agent acknowledged it, and whether it was violated explicitly or drifted past implicitly. Reactive approaches that discover constraint gaps mid-trajectory consistently underperform.
Causal graph construction. Build a dependency graph of the agent run rather than reading a flat log sequentially. Which outputs fed which inputs. Which decisions were contingent on which prior results. The timeline tells you sequence. The graph tells you causality. Failures in multi-agent systems almost always have a clear causal path in the dependency graph that's invisible in the execution timeline.
Historical pattern matching. The most effective RCA approaches don't analyze failures in isolation — they retrieve against a structured record of historical postmortems, asking whether this pattern has been seen before. A failure that looks novel in isolation often matches a known failure mode when compared against a well-maintained history. Tribal knowledge, operationalized.
The Memory Gap — Again
Every approach above depends on something most agent architectures don't provide: a structured, queryable, auditable record of what the agent knew and believed at each point in its execution.
Not the execution trace. The belief state — typed, versioned memory that was influencing the agent's decisions at the moment things went wrong.
Current agent memory systems consistently fall short on the dimensions that matter most for diagnosis: temporal and causal dynamics, selective forgetting, long-range understanding across extended interactions. The memory infrastructure most agents are running on wasn't designed for the kind of inspection that serious RCA requires.
OTel gives you the execution trace. Structured logging gives you the reasoning artifacts. Neither gives you typed, versioned, governed memory state with policy attached. Without it, RCA in high-stakes agent systems is still mostly manual reconstruction — forensic work performed after the fact by engineers inferring what the agent believed from indirect evidence.
That's not an operational model that scales. And it's the gap the next generation of agent infrastructure needs to close.
Next: the feedback loop — how you take everything built across this series and turn it into a system that actually gets better over time.