In Part 2 we built the observability foundation — the full execution trace, session replay, step-level instrumentation. Now we use it. Anomaly detection in agent systems means catching failures before users surface them. That requires understanding the two distinct ways agents fail.
Intra-Agent Failures
Intra-agent failures occur within a single agent's reasoning loop. They're the failures most teams think about first, and they come in several forms:
Hallucination under tool failure. The most reliable signal: a tool invocation with a thin or null result, immediately followed by a confident completion. The agent received nothing useful and kept going anyway. This pattern is detectable in the trace — if you're capturing tool results alongside completions.
Reasoning loops. An agent cycling through the same decision or tool call without making progress. Detectable via repeated tool invocations with identical or near-identical arguments within a single session.
Task specification drift. The agent's working interpretation of its objective has diverged from the original task — usually through accumulated context that gradually overwrites the original framing.
Prompt injection. External content — from retrieved documents, tool outputs, or user inputs — that redirects agent behavior. Requires both trace coverage and content inspection to catch reliably.
Inter-Agent Failures
Inter-agent failures emerge from interactions between agents rather than individual breakdowns. They're harder to detect because no single agent misbehaves — the failure lives in the handoff.
Stale context propagation. A downstream agent receives context that was valid when an upstream agent generated it but has since been superseded. The downstream agent has no way to know this without explicit provenance metadata on the context itself.
Semantic drift across handoffs. Each agent in a pipeline slightly reinterprets what it receives. Small reinterpretations compound. By the time output reaches the final agent, the original intent may be unrecognizable — but no individual step looked obviously wrong.
Communication anomalies. Message storms, deadlocks, and coordination failures in orchestrator/subagent architectures. These look like infrastructure failures but originate in agent logic.
Three Detection Layers
Effective anomaly detection in production agent systems uses three layers, each catching what the others miss:
Rule-based detection is fastest. Explicit checks for known failure signatures — null tool results followed by confident completions, loop detection via repeated calls, threshold violations on latency or token cost. High precision, low recall. Catches the obvious failures immediately.
Statistical baselines catch what rules miss. Once you have historical run data, you can establish what normal behavior looks like for your specific agent and flag deviations — unusual tool invocation sequences, completion length outliers, latency spikes that don't correlate with input complexity.
LLM-as-judge evaluation handles semantic failures that neither rules nor statistics can characterize. Run asynchronously after completion — not in the hot path — scoring outputs against rubrics that capture what correct agent behavior actually looks like for your use case.
The Gap Nobody Talks About
Most frameworks treat agent memory as opaque. The execution trace shows you what the agent did. The completion shows you what it said. But the belief state at the time of failure — what the agent actually knew and believed when it made the wrong call — is the most diagnostic signal, and most architectures don't preserve it.
You can detect that a failure occurred. Understanding why it occurred requires access to the memory state that was driving the agent's decisions at the moment things went wrong. Without that, anomaly detection tells you there's a problem. It can't tell you what caused it.
That's root cause analysis — and that's Part 4.