Microsoft Research Profiles AgentRx for Diagnosing AI Agent Failures
The research framework converts failed agent traces into evidence-backed constraints, but its public records disagree on benchmark size.
Edited by Tyronne Panaino
Microsoft Research listed AgentRx on October 2, 2026 as work for the 2026 Conference on Empirical Methods in Natural Language Processing. The framework is meant to help AI-agent developers trace a failed run to the step where it became unrecoverable, an increasingly practical problem for teams operating long, probabilistic workflows across tools and multiple agents.
The work matters to agent engineers, evaluators and reliability teams because a final failure rarely explains its own cause. AgentRx turns an execution trajectory into a sequence of constraints, checks those constraints step by step, and produces an evidence-linked validation log before an LLM judge assigns a critical step and failure category. The supported contribution is diagnostic: it does not make the underlying agent deterministic or prove that an automated diagnosis is always correct.
From an opaque trace to a reviewable diagnosis
The Microsoft Research publication page says the researchers manually annotated failed trajectories with a critical failure step and a category drawn from a cross-domain taxonomy. Those annotations give the framework a target for evaluating whether its automated attribution agrees with human-labelled failures.
AgentRx then synthesizes constraints from the task and trajectory, evaluates them in sequence, and records violations together with supporting evidence. An LLM-based judge uses that log to localize both the critical step and the failure category. That structure is important because it gives a reviewer more than a bare model verdict: the diagnostic path can be inspected against the trace.
The evaluated material spans structured API workflows, incident management and open-ended web or file tasks. Those domains cover several recurring agent patterns, including deterministic tool calls, operational response and less constrained computer work. They do not establish universal performance across every model, orchestration system or production environment.
The public records describe different benchmark versions
The sources do not present one stable benchmark count. Microsoft Research's summary says the release contains 115 failed trajectories, while the current arXiv v2 record says 170 trajectories across 11 task settings. ArXiv records the first submission on February 2, a revision on August 31, and acceptance to EMNLP Findings 2026. The Microsoft Research archive page dates its listing to October 2.
The most cautious reading is that the Microsoft page and arXiv abstract reflect different research versions. That is an inference from the conflicting counts, not an explanation supplied by either source. Readers comparing results should therefore identify the version they are using instead of combining the 115- and 170-trajectory descriptions.
The arXiv abstract reports an average 75% improvement in step localization over prior work. That figure is an author-reported benchmark result, not an independent reproduction. The abstract also does not show whether the same gain will hold for a team's private tools, permissions, prompts, model mix or failure distribution.
What engineering teams can take from the work
AgentRx offers a useful design pattern even before broader replication: separate failure diagnosis into explicit constraints, sequential checks, evidence capture and a final attribution decision. That makes it easier to ask whether the claimed critical step is supported by the trace and whether a category is actionable for testing or remediation.
The next verifiable checkpoint is a reconciled public benchmark version, followed by independent evaluation on unseen production trajectories. Useful evidence would include error analysis for cases where the judge selects the wrong step, sensitivity to different underlying models, and the cost of generating and checking constraints on long traces.
Status
Learning. Internal confidence is medium because Microsoft Research and the authors' arXiv record directly document the method, but they describe the same work, disagree on the benchmark size and do not provide independent production validation.
Sources
- Microsoft Research archive — AgentRx listing dated October 2
- Microsoft Research — AgentRx publication page
- arXiv — AgentRx v2 abstract and submission history
Update note: Last reviewed 2026-10-06. We will revise this post if the public benchmark records are reconciled or independent evaluations materially change the evidence.
Sources
- Microsoft Research archive — AgentRx listing — official
- Microsoft Research — AgentRx publication page — official
- arXiv — AgentRx v2 — research
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.