Microsoft Releases AX Playbook for Evaluating Coding Agents
The method links runnable tests and agent trajectories to changes in documentation, SDKs and extensions, with a companion skill limited to the playbook.
Edited by Tyronne Panaino
Microsoft released the Agent Experience Practitioner Playbook on October 9, presenting a repeatable method for teams that want AI coding agents to use their technology correctly. The official Microsoft Developer article says the method draws on hundreds of agent sessions and connects evaluation results to changes in the sources agents use, including documentation, tools and extensions.
The audience is broader than teams building coding agents themselves. Microsoft addresses people who maintain SDKs, APIs, services, command-line tools, MCP servers, skills, plugins and instruction sets. Their practical problem is that a coding agent can produce plausible code while selecting an outdated version, deprecated authentication pattern or unsuitable setup. The playbook reframes that behavior as something a product team can measure and improve.
From a score to a repair loop
Microsoft describes a workflow that runs from evaluation setup to a shipped fix. Its central distinction is between criteria that judge whether an answer is meaningful and executable gates that establish whether generated code actually runs. That matters because a high score can conceal a broken implementation when the check measures only whether a term or platform name appeared.
The method also separates the readout from the trajectory. A readout records what failed; the trajectory shows how the agent reached the result. Inspecting that path can distinguish an extension that never loaded from one that loaded but was not called, or one that was called and then applied incorrectly. Those failures point to different remedies, so a single aggregate score is not enough to decide what to change.
The operational loop is therefore: define a realistic task, write calibrated criteria, prove the result can run, inspect the agent's path, identify the responsible source surface, change that surface and rerun the evaluation. The value is not a universal benchmark rank. It is a repeatable way to test whether a documentation, SDK or extension change altered agent behavior in the intended direction.
What Microsoft says changed in practice
The announcement grounds the method in Microsoft product work rather than describing it only as theory. Microsoft says its evaluations contributed to dozens of shipped documentation and extension fixes, including 46 improvements to the Azure Cosmos DB Agent Kit. That figure is a first-party account, not an independently reproduced outcome, but it shows the kind of artifact the process is intended to produce: a concrete product change tied back to an observed agent failure.
Microsoft also says the method is not tied to one evaluation platform. A team can use an internal system or another tool if it supports the capabilities the playbook expects. This makes the framework potentially useful as a common process across products with different test infrastructure, while leaving each team responsible for the quality of its scenarios, criteria and runnable gates.
The companion skill and its limits
Alongside the document, Microsoft released an AX Practitioner skill intended to help users apply the method while working. The company says the skill answers from the playbook and reports an average score of 95% across 330 questions measured against the playbook's content. Those are vendor-reported evaluation results. The announcement does not publish the full question set, scoring rubric, error analysis, uncertainty range or an independent reproduction, so the figures should not be read as a general measure of agent accuracy.
The same evidence limit applies to the playbook's wider impact. Microsoft provides examples and a process, but the article does not establish how much the method improves correctness across unrelated products, teams or models. The next useful checkpoint will be public evaluations from adopters that disclose their tasks, executable gates, before-and-after results and the exact source changes they shipped.
Status
Learning. The playbook and companion skill are confirmed through Microsoft's official developer channel. Internal confidence is medium because the workflow, improvement count and skill score come from one first-party announcement without independent outcome validation.
Sources
Update note: Last reviewed 2026-10-09. We will revise this explainer if Microsoft publishes evaluation artifacts or independent adopters report reproducible results.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.