The News, Explained
Microsoft released the Agent Experience (AX) Practitioner Playbook on October 9, 2026. AX describes how an AI coding agent discovers, selects, and correctly uses a particular SDK, API, or service. The method starts from a practical problem: an agent can produce plausible code that compiles while choosing an old SDK version or a deprecated authentication pattern. Source
The playbook focuses on surfaces a technology team can change rather than retraining the model. Those surfaces include documentation, MCP servers, skills, plugins, instructions, command-line tools, and APIs. MCP is a standard interface for connecting an AI system to external data and tools. The method is intended for people who build SDKs, APIs, services, and agent extensions, as well as people who document or advocate for them, and it does not depend on one evaluation system. Source
The workflow begins with a realistic task scenario and pairs success criteria with execution gates. Meaning-based criteria can check whether the agent chose the right SDK and recommended pattern, while an execution gate proves that the generated project builds and runs. Microsoft describes evaluations that gave perfect scores to code that never compiled and checks for using a platform that passed whether or not the platform was actually used. The playbook therefore calls for calibrating criteria before trusting them, versioning them as a product changes, and avoiding the trap of handing all criterion writing to the model being evaluated. Source
The next step examines the agent’s trajectory—the record of which information and tools it selected—instead of looking only at the final answer. An extension that never loaded, one that loaded but was not called, and one that was called but applied incorrectly may look like the same failure in the output, but they require different fixes. The playbook uses nine recurring failure patterns to narrow the cause, test a documentation or extension change as a hypothesis, and bring the resulting evidence to the team that owns the source. Source
Microsoft says evaluations on technologies including Cosmos DB and SharePoint Framework produced multiple shipped fixes, including 46 improvements to the Azure Cosmos DB Agent Kit. It also released an AX Practitioner skill for learning and applying the method in conversation. The skill is designed to answer from the playbook and ask before using another source when a question falls outside it. Microsoft reports a 95% average score across 330 questions about the playbook. That vendor-reported figure measures alignment with the playbook, not a 95% code-accuracy rate across technologies. Source
OYOPICK’s Take
The useful shift is from treating every coding-agent mistake as a vague model problem to separating failures in documentation and tools that a product team can change and measure. Checking both the output and the execution path makes repeated failures easier to diagnose. A practical extension would be to run representative agent tasks as regression tests alongside human documentation review whenever a product ships a significant change.
Agent scores should not become the sole objective for documentation. Instructions that are difficult for people to read or tools with overly broad permissions could raise a narrow score while weakening maintainability and safety. Evaluation criteria should cover recommended versions and execution success together with least privilege, secure patterns, and error messages that people can understand.
As AI becomes more capable with development tools, we hope human expertise expands into designing evaluation criteria and review paths rather than being pushed aside. Microsoft’s examples and skill score are a starting point. Each team still needs to calibrate the method on representative work in its own environment and keep human approval for high-impact changes.