How Microsoft's run-assert-eval Tests Agent Controls Against the Same Failures
The open-source skill links threat discovery, generated evaluations and runtime policy, then reruns the unchanged test to measure the control.
Edited by Tyronne Panaino
Microsoft introduced run-assert-eval on September 24, 2026 as an open-source skill for teams building AI agents in VS Code. The workflow starts with an agent description, discovers relevant failure modes, turns selected risks into evaluations, drafts runtime policy and reruns the evaluation against the governed agent.
The practical change is not another standalone safety metric. Microsoft is connecting four jobs that teams often perform separately: threat discovery with Clarity, behavior definition and testing with ASSERT, policy generation through the Agent Control Specification, and a controlled comparison of the baseline and governed agent. That makes the release relevant to developers and risk teams that need evidence a control addressed the failure it was designed to stop.
What the workflow changes
Clarity sits at the start of the loop. It threat-models the agent and produces candidate failure modes before the team assumes its written requirements are complete. A person reviews that list and selects which risks should move into measurement. That review boundary matters because the skill is designed to surface more risks than a team will necessarily evaluate in one pass.
ASSERT then converts each selected risk into a narrow behavior and a test suite. Microsoft says the workflow can research prior evaluation methods to propose useful dimensions for the test set. In the company's billing-support example, those dimensions separate different access paths and different ways a user might ask for another account's information. The purpose is to show where a control fails, not merely whether any failure appeared somewhere in a mixed set.
Once a failure is measured, the skill can generate and validate a draft Agent Control Specification policy. The source describes Rego rules paired with a manifest that identifies where in the agent lifecycle enforcement should occur. Policy generation is not approval: Microsoft explicitly keeps review of the policy, manifest, intervention point and target wiring with people before a governed run begins.
Why the second run is the important part
The comparison reuses the same behavior definition, test cases and judge for the governed run. Only the intended policy intervention changes. That is a stronger experimental shape than regenerating the evaluation after a fix, because a different score is less likely to be explained by a different test or judging setup.
In Microsoft's worked example, account-scoping rules were placed before and after tool calls. The pre-tool control could deny a request before a tool reached another customer's record; the post-tool control could withhold a result that should not enter the model context. The example uses deterministic account identifiers rather than asking another model to decide whether a request looks suspicious.
Microsoft reports that its billing-support example reduced impermissible behavior across the tested splits while also reducing permissible-behavior violations. Those results remain a vendor-run demonstration, not an independent benchmark or a guarantee that another agent, judge or control will behave the same way. The useful product delta is the repeatable loop and preserved comparison, not the headline score from one example.
What teams still have to prove
The skill ships in the ASSERT repository, and Microsoft says ASSERT and ACS are available under the MIT license. The release includes worked domains and risk suites, but the announcement does not establish coverage for every framework, tool boundary or production environment.
Teams still need to validate whether Clarity finds the risks that matter in their own system, whether the generated behavior definitions are precise, whether the policy is enforced at the correct interception point, and whether the judge agrees with qualified human review. They also need to test operational concerns outside the worked example, including tool failures, policy bypasses, latency and the handling of state that crosses sessions.
The next useful evidence will be independent replications across different agent stacks and published comparisons between automated judging and domain-expert review. For now, run-assert-eval is best understood as an available workflow for producing a traceable find-measure-control-retest cycle, with human approval retained at the consequential policy step.
Status
Learning. Internal confidence is medium because one official Microsoft article documents the release and worked example, while the reported results have not been independently reproduced in this run.
Sources
Update note: Last reviewed 2026-09-28. We will revise this post if Microsoft changes availability, publishes broader evaluations or independent testing challenges the stated workflow.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.