Learn learning medium confidence

OpenAI Tests GPT-6 Astra on Ironclad Contracting Workflows

The research evaluation uses detailed criteria across 11 simulated tasks, but its vendor-reported scores and simulated times do not establish customer productivity or general legal reliability.

Edited by Tyronne Panaino

OpenAI published a research evaluation on October 6 that tests GPT-6 Astra on 11 contracting tasks developed with Ironclad. The results show Astra scoring higher than GPT-5.6 Sol in OpenAI's internal setup, but the source explicitly limits the finding to those tasks and says the reported times are simulations rather than measured customer savings.

The OpenAI publication on the Ironclad collaboration frames the work as training and evaluation for computer-use agents in specialized professional software. Ironclad is the first named partner in a program that OpenAI says will work with a small number of software companies on difficult business workflows.

The evaluation turns workflows into scored tasks

OpenAI and Ironclad identified 11 tasks spanning legal, commercial and procurement work. Examples in the source include configuring nondisclosure agreements, building procurement approval processes and updating a reusable legal clause based on a selected jurisdiction. OpenAI estimated that an experienced user would take an average of about 30 to 40 minutes per task.

Each task was checked against between 8 and 50 criteria, depending on complexity. That design matters because completing one visible action is not enough for a multi-step workflow: the final configuration has to satisfy the collection of rules represented in the scoring rubric. Ironclad provided hosted product environments, while OpenAI created synthetic training tasks and used reinforcement learning to improve model performance through practice and feedback.

The source says those synthetic tasks were built from contracts publicly available through the SEC's EDGAR database after filters intended to remove personal information. It also says OpenAI customer data, OpenAI's internal contracts and nonpublic Ironclad customer contracts were not used for training or evaluation. Those statements describe the vendor's reported data boundary; this run did not independently audit the filtering process.

Astra scored higher under different reasoning settings

OpenAI compared GPT-6 Astra using Max reasoning with GPT-5.6 Sol using High reasoning, the setting where each model scored highest. Across the 11 tasks, Astra recorded a mean rubric score of 55.0%, compared with 41.6% for Sol. The estimated average time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol.

Those figures describe a specific internal research evaluation, not a broad benchmark for contracting or legal work. The models used different reasoning settings, the task creators worked with the software vendor and the evaluation covered one hosted product environment. The comparison may be useful for tracking progress inside this test, but it does not establish how either model performs across other contracting platforms, organizations or legal systems.

OpenAI also reports that an internal model used during Astra's development reached 63.7% on the tasks. That model is not the released comparison target, so the figure is best read as a research direction rather than a capability available to customers.

Simulated time is not measured productivity

The source's footnote says the 11 tasks do not cover all Ironclad workflows. It also says the times are simulated estimates based on assumed processing and generation speeds, not measured customer time savings. That distinction prevents the 19.2-minute figure from being treated as evidence that a legal or procurement team will save a particular amount of time.

A production assessment would also need to examine failure severity, permissions, exception handling, audit trails and human review. The published mean score shows that Astra still missed a substantial share of the scored criteria. A higher average is evidence of relative progress inside the test, not evidence that an agent can safely operate without oversight.

What the collaboration demonstrates

The useful lesson is methodological. A professional-software evaluation becomes more informative when a task has an environment, a realistic sequence and explicit criteria for the completed state. That is more specific than judging whether a model produced a plausible explanation of what it would do.

At the same time, the evidence remains first-party. This run found no independent reproduction of the tasks, scores, timing estimates or data filters. The next strong checkpoint would be an external evaluation with published task artifacts, identical reasoning conditions, failure analysis and measured results from representative users.

Status

Learning. The official OpenAI publication supports the evaluation design and reported results, while the narrow task set, simulated timing and lack of independent reproduction keep internal confidence at medium.

Sources

Update note: Last reviewed 2026-10-07. We will revise this post if the task set, evaluation artifacts, independent reproductions or measured customer outcomes become available.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Learn coverage