🔍 Read the full analysis: What OpenAI’s Agent Training Could Mean For Software Users And Ironclad on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI says it trained GPT-6 Astra in hosted copies of Ironclad’s contract-management software using 11 legal, commercial and procurement tasks. Astra met an average 55% of each task’s rubric criteria; its estimated completion times were simulated, and the results do not establish customer productivity gains or readiness for unsupervised use.
OpenAI published results on October 6 from training its frontier model GPT-6 Astra in hosted copies of Ironclad’s contract-management software, using 11 tasks based on legal, commercial and procurement workflows. The model met an average 55% of the tasks’ evaluation criteria, according to OpenAI, while its reported time estimates were simulated rather than measured customer savings. The work offers an early look at how software vendors might help train agents for specialized business tasks, but it does not show that the agent is ready to handle contract work without human review.
Ironclad is a contract-management software company, not the name of a new agent framework. OpenAI’s post, titled “Advancing computer use with Ironclad,” describes a collaboration in which Ironclad staff and OpenAI employees who use the product selected 11 tasks. Examples included setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task.
OpenAI says each task was graded against a rubric of 8 to 50 criteria, depending on complexity. The reported 55% result is the average share of criteria met, not the percentage of tasks completed successfully. OpenAI compared GPT-6 Astra, run at a “max” setting, with GPT-5.6 Sol at a “high” setting: the models met an average 55.0% and 41.6% of criteria, respectively. An internal model used during Astra’s development scored 63.7%. On one example task, Astra met about 94% of the criteria, but that single result is not the overall score.
OpenAI reported estimated times of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol. Its post says those figures are simulated estimates based on assumed processing and generation speeds, not measured time savings for customers. The estimates cover the 11 research tasks; they do not establish how the model performs across Ironclad’s full product or in routine customer operations.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Accuracy Matters
The findings matter because business software workflows can depend on several rules being followed together. OpenAI’s example describes a procurement process that may require Finance approval above a spending threshold, Security review for certain requests and Legal review when terms are nonstandard. Meeting two requirements while missing a third is not necessarily a useful partial result: it could send a purchase forward without a required control.
That makes the 55% average rubric score a measure of progress in a controlled test, not evidence that Astra can safely run contract processes on its own. OpenAI’s post says human oversight remains necessary, and Ironclad’s CTO, Sunita Verma, emphasized preserving the controls teams rely on. For software buyers, the immediate relevance is practical: agent demonstrations and aggregate scores do not replace checking which specific rules were met, which were missed, and who must review the result.
For vendors, the collaboration suggests a possible route to make agents more capable in specialized products: provide real workflows, expert knowledge, a secure environment and research data that can be used safely. There is also a strategic question. If customers increasingly give instructions through an agent rather than a product’s screens, a vendor’s lasting value may depend more on its business rules, data, audit records and controls than on its interface. That is an implication of the approach, not a stated outcome of this test.
contract management software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Ironclad Entered the Test
OpenAI says Ironclad provided hosted copies of its product for model practice. For synthetic training tasks, OpenAI says it used publicly filed contracts from the U.S. Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company also says the work used no OpenAI customer data, no OpenAI internal contracts and no non-public Ironclad customer data.
The project is presented as research into whether models can learn a company’s business rules, complete multi-step work inside specialized software, and check completed work against the original requirements. OpenAI describes GPT-6 Astra as the first frontier model trained in this way. In the post’s final section, it invites a small number of software companies to discuss similar partnerships. It asks potential partners to bring concrete examples of tasks agents cannot reliably complete, people with deep knowledge of the work, a secure test environment and data suitable for research.
The post’s invitation makes the work relevant beyond contract management, but the published measurements are limited to the 11 selected Ironclad tasks. They do not show how an agent would perform across other business applications, organizations or live customer workflows.
As an affiliate, we earn on qualifying purchases.
What the Test Does Not Establish
The results do not show whether GPT-6 Astra can reliably complete all requirements in a workflow, or how often it misses particular high-impact rules. OpenAI’s average rubric score combines criteria across tasks, and the published summary does not establish that every task passed a threshold appropriate for use in live contracting. A high score on one showcase task does not resolve that gap.
It is also unclear whether the estimated times would translate into real time savings after users check outputs, correct errors or rerun tasks. OpenAI explicitly labels the figures as simulations, so they should not be treated as measured productivity gains. The source does not establish deployment plans, customer adoption, commercial terms for possible research partners, or whether other vendors will participate.
OpenAI says the work did not use non-public Ironclad customer data, but the available account does not provide a full technical description of the data handling, security controls or evaluation process. Those details would be needed to assess the research methods and how its safeguards might apply in other settings.
AI-powered contract drafting tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
OpenAI’s Proposed Partner Research
OpenAI says it is seeking a small number of software-company partners to identify tasks current agents cannot reliably handle. The company says candidates should provide a specific failing example, subject-matter experts, a secure environment and data that can safely support research. The post does not give a timetable for selecting partners or publishing further results.
For any later evaluation to clarify readiness, useful details would include task-by-task results, the types of criteria missed, how often human reviewers had to intervene, and performance in realistic workflows. Until such evidence is available, the Ironclad study is best understood as a bounded research test rather than a demonstration of verified customer savings or an agent that can manage contracts without supervision.
business workflow automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They evaluated GPT-6 Astra on 11 selected legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.
Does the 55% score mean Astra completed 55% of the tasks?
No. OpenAI says 55% is the average share of evaluation criteria met across the tasks, not the share of tasks completed successfully.
Did Astra save customers time?
The published time figures are simulated estimates, not measured customer results. OpenAI says they are based on assumed processing and generation speeds and apply to the 11 research tasks.
Can companies use Astra to manage contracts without human review?
The reported results do not establish that. The average score leaves criteria unmet, and OpenAI’s post says human oversight still matters for workflows that must preserve business rules and controls.
What data did OpenAI say it used?
OpenAI says it created synthetic training tasks from publicly filed SEC EDGAR contracts filtered to remove personal information. It says the project did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
