What OpenAI’s Agent Training Could Mean For Software Users And Ironclad
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What OpenAI’s Agent Training Could Mean For Software Users And Ironclad on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get hardware and tech essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra in hosted copies of Ironclad’s contract-management software using 11 legal, commercial and procurement tasks. Astra met an average 55% of each task’s rubric criteria; its estimated completion times were simulated, and the results do not establish customer productivity gains or readiness for unsupervised use.

OpenAI published results on October 6 from training its frontier model GPT-6 Astra in hosted copies of Ironclad’s contract-management software, using 11 tasks based on legal, commercial and procurement workflows. The model met an average 55% of the tasks’ evaluation criteria, according to OpenAI, while its reported time estimates were simulated rather than measured customer savings. The work offers an early look at how software vendors might help train agents for specialized business tasks, but it does not show that the agent is ready to handle contract work without human review.

Ironclad is a contract-management software company, not the name of a new agent framework. OpenAI’s post, titled “Advancing computer use with Ironclad,” describes a collaboration in which Ironclad staff and OpenAI employees who use the product selected 11 tasks. Examples included setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task.

OpenAI says each task was graded against a rubric of 8 to 50 criteria, depending on complexity. The reported 55% result is the average share of criteria met, not the percentage of tasks completed successfully. OpenAI compared GPT-6 Astra, run at a “max” setting, with GPT-5.6 Sol at a “high” setting: the models met an average 55.0% and 41.6% of criteria, respectively. An internal model used during Astra’s development scored 63.7%. On one example task, Astra met about 94% of the criteria, but that single result is not the overall score.

OpenAI reported estimated times of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol. Its post says those figures are simulated estimates based on assumed processing and generation speeds, not measured time savings for customers. The estimates cover the 11 research tasks; they do not establish how the model performs across Ironclad’s full product or in routine customer operations.

At a glance
reportWhen: Published October 6; further partner re…
The developmentOpenAI published results from training and evaluating GPT-6 Astra on selected Ironclad contract-management workflows and invited other software companies to explore similar research partnerships.
Crypto market snapshot
Fear & Greed Index
64/100 — Greed
Bitcoin BTC$81,695▼ 1.9%
Ethereum ETH$2,461▼ 4.1%
Tether USDT$0.9993▼ 0.0%
BNB BNB$730.41▼ 5.3%
XRP XRP$1.37▼ 3.4%
USDC USDC$0.9996▼ 0.0%
Solana SOL$109.14▼ 5.7%
TRON TRX$0.3329▼ 0.7%
Live data · CoinGecko · alternative.me (24h change)
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Accuracy Matters

The findings matter because business software workflows can depend on several rules being followed together. OpenAI’s example describes a procurement process that may require Finance approval above a spending threshold, Security review for certain requests and Legal review when terms are nonstandard. Meeting two requirements while missing a third is not necessarily a useful partial result: it could send a purchase forward without a required control.

That makes the 55% average rubric score a measure of progress in a controlled test, not evidence that Astra can safely run contract processes on its own. OpenAI’s post says human oversight remains necessary, and Ironclad’s CTO, Sunita Verma, emphasized preserving the controls teams rely on. For software buyers, the immediate relevance is practical: agent demonstrations and aggregate scores do not replace checking which specific rules were met, which were missed, and who must review the result.

For vendors, the collaboration suggests a possible route to make agents more capable in specialized products: provide real workflows, expert knowledge, a secure environment and research data that can be used safely. There is also a strategic question. If customers increasingly give instructions through an agent rather than a product’s screens, a vendor’s lasting value may depend more on its business rules, data, audit records and controls than on its interface. That is an implication of the approach, not a stated outcome of this test.

Amazon

contract management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Entered the Test

OpenAI says Ironclad provided hosted copies of its product for model practice. For synthetic training tasks, OpenAI says it used publicly filed contracts from the U.S. Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company also says the work used no OpenAI customer data, no OpenAI internal contracts and no non-public Ironclad customer data.

The project is presented as research into whether models can learn a company’s business rules, complete multi-step work inside specialized software, and check completed work against the original requirements. OpenAI describes GPT-6 Astra as the first frontier model trained in this way. In the post’s final section, it invites a small number of software companies to discuss similar partnerships. It asks potential partners to bring concrete examples of tasks agents cannot reliably complete, people with deep knowledge of the work, a secure test environment and data suitable for research.

The post’s invitation makes the work relevant beyond contract management, but the published measurements are limited to the 11 selected Ironclad tasks. They do not show how an agent would perform across other business applications, organizations or live customer workflows.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Test Does Not Establish

The results do not show whether GPT-6 Astra can reliably complete all requirements in a workflow, or how often it misses particular high-impact rules. OpenAI’s average rubric score combines criteria across tasks, and the published summary does not establish that every task passed a threshold appropriate for use in live contracting. A high score on one showcase task does not resolve that gap.

It is also unclear whether the estimated times would translate into real time savings after users check outputs, correct errors or rerun tasks. OpenAI explicitly labels the figures as simulations, so they should not be treated as measured productivity gains. The source does not establish deployment plans, customer adoption, commercial terms for possible research partners, or whether other vendors will participate.

OpenAI says the work did not use non-public Ironclad customer data, but the available account does not provide a full technical description of the data handling, security controls or evaluation process. Those details would be needed to assess the research methods and how its safeguards might apply in other settings.

Amazon

AI-powered contract drafting tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI’s Proposed Partner Research

OpenAI says it is seeking a small number of software-company partners to identify tasks current agents cannot reliably handle. The company says candidates should provide a specific failing example, subject-matter experts, a secure environment and data that can safely support research. The post does not give a timetable for selecting partners or publishing further results.

For any later evaluation to clarify readiness, useful details would include task-by-task results, the types of criteria missed, how often human reviewers had to intervene, and performance in realistic workflows. Until such evidence is available, the Ironclad study is best understood as a bounded research test rather than a demonstration of verified customer savings or an agent that can manage contracts without supervision.

Amazon

business workflow automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They evaluated GPT-6 Astra on 11 selected legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.

Does the 55% score mean Astra completed 55% of the tasks?

No. OpenAI says 55% is the average share of evaluation criteria met across the tasks, not the share of tasks completed successfully.

Did Astra save customers time?

The published time figures are simulated estimates, not measured customer results. OpenAI says they are based on assumed processing and generation speeds and apply to the 11 research tasks.

Can companies use Astra to manage contracts without human review?

The reported results do not establish that. The average score leaves criteria unmet, and OpenAI’s post says human oversight still matters for workflows that must preserve business rules and controls.

What data did OpenAI say it used?

OpenAI says it created synthetic training tasks from publicly filed SEC EDGAR contracts filtered to remove personal information. It says the project did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Prosumer’s Approach To Desktop Automation Using Watch-Once Commands

A new approach enables power users to automate desktop tasks via watch-once spoken commands, promising simplified workflows and shared libraries.

What Is Simple Moving Average

Discover the critical role of the Simple Moving Average in trading and how it can transform your market analysis—find out more inside!

What Blobspace Means for Ethereum Users

Narrowed network congestion and enhanced scalability make Blobspace a game-changer for Ethereum users, with more to discover about its potential.

U.S. vs. China: DeepSeek and Bitcoin at the Center of Global Trade Power Play

Keen insights into DeepSeek’s AI impact reveal a looming battle in global trade; will the U.S. reclaim its technological edge or falter?