🔍 Read the full analysis: A Look At Ironclad’s Fine Print Amid OpenAI’s Agent Training In Software on ThorstenMeyerAI.com
TL;DR
OpenAI described training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while estimated completion times were simulated and OpenAI says human oversight remains necessary.
OpenAI said on Oct. 6 that it trained and evaluated a frontier model on workflows inside Ironclad’s contract-management software, reporting that GPT-6 Astra met an average of 55% of the criteria across 11 tasks. The results describe research, not demonstrated customer productivity gains: OpenAI says its time estimates were simulated, and the company says people still need to check the work.
The tasks were selected by Ironclad staff and OpenAI employees who use the product. They covered legal, commercial and procurement work, including creating nondisclosure agreements, setting up procurement approval processes and revising a reusable contract clause according to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task.
OpenAI scored each task against a rubric of 8 to 50 criteria, depending on its complexity. GPT-6 Astra met an average 55% of those criteria, compared with 41.6% for GPT-5.6 Sol at the stated high setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria, but that result does not describe its average performance.
The 55% figure is the share of criteria met on average, not the percentage of tasks completed successfully. OpenAI also reported estimated times of 19.2 minutes per attempt for Astra and 37 minutes for GPT-5.6 Sol. It says these figures are simulated estimates based on assumed processing and generation speeds, rather than measured time savings for customers. They apply to the 11 research tasks, not to Ironclad workflows generally.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed „Ironclad“ was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are „simulated estimates … not measured customer time savings,“ per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflows Need Oversight
The results point to a possible way to train AI agents for specialized business software: test them against concrete workflows and score them on whether they follow the requirements. But in contract and procurement work, a partial result may still create a serious process failure. A workflow that misses a required Finance, Security or Legal approval is not necessarily useful just because it satisfied other checklist items.
OpenAI’s own account acknowledges that an agent can lose track of a business rule partway through a task and says human oversight remains necessary. Ironclad CTO Sunita Verma also emphasized that agents must preserve the controls teams rely on. The reported average therefore signals progress on a difficult task, but does not show that the system can be trusted to run these processes unsupervised.
The collaboration may also matter to software vendors and their customers. If models learn to operate products through real workflows, customers could increasingly use an agent instead of navigating a product’s screens. That possibility puts added weight on the software’s underlying business rules, records, audit trails and controls—the parts that must remain dependable even if the interface changes.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Tests Were Built
OpenAI says Ironclad provided hosted copies of its product where models could practise. The training tasks were generated synthetically from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database. OpenAI says it filtered those documents to remove personal information and did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
The work is presented as a way to train models to follow company rules, complete multi-step tasks in specialized software and check the output against the original requirements. OpenAI identifies GPT-6 Astra as its first frontier model trained in this way. The company’s post also frames a full contracting platform as necessary, rather than suggesting that an agent alone replaces the system’s controls and records.
OpenAI says it is inviting a small number of software companies to work on tasks that current agents cannot complete reliably. It asks prospective partners to bring a specific failing example, people with deep knowledge of the work, a secure testing environment and data that can be used safely for research.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
The published results do not establish how Astra would perform across Ironclad’s full range of customer workflows, or how often it would miss each individual requirement. An average criteria score can conceal whether errors involve minor details or mandatory approvals. OpenAI’s example of a high score on one task does not resolve that gap.
It is also unclear whether the simulated time estimates would translate into faster completed work after a person checks every requirement and corrects mistakes. The source material does not provide measured customer deployments, a broader comparison with real users or evidence that these tasks can be run without supervision. OpenAI says the training used no non-public Ironclad customer data, but the details of any future partnerships and their safeguards have not been specified here.
As an affiliate, we earn on qualifying purchases.
The Next Test Is Reliable Deployment
OpenAI says it plans to work with a small number of software companies on difficult tasks where current agents fail. The proposed partners would supply real examples of those failures, subject-matter expertise, secure test environments and research-appropriate data. OpenAI has not identified additional partners or announced a timetable in the source material.
For businesses considering agents in contracting, procurement or other controlled workflows, the reported scores are a reason to ask for task-level evidence rather than rely on a single average. Buyers will need to know which criteria were missed, how approvals are enforced, what a human must review and how errors are recorded. Until those questions are answered with deployment evidence, the Ironclad tests remain a research result—not proof of unsupervised, time-saving performance.
procurement approval workflow software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They evaluated an AI model on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.
Does GPT-6 Astra’s 55% score mean it completed 55% of tasks?
No. OpenAI says 55% is the average share of rubric criteria met across the tasks, not the share of tasks completed successfully.
Did the reported time estimates measure customer savings?
No. OpenAI describes the 19.2-minute and 37-minute figures as simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings.
What data did OpenAI say it used?
OpenAI says it created synthetic training tasks from public contracts filed in the SEC’s EDGAR database, with personal information filtered out. It says it did not use OpenAI customer data, internal contracts or non-public Ironclad customer data.
Can companies deploy these agents without human review?
The reported results do not support that conclusion. OpenAI says human oversight remains necessary, and an average criteria score does not show that every required approval or business rule was followed.
Source: ThorstenMeyerAI.com