AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Look At Ironclad’s Fine Print Amid OpenAI’s Agent Training In Software on ThorstenMeyerAI.com

TL;DR

OpenAI described training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while estimated completion times were simulated and OpenAI says human oversight remains necessary.

OpenAI said on Oct. 6 that it trained and evaluated a frontier model on workflows inside Ironclad’s contract-management software, reporting that GPT-6 Astra met an average of 55% of the criteria across 11 tasks. The results describe research, not demonstrated customer productivity gains: OpenAI says its time estimates were simulated, and the company says people still need to check the work.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They covered legal, commercial and procurement work, including creating nondisclosure agreements, setting up procurement approval processes and revising a reusable contract clause according to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes for each task.

OpenAI scored each task against a rubric of 8 to 50 criteria, depending on its complexity. GPT-6 Astra met an average 55% of those criteria, compared with 41.6% for GPT-5.6 Sol at the stated high setting. An internal OpenAI model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria, but that result does not describe its average performance.

The 55% figure is the share of criteria met on average, not the percentage of tasks completed successfully. OpenAI also reported estimated times of 19.2 minutes per attempt for Astra and 37 minutes for GPT-5.6 Sol. It says these figures are simulated estimates based on assumed processing and generation speeds, rather than measured time savings for customers. They apply to the 11 research tasks, not to Ironclad workflows generally.

At a glance
reportWhen: Published Oct. 6; further partner work…
The developmentOpenAI published details on Oct. 6 of a collaboration with Ironclad to train and evaluate AI agents on workflows inside contract-management software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed „Ironclad“ was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What „55%“ means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are „simulated estimates … not measured customer time savings,“ per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why „a full contracting platform remains essential.“ Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, „Advancing computer use with Ironclad“ (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of „Ironclad“ in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflows Need Oversight

The results point to a possible way to train AI agents for specialized business software: test them against concrete workflows and score them on whether they follow the requirements. But in contract and procurement work, a partial result may still create a serious process failure. A workflow that misses a required Finance, Security or Legal approval is not necessarily useful just because it satisfied other checklist items.

OpenAI’s own account acknowledges that an agent can lose track of a business rule partway through a task and says human oversight remains necessary. Ironclad CTO Sunita Verma also emphasized that agents must preserve the controls teams rely on. The reported average therefore signals progress on a difficult task, but does not show that the system can be trusted to run these processes unsupervised.

The collaboration may also matter to software vendors and their customers. If models learn to operate products through real workflows, customers could increasingly use an agent instead of navigating a product’s screens. That possibility puts added weight on the software’s underlying business rules, records, audit trails and controls—the parts that must remain dependable even if the interface changes.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Tests Were Built

OpenAI says Ironclad provided hosted copies of its product where models could practise. The training tasks were generated synthetically from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database. OpenAI says it filtered those documents to remove personal information and did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

The work is presented as a way to train models to follow company rules, complete multi-step tasks in specialized software and check the output against the original requirements. OpenAI identifies GPT-6 Astra as its first frontier model trained in this way. The company’s post also frames a full contracting platform as necessary, rather than suggesting that an agent alone replaces the system’s controls and records.

OpenAI says it is inviting a small number of software companies to work on tasks that current agents cannot complete reliably. It asks prospective partners to bring a specific failing example, people with deep knowledge of the work, a secure testing environment and data that can be used safely for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Show

The published results do not establish how Astra would perform across Ironclad’s full range of customer workflows, or how often it would miss each individual requirement. An average criteria score can conceal whether errors involve minor details or mandatory approvals. OpenAI’s example of a high score on one task does not resolve that gap.

It is also unclear whether the simulated time estimates would translate into faster completed work after a person checks every requirement and corrects mistakes. The source material does not provide measured customer deployments, a broader comparison with real users or evidence that these tasks can be run without supervision. OpenAI says the training used no non-public Ironclad customer data, but the details of any future partnerships and their safeguards have not been specified here.

Amazon

electronic NDA signing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Test Is Reliable Deployment

OpenAI says it plans to work with a small number of software companies on difficult tasks where current agents fail. The proposed partners would supply real examples of those failures, subject-matter expertise, secure test environments and research-appropriate data. OpenAI has not identified additional partners or announced a timetable in the source material.

For businesses considering agents in contracting, procurement or other controlled workflows, the reported scores are a reason to ask for task-level evidence rather than rely on a single average. Buyers will need to know which criteria were missed, how approvals are enforced, what a human must review and how errors are recorded. Until those questions are answered with deployment evidence, the Ironclad tests remain a research result—not proof of unsupervised, time-saving performance.

Amazon

procurement approval workflow software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They evaluated an AI model on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.

Does GPT-6 Astra’s 55% score mean it completed 55% of tasks?

No. OpenAI says 55% is the average share of rubric criteria met across the tasks, not the share of tasks completed successfully.

Did the reported time estimates measure customer savings?

No. OpenAI describes the 19.2-minute and 37-minute figures as simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings.

What data did OpenAI say it used?

OpenAI says it created synthetic training tasks from public contracts filed in the SEC’s EDGAR database, with personal information filtered out. It says it did not use OpenAI customer data, internal contracts or non-public Ironclad customer data.

Can companies deploy these agents without human review?

The reported results do not support that conclusion. OpenAI says human oversight remains necessary, and an average criteria score does not show that every required approval or business rule was followed.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

What’s The Cost Of Changing AI Providers? The Claude Example

A report says Meta and Microsoft are shifting internal work from Claude. The changes highlight the costs of switching AI providers.

Creative industries. The bifurcated reality.

New data shows a bifurcation in creative jobs, with top-tier professionals augmenting work and mid-tier roles contracting due to AI substitution in 2025-2026.

Siemens’ AI Strategy: Revolutionizing The Manufacturing Landscape

Siemens unveils a strategic shift toward physical AI, partnering with NVIDIA to embed AI across manufacturing and industrial processes, aiming to reshape the sector.

Build, Rent, or Quantize: Cutting Your Memory Bill Without Cutting Capability

Exploring how AI practitioners can reduce memory expenses through building, renting, or quantizing models, with a focus on recent advances in compression techniques.