🔍 Read the full analysis: What’s In The Fine Print As OpenAI Trains Agents Inside Your Software? on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training a frontier model inside hosted copies of contract-management software company Ironclad’s product, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, according to OpenAI, while the reported time estimates were simulated rather than measured customer savings. The work also signals that OpenAI is seeking software partners to help train agents on specialized workflows.
OpenAI has described training its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software, reporting that Astra met an average 55% of the criteria used to grade 11 tasks. The October 6 post also invites a small number of software companies to partner on similar research, pointing toward a model-training approach built around agents working inside specialized business products.
The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and revising a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task. Depending on complexity, each task was assessed against 8 to 50 criteria; the score reflects the share of criteria met, not the share of tasks fully completed.
OpenAI reported that GPT-6 Astra, run at its maximum setting, met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. An internal model used in Astra’s development reached 63.7%, according to the post. On one showcase task, Astra met about 94% of the criteria. OpenAI also estimated 19.2 minutes per Astra attempt, versus 37 minutes for GPT-5.6 Sol.
Those time figures are simulated estimates, not observed customer results, OpenAI said. They use assumed processing and generation speeds and cover the 11 research tasks, not Ironclad workflows generally. OpenAI said the training used synthetic tasks derived from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Scores Matter
The results show both the potential and the current limits of agents operating in business software. A model that can set up a workflow or modify a clause may reduce manual work over time, but 55% of criteria met does not establish that the workflow is safe or complete. In contract and procurement operations, a missed approval rule can matter more than several correctly completed steps.
OpenAI’s own example describes a procurement process that may need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. A system that misses one of those conditions could route a purchase incorrectly. That makes the grading rubric and the identity of failed criteria more useful to prospective buyers than an average score alone. The post says human oversight remains necessary when an agent may lose track of a business rule during a task.
For software vendors, the partnership model offers a way to influence how agents handle their products and to identify failure points in real workflows. It also carries a strategic question: if customers increasingly use an agent rather than product screens, the vendor’s value may depend more on business rules, records, audit trails and controls than on its interface. That is an implication of the direction described, not a demonstrated outcome of this study.
contract management software with AI integration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Tests Worked
The October 6 post, titled “Advancing computer use with Ironclad,” was one of two OpenAI publications that day. The source article says the other, about 722 mathematics manuscripts, attracted more attention. Some AI-news trackers interpreted “Ironclad” as the name of an agent framework; it refers instead to the contract-management software company.
OpenAI says Ironclad supplied hosted product copies where models could practice. The work was designed to test whether models could understand company rules, complete multi-step tasks in specialized software and check completed work against the task’s requirements. The 11 tasks were a research set, not a claim that Astra can reliably handle all work in Ironclad or other enterprise applications.
The partnership invitation is a further part of the announcement. OpenAI says it is seeking a small number of software companies that can provide a concrete example of a task current agents fail, subject-matter experts, a secure test environment and data that can be used safely for research. The post frames this as a way to develop agents for work that remains difficult in specialized products.
“Agents have to preserve “the controls teams rely on.””
— Ironclad CTO Sunita Verma
What the Benchmark Leaves Open
The reported averages do not identify which criteria Astra missed on each task, how often an error would affect a real contract or whether a human reviewer would catch it. The source material does not provide evidence that the model has been deployed for customer work or that it has produced measured productivity gains. The estimated task times should not be read as demonstrated savings.
It is also unclear how performance would change across a wider range of contracts, company policies, software configurations or less controlled environments. OpenAI’s description covers 11 selected tasks in hosted copies of Ironclad’s product; it does not establish reliability across the full product or other vendors’ systems. The data-use statements are OpenAI’s account of the research, and the supplied material does not describe an independent audit of them.
Finally, the post does not specify how many software companies will be selected, what partnership terms would apply or how resulting model capabilities would be made available. The strategic effect on software vendors—whether agents complement product interfaces or displace some customer interaction—remains an open question.
What OpenAI’s Partner Search Requires
OpenAI says it is looking for a small group of software-company partners. Applicants are asked to bring a specific task agents currently fail, people with deep knowledge of the work, a secure testing environment and data suitable for research. The post does not give a selection timeline or name additional partners.
For companies evaluating agents, the immediate next step is to seek task-level evidence rather than rely on an average score: which rules were tested, what failed, how sensitive actions are reviewed and what records are kept. OpenAI’s work with Ironclad establishes a research direction, but the published figures do not show that agents are ready to complete consequential contract or procurement work without human checks.
Key Questions
What is Ironclad in this announcement?
Ironclad is a contract-management software company. OpenAI’s post describes testing agents in hosted copies of Ironclad’s product, not introducing a framework named Ironclad.
What does Astra’s 55% score mean?
OpenAI says Astra met an average of 55% of the grading criteria across the 11 tasks. It does not mean Astra completed 55% of the tasks, and the supplied results do not show that each workflow was fully correct.
Did OpenAI measure customer time savings?
No. The reported 19.2-minute estimate for Astra was simulated using assumed processing and generation speeds. OpenAI says it is not measured time saved by customers.
What data did OpenAI say it used?
OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It also said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can companies partner with OpenAI on similar work?
OpenAI said it is inviting a small number of software companies to participate. It asks prospective partners to bring a difficult task example, experts, a secure test environment and data that can safely be used for research; further selection details have not been provided.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
