What’s In The Fine Print As OpenAI Trains Agents Inside Your Software?
KIDieser Beitrag wurde mit Unterstützung künstlicher Intelligenz (KI) erstellt.

🔍 Read the full analysis: What’s In The Fine Print As OpenAI Trains Agents Inside Your Software? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model inside hosted copies of contract-management software company Ironclad’s product, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, according to OpenAI, while the reported time estimates were simulated rather than measured customer savings. The work also signals that OpenAI is seeking software partners to help train agents on specialized workflows.

OpenAI has described training its GPT-6 Astra model on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software, reporting that Astra met an average 55% of the criteria used to grade 11 tasks. The October 6 post also invites a small number of software companies to partner on similar research, pointing toward a model-training approach built around agents working inside specialized business products.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and revising a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task. Depending on complexity, each task was assessed against 8 to 50 criteria; the score reflects the share of criteria met, not the share of tasks fully completed.

OpenAI reported that GPT-6 Astra, run at its maximum setting, met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. An internal model used in Astra’s development reached 63.7%, according to the post. On one showcase task, Astra met about 94% of the criteria. OpenAI also estimated 19.2 minutes per Astra attempt, versus 37 minutes for GPT-5.6 Sol.

Those time figures are simulated estimates, not observed customer results, OpenAI said. They use assumed processing and generation speeds and cover the 11 research tasks, not Ironclad workflows generally. OpenAI said the training used synthetic tasks derived from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

At a glance
reportWhen: Published October 6; research results a…
The developmentOpenAI published details of a research partnership with Ironclad that trains agents on workflows inside the contract-management product and invites other software companies to participate.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Scores Matter

The results show both the potential and the current limits of agents operating in business software. A model that can set up a workflow or modify a clause may reduce manual work over time, but 55% of criteria met does not establish that the workflow is safe or complete. In contract and procurement operations, a missed approval rule can matter more than several correctly completed steps.

OpenAI’s own example describes a procurement process that may need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. A system that misses one of those conditions could route a purchase incorrectly. That makes the grading rubric and the identity of failed criteria more useful to prospective buyers than an average score alone. The post says human oversight remains necessary when an agent may lose track of a business rule during a task.

For software vendors, the partnership model offers a way to influence how agents handle their products and to identify failure points in real workflows. It also carries a strategic question: if customers increasingly use an agent rather than product screens, the vendor’s value may depend more on business rules, records, audit trails and controls than on its interface. That is an implication of the direction described, not a demonstrated outcome of this study.

Amazon

contract management software with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Tests Worked

The October 6 post, titled “Advancing computer use with Ironclad,” was one of two OpenAI publications that day. The source article says the other, about 722 mathematics manuscripts, attracted more attention. Some AI-news trackers interpreted “Ironclad” as the name of an agent framework; it refers instead to the contract-management software company.

OpenAI says Ironclad supplied hosted product copies where models could practice. The work was designed to test whether models could understand company rules, complete multi-step tasks in specialized software and check completed work against the task’s requirements. The 11 tasks were a research set, not a claim that Astra can reliably handle all work in Ironclad or other enterprise applications.

The partnership invitation is a further part of the announcement. OpenAI says it is seeking a small number of software companies that can provide a concrete example of a task current agents fail, subject-matter experts, a secure test environment and data that can be used safely for research. The post frames this as a way to develop agents for work that remains difficult in specialized products.

“Agents have to preserve “the controls teams rely on.””

— Ironclad CTO Sunita Verma

What the Benchmark Leaves Open

The reported averages do not identify which criteria Astra missed on each task, how often an error would affect a real contract or whether a human reviewer would catch it. The source material does not provide evidence that the model has been deployed for customer work or that it has produced measured productivity gains. The estimated task times should not be read as demonstrated savings.

It is also unclear how performance would change across a wider range of contracts, company policies, software configurations or less controlled environments. OpenAI’s description covers 11 selected tasks in hosted copies of Ironclad’s product; it does not establish reliability across the full product or other vendors’ systems. The data-use statements are OpenAI’s account of the research, and the supplied material does not describe an independent audit of them.

Finally, the post does not specify how many software companies will be selected, what partnership terms would apply or how resulting model capabilities would be made available. The strategic effect on software vendors—whether agents complement product interfaces or displace some customer interaction—remains an open question.

What OpenAI’s Partner Search Requires

OpenAI says it is looking for a small group of software-company partners. Applicants are asked to bring a specific task agents currently fail, people with deep knowledge of the work, a secure testing environment and data suitable for research. The post does not give a selection timeline or name additional partners.

For companies evaluating agents, the immediate next step is to seek task-level evidence rather than rely on an average score: which rules were tested, what failed, how sensitive actions are reviewed and what records are kept. OpenAI’s work with Ironclad establishes a research direction, but the published figures do not show that agents are ready to complete consequential contract or procurement work without human checks.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes testing agents in hosted copies of Ironclad’s product, not introducing a framework named Ironclad.

What does Astra’s 55% score mean?

OpenAI says Astra met an average of 55% of the grading criteria across the 11 tasks. It does not mean Astra completed 55% of the tasks, and the supplied results do not show that each workflow was fully correct.

Did OpenAI measure customer time savings?

No. The reported 19.2-minute estimate for Astra was simulated using assumed processing and generation speeds. OpenAI says it is not measured time saved by customers.

What data did OpenAI say it used?

OpenAI said it created synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It also said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can companies partner with OpenAI on similar work?

OpenAI said it is inviting a small number of software companies to participate. It asks prospective partners to bring a difficult task example, experts, a secure test environment and data that can safely be used for research; further selection details have not been provided.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring KIRIN EXPRESS’s Floating-Seam Journey And Its Image-Free Workflow

A browser-based exhibition imagines a 603 km/h maglev through scroll-linked motion, interactive controls and visuals drawn without image assets.

Anthropic Claude Subscription Plan Rated More Valuable Than OpenAI’s

The Register reported a study favored Claude over OpenAI’s subscription, but the study, methods and plans compared are not identified.

10 Best Ultrawide Monitors For Productivity And Play In 2027

A specification-based guide to 10 ultrawide monitors, comparing workspace, resolution, USB-C connectivity and gaming refresh rates.

The SSD Squeeze: Why Storage Joined the Party

Storage prices are rising sharply due to supply constraints, AI-driven demand, and factory competition, impacting enterprise and consumer markets.