OpenAI and contract-management company Ironclad are turning real contracting workflows into training and evaluation tasks for computer-use agents, with GPT-6 Astra the first OpenAI frontier model trained on the resulting work.
The October 6 research announcement is not a general release of a legal automation product. It is a collaboration aimed at teaching and measuring whether models can complete long, structured tasks inside business software where the hard part is often keeping track of rules, dependencies and state across many steps.
The tasks come from actual contracting workflows
OpenAI says Ironclad worked with its researchers to translate complex legal-operations processes into computer-use tasks. Examples include configuring agreement workflows, setting approval conditions and creating reusable contract terms inside Ironclad’s platform.
That is a different benchmark shape from asking a model one isolated question. A computer-use agent has to navigate an interface, understand which settings affect later steps and avoid breaking the workflow while it works.
The practical point is less about contracts specifically and more about environment complexity. Business applications tend to contain nested menus, permissions, conditional logic and long chains of actions. Turning those workflows into repeatable training tasks gives a model a way to practise the kind of persistence that normal text benchmarks do not measure well.
OpenAI reports a sizeable evaluation gain
OpenAI says GPT-6 Astra is the first frontier model trained on the Ironclad-derived tasks. On OpenAI’s research evaluation, the company reports that Astra’s average score was 32 percent higher than GPT-5.6 Sol, while estimated time per attempt was 48 percent lower.
Those figures should be read as OpenAI’s reported results on its own research setup, not as a universal performance guarantee for every contracting workflow. The announcement does not establish that Astra can independently handle legal work without supervision, nor does it turn benchmark success into legal judgement.
The collaboration is useful because it tests something narrower and more concrete: can a model operate software accurately enough to complete multi-step configuration work that resembles what professionals actually do?
Ironclad already uses generative AI in contract work
The new research sits on top of an existing relationship between Ironclad and OpenAI. An earlier OpenAI case study described Ironclad using GPT-4-based systems for contract review and analysis, including extracting information and helping users work through agreement language.
That background matters because the October project is not beginning with a blank interface. Ironclad already has structured workflows and domain-specific software that can be turned into tasks with clear start and end states.
The wider implication is about how computer-use models may be trained. Rather than relying only on generic web navigation, OpenAI is using specialised software partners to supply difficult, auditable workflows that resemble real knowledge work.
There is still a large gap between completing a benchmark task and being trusted with consequential legal operations. Permissions, review, accountability and domain judgement remain part of the deployment problem. But the Ironclad project shows where OpenAI is putting some of the training effort: not just into making agents click faster, but into making them follow complicated business logic without losing the thread.
Reporting notes