OpenAI Ironclad: GPT-6 Astra Contracting Results
GPT-6 Astra outscored GPT-5.6 Sol in OpenAI's Ironclad contracting study; its timing results are simulated.
OpenAI Ironclad research published October 6 compared GPT-6 Astra with GPT-5.6 Sol on 11 contracting tasks. Astra scored higher with lower simulated time; this is a research evaluation, not a customer rollout. OpenAI’s research announcement
OpenAI Ironclad results: scores and estimated time
The comparison used Astra at Max reasoning and Sol at High reasoning.
| Metric across the 11 tasks | GPT-5.6 Sol | GPT-6 Astra |
|---|---|---|
| Mean rubric score | 41.6% | 55.0% |
| Estimated minutes per attempt | 37.0 | 19.2 |
OpenAI reports a 32% relative score gain and 48% lower estimated time. Timings are simulated using assumed processing and generation speeds, not measured customer savings. Evaluation and timing scope
The score difference is 13.4 percentage points. Dividing that difference by the older score produces the relative improvement; it does not mean the newer model completed 32 additional tasks out of 100. Likewise, a lower estimated attempt time says nothing by itself about how much time a reviewer needs to correct a result. Keep score, elapsed execution time and human review effort as separate measurements when comparing your own runs.
For the model’s general capabilities and interfaces, use the GPT-6 Astra model page. This dated research result answers a narrower question: how two specified configurations performed on this particular contracting evaluation.
Make the rubric describe a working process
OpenAI’s evaluation guidance recommends task-specific tests, representative inputs, clear objectives and human calibration of automated scoring. It also recommends testing repeatedly as an application changes rather than treating one successful demonstration as permanent acceptance. These are methodological principles; following them does not require adopting a particular hosted evaluation product. Evaluation best practices
For a contracting workflow, a useful editorial distinction is between an action and an outcome. Clicking an approval option is an action. Producing the intended approval path for each request is an outcome. A checklist that counts clicks could reward a visually convincing configuration while missing the business rule it was supposed to implement.
Our recommendation is to write acceptance criteria before the trial. Identify the inputs, the expected path, the person authorized to approve it and the saved evidence that proves the result. Give critical controls their own pass-or-fail checks instead of allowing many harmless successes to conceal one incorrect approval route. This is a proposed evaluation design, not a scoring system used in OpenAI’s published study.
A renewal-review example you can adapt
The following is an untested, fictional acceptance exercise, not an Ironclad feature description. Suppose your organization wants a draft renewal-review queue: agreements expiring within an internally chosen number of days need an assigned owner, and records missing an owner need a separate exception queue.
Create a disposable test workspace with invented agreements and expiration dates. Ask the agent to describe the intended routing first, then configure a draft within that test environment if its available tools permit it. Use an instruction such as: “Prepare the renewal-review queue from these requirements. Do not activate reminders or change agreements. Return the routing decisions and the evidence needed for a reviewer to test each branch.”
Test four records: inside the review window with an owner; outside it with an owner; inside it without an owner; and outside it without an owner. For each, write down the expected queue before running the case. Add a record exactly on the boundary date and decide whether that date belongs inside or outside the window. Decide how a missing expiration date should be handled as well. These are your own policy inputs, not rules supplied by OpenAI or Ironclad.
Record the actual path alongside the expected one. If a branch differs, preserve the request and configuration evidence before correcting it. Then rerun the affected branch and the cases that previously passed. The result should show which requirement changed and whether that change disturbed another route. This makes a second trial comparable to the first without relying on the agent’s narrative of what it did.
The document-reconciliation example offers a related pattern for checking consistency between requirements and a finished artifact. Adapt that checking structure, not its subject matter, to your contracting test.
Computer-use setup is separate from the research
OpenAI’s current Computer Use documentation describes a desktop capability for ChatGPT Work and Codex on macOS and Windows in supported regions. It requires the Computer Use plugin; macOS also requires Screen Recording and Accessibility permissions. App approvals, operating-system permissions and the task’s file or command permissions remain separate. Windows tasks use the active desktop. When a suitable structured integration exists, OpenAI recommends it for repeatable operations and data access. Computer Use setup and permissions
Those instructions describe the general feature, not access to Ironclad’s research environment. Before a local trial, our recommendation is to choose the interface you actually need, confirm the account is a test account and close unrelated sensitive applications. Granting an agent permission to operate an app should not be treated as approval to activate a process, issue a contract or accept legal terms. Name the permitted actions and the stopping point in the task.
Keep the evidence and the handoff together
For your own pilot, prepare a short handoff containing the requirements, test inputs, resulting configuration, branch-by-branch results and unresolved decisions. Ask the process owner to verify the actual saved state rather than signing off on a polished summary alone. Review effort is worth recording alongside execution time because a fast draft and a usable workflow are different deliverables.
The useful next step is a small, reviewable test with explicit acceptance criteria. The Ironclad computer use evaluation supplies a concrete reason to test business-rule preservation; your organization’s evidence must establish whether its own process works.