FIELD NOTE

OpenAI Third-Party AI Safety Assessments: The Plan

OpenAI proposes four priorities and seven operating principles for independent third-party AI safety assessments of frontier systems.

OpenAI’s third-party AI safety assessments plan defines what outside reviewers should examine and how those reviews should be run. Published September 22, 2026, it covers safety cases, safeguards, capability evaluations and critical misalignment incidents. The document focuses on independent private and nonprofit organizations performing technical safety assessments; government testing is complementary and may use different approaches. Read OpenAI’s assessment principles.

What third-party AI safety assessments would cover

OpenAI identifies four priority areas rather than one generic audit. The first is a review of safety cases across training, evaluation, internal deployment and external deployment. A safety case is the evidence-backed argument that risks are adequately managed for a specified activity. It connects individual claims to evidence and exposes assumptions, uncertainty and remaining risks.

The second priority is testing critical safeguards. OpenAI lists model-level, enforcement and security safeguards, along with misalignment monitors. An assessor might examine whether safeguards resist authorized adversarial testing, whether cyber defenses contain harmful agent actions under realistic conditions, whether monitoring has critical gaps and whether controls match the model’s capabilities.

The third priority is reviewing capability evaluations. OpenAI names the Preparedness Framework categories of chemical and biological risks, cybersecurity and AI self-improvement, plus alignment evaluations for severe misalignment. Reviewers can ask whether a test covers its stated threshold, whether saturated tests are refreshed and whether important behaviors or conditions are missing.

The fourth priority is an independent investigation of critical misalignment incidents, including behavior such as acting without authorization or evading oversight. The proposed questions are practical: what happened, what contributed to it and whether safeguards or remediation would prevent a similar incident. Relevant expertise may include cyber forensics, alignment research and large-scale chain-of-thought analysis.

These tracks can involve different specialists; OpenAI says no single third party should cover every urgent frontier-safety question. Assessments may last weeks or several months. Although one can inform deployment, the program described here is generally longer-term and launch-agnostic.

Safety claims need a defined scope and evidence

The proposal distinguishes a safety claim from a safety case. A safety claim is a specific assertion about capability, behavior or safeguards that can be checked against evidence. It should identify the risk, operating conditions, assumptions and limits. A safety case organizes multiple claims into an argument about why risk is adequately managed for a particular activity.

A broad label such as “safe” gives an assessor little to test. A claim such as “this safeguard detects a defined class of harmful actions under these operating conditions” creates a reviewable boundary. Evidence can then support, weaken or leave the claim unresolved without implying a universal verdict on the whole model.

OpenAI proposes that claims be clearly scoped and preregistered before work begins. The parties should state whether a claim comes from the lab or an independent question raised by the assessor. They should identify exclusions caused by unavailable data, insufficient expertise or urgent time limits, and define how significant risks found outside the scope may trigger further investigation.

For readers, the first useful check is the claim inventory: what was tested, in which environment, against which criteria and with which evidence. Our source-backed research brief skill provides a reusable way to keep supported claims, conflicts and unknowns separate. It is not an AI safety assessment standard, but its evidence structure is useful when reviewing one.

Access and independence are part of the result

OpenAI says assessors should receive access proportionate to the agreed claims, within legal, security and intellectual-property constraints. Previous engagements have included information about technical safeguards, visible chain-of-thought access, confidential data and internal-deployment access for incident response and monitor red teaming. The proposal also allows indirect or privacy-preserving methods when direct access is impractical.

A report should explain enough about access for readers to understand what the assessor could observe. A black-box test, a grey-box adversarial review and an investigation with confidential internal records answer different questions. Limited access can narrow the conclusions the evidence supports.

Assessors should disclose conflicts, including financial incentives, developer relationships and prior involvement in the work under review. OpenAI suggests safeguards such as recusal or exclusion periods so commercial pressure and compensation do not determine findings. Relevant technical expertise must also be demonstrated rather than inferred from organizational independence.

Security requirements can limit how evidence is handled. The principles call for information-security practices and enforceable confidentiality controls proportionate to the accessed systems and records. Where an assessor cannot secure material in its own environment, OpenAI says company-managed devices or premises may be appropriate. Readers should distinguish protected detail from a report that omits its methods and uncertainty.

Findings should lead to action without losing editorial control

OpenAI proposes that assessments identify specific gaps and give labs enough detail to address them. Where appropriate, a lab may receive a reasonable remediation period before publication. Reports should also give practical recommendations to agent developers, deployers and defenders when relevant.

That remediation window is paired with editorial independence. Reports should be as open as possible while protecting sensitive information. Clear redaction policies should let a lab request removal of protected details, while the assessor can note substantive redactions and their effect. Confidential reporting to a board or other appropriate governance body can support accountability when full disclosure is not possible.

A useful reading checklist follows from those principles: preregistered claims, assessor expertise and conflicts, access level, methods and criteria, uncertainty, out-of-scope risks, remediation status and material redactions. This is our editorial checklist derived from the proposal, not a certification program announced by OpenAI.

How this differs from misalignment reporting

OpenAI’s proposal does not announce a completed independent assessment, name a universal auditor or certify any model. It says the company is discussing proposals with multiple third parties and intends to help develop shared practices and international standards. The publication is a framework for future work, not evidence that every listed review has happened.

It also serves a different purpose from the model-misalignment disclosure process. Our coverage of the OpenAI model misalignment reporting framework explains how qualifying incidents move toward public disclosure. The new principles address how external organizations could examine safety claims, safeguards, evaluations and selected incidents with defined access and independence. An incident report investigates a concerning event; a third-party assessment can also test preventive claims outside a specific incident.

The material change is a concrete operating model for independent scrutiny. The proposal tells reviewers and labs what questions deserve priority, what access and expertise matter, and how to preserve confidentiality and editorial independence. Its value will depend on later engagements publishing enough scope, evidence and uncertainty for outsiders to judge what was established.