Cryptelio

Companies

OpenAI Introduces Third-Party Evaluations for AI Model Safety

Cryptelio Editorial Published 23 Sep 2026 · 09:15 UTC

On September 22, OpenAI announced a significant shift in its approach to AI safety by allowing independent evaluators access to its models during earlier stages of development. This initiative aims to identify safety issues before they escalate into larger problems.

The expanded evaluation program focuses on four key areas: assessing OpenAI's safety cases, evaluating the resilience of safeguards against adversarial threats, aligning assessments with the company's Preparedness Framework, and investigating incidents of misalignment. OpenAI has laid out seven guiding principles for these assessments, emphasizing scientific rigor, independence, and responsible publication of findings.

Organizations like METR and Redwood Research, known for their expertise in AI safety, are in discussions to partner with OpenAI for this initiative. This move builds on commitments made by CEO Sam Altman earlier in September, signaling a deeper integration of evaluators within OpenAI's operational structure.

The new framework represents a proactive approach to safety, moving evaluations upstream in the development timeline, which could serve as a model for industry-wide safety standards. However, the challenge remains in balancing transparency with commercial interests, particularly if independent assessments reveal findings that conflict with OpenAI's business goals.

FAQ

What is the purpose of OpenAI's third-party evaluations?

The purpose of OpenAI's third-party evaluations is to identify safety issues in AI models during earlier stages of development, before they escalate into larger problems.

What areas does the expanded evaluation program focus on?

The expanded evaluation program focuses on assessing OpenAI's safety cases, evaluating the resilience of safeguards against adversarial threats, aligning assessments with the company's Preparedness Framework, and investigating incidents of misalignment.

Who are the organizations partnering with OpenAI for these evaluations?

Organizations like METR and Redwood Research, known for their expertise in AI safety, are in discussions to partner with OpenAI for this initiative.

What are the guiding principles for the assessments?

OpenAI has laid out seven guiding principles for the assessments, emphasizing scientific rigor, independence, and responsible publication of findings.

What challenge does OpenAI face with this new evaluation framework?

The challenge OpenAI faces is balancing transparency with commercial interests, especially if independent assessments reveal findings that conflict with OpenAI's business goals.

Read story →