Cryptelio

Hacks & Exploits

Hacktron AI Researchers Breach OpenAI Using Anthropic's Claude Model

Cryptelio Editorial Published 18 Sep 2026 · 18:17 UTC

Researchers at Hacktron AI have successfully breached OpenAI, utilizing Anthropic's Claude AI model to exploit vulnerabilities in the company's systems. The breach, which took less than 72 hours to execute, culminated in a $6,500 bounty awarded by OpenAI after the team demonstrated access to its private source code.

Details of the Breach

The attack began with an image upload feature on OpenAI's community help forum, which is powered by third-party software called Discourse. A safety filter designed to screen uploaded files failed to recognize certain photo formats, allowing them to bypass security checks. These files subsequently reached an image-processing library with a known memory-corruption flaw.

Hacktron's team, consisting of Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini, attempted to exploit this flaw in late July. Initially, Claude's earlier model struggled against a security safeguard that randomized memory locations. However, a newer version of Claude successfully generated functional attack code tailored to the forum's configuration.

While this initial access was limited to the forum's servers, a second vulnerability in OpenAI's single sign-on system allowed the researchers to hijack an employee's session, granting them direct access to OpenAI's private code repository. Following the incident, Discourse patched the image upload vulnerability, rating its severity at 8.8 out of 10, while OpenAI addressed the authentication flaw within approximately 14 hours of the report.

Broader Implications

This breach raises significant security concerns for OpenAI, particularly as the company faces scrutiny regarding its valuation and investor confidence. The incident reflects ongoing challenges in cybersecurity within the tech industry, echoing previous breaches and vulnerabilities faced by leading companies.

As observers monitor OpenAI's response, including potential updates from CEO Sam Altman, the incident may influence market perceptions and future funding opportunities for the company.

New Facts from Hacktron AI Researchers Breach OpenAI

  • The breach occurred in July 2026, with public disclosure on September 18, 2026.
  • The Hacktron AI research team included Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini.
  • Access was gained to employee ChatGPT and Codex accounts, as well as private GitHub repositories.
  • The attack exploited a heap buffer overflow in the libheif library and an overly permissive single sign-on (SSO) token.
  • A specialized version of Anthropic’s Claude Opus model was used to develop the exploits.
  • The full attack chain was executed within 72 hours of starting work.
  • OpenAI patched the vulnerabilities within approximately 14 hours after being notified.
  • Hacktron was awarded a $6,500 bounty for the responsible disclosure of the vulnerabilities.
  • This incident highlights the dual-use problem of AI, as the researchers utilized a legitimate AI model for security research.
  • The libheif vulnerability is significant as it is embedded in numerous software products across the tech industry.

FAQ

What was the main method used by Hacktron AI to breach OpenAI?

Hacktron AI utilized Anthropic's Claude AI model to exploit vulnerabilities in OpenAI's systems, specifically targeting an image upload feature on OpenAI's community help forum.

How long did it take for Hacktron AI to execute the breach?

The breach took less than 72 hours to execute.

What was the outcome of the breach for Hacktron AI?

Hacktron AI was awarded a $6,500 bounty by OpenAI after demonstrating access to its private source code.

What vulnerabilities were exploited during the breach?

The researchers exploited a flaw in the image upload feature that allowed certain photo formats to bypass security checks, and a second vulnerability in OpenAI's single sign-on system that enabled them to hijack an employee's session.

What actions were taken by Discourse and OpenAI following the breach?

Discourse patched the image upload vulnerability, rating its severity at 8.8 out of 10, while OpenAI addressed the authentication flaw within approximately 14 hours of the report.

Read story →