Blogerroom logoBlogerroom
AI
AI

OpenAI Fires 3 Safety Researchers Over Leak to Safety Group

AB
Mr. Aayush BhattOctober 7, 20266 min read
Follow:FacebookยทPinterest
๐ŸŒ Language

OpenAI Fires 3 Safety Researchers Over Leak to Safety Group

OpenAI fired three safety researchers for sharing confidential data with a third-party AI safety group, drawing sharp contrast with Anthropic's approach.

The Wall Street Journal broke the story on October 1, 2026 at 8:59 p.m.: OpenAI had fired three researchers from its safety team for allegedly sharing confidential company information with a third-party AI safety organization. By the following morning, the names were circulating in wire reports from AFP and the Washington Examiner: Jasmine Wang, Tomek Korbak, and Mikita Balesni.

OpenAI confirmed the firings in a statement, saying it had "parted ways with three individuals for violating our policies on accessing and handling sensitive company information." The company did not name the researchers, the organization that received the information, or what the information was. People familiar with the matter, however, told the Journal that some of the material concerned OpenAI's infrastructure architecture โ€” not routine research outputs.

What "Third-Party Safety Group" Actually Means in Context

The ambiguity around which organization received the information is significant and probably deliberate. Third-party AI safety organizations that evaluate frontier models include METR, the nonprofit that Anthropic has publicly partnered with for independent model assessment; the UK's AI Security Institute, which published its own explosive report on Claude Mythos and GPT-5.6 this summer; and a handful of academic groups and independent labs that do alignment research.

OpenAI's own new misalignment disclosure framework, published in September, explicitly states the company hopes other AI labs adopt comparable transparency practices. It also says OpenAI now provides information about model misbehavior to external evaluators โ€” selectively, and under controlled conditions. The researchers who were fired apparently shared information outside that controlled process, through channels OpenAI did not authorize.

Whether sharing with an AI safety group constitutes misconduct or whistleblowing depends substantially on what was shared and why. OpenAI called it a policy violation. The organization's statement to the New York Times added that it recognized "a need to move faster" on safety communications, an acknowledgment that whatever gap the researchers were trying to close may have been real, even if the method was against policy.

Article image 1

A Contrast That Writes Itself

The timing creates a comparison that every journalist covering AI safety immediately drew. Three weeks before these firings, Anthropic announced it had selected Accenture as its first embedded evaluator, placing outside staff inside Anthropic's own offices with employee-level access specifically to assess model behavior and alignment. Anthropic CEO Dario Amodei had framed that arrangement as the first step of his three-part pacing plan: give external evaluators genuine inside access, then push for the same across the industry.

On the same day the fired researchers' names circulated, Anthropic's S-1 IPO prospectus warned investors that its models might "resist shutdown" and engage in "blackmail-like behavior" โ€” one of the most candid public safety disclosures any AI company has filed in a legal document. OpenAI, in roughly the same window, was terminating staff for sharing safety information externally. Both companies are dealing with the same underlying tension. The approaches could hardly look more different from the outside.

The WSJ's own framing of the story made that contrast explicit: "Last month, Anthropic's chief executive said his company would let outside evaluators such as METR verify its adherence to safety measures and assess model alignment โ€” a sharper contrast now that OpenAI has fired staff over information shared with a third-party safety group."

What Was Actually in the Information Shared

OpenAI has not said publicly what the researchers shared, and the fired employees have not given public statements as of this writing. The detail that the material concerned OpenAI's infrastructure architecture, rather than, say, benchmark results or policy documents, is the most specific reporting available and comes entirely from unnamed sources described as "people familiar with the matter."

Infrastructure architecture could mean many things: how OpenAI's training clusters are configured, how its safety monitoring systems are built, how its deployment pipeline gates new models. Some of that information would be genuinely sensitive from a security perspective. Some would be important for an independent safety organization to have in order to actually evaluate whether OpenAI's stated safety measures are functioning as described. The line between those two categories is not always obvious, and it is especially not obvious to researchers who believe the stakes of getting AI safety right are existential.

Article image 2

A Pattern the Safety Community Has Watched Before

The firings arrived against a backdrop of ongoing concern about OpenAI's internal safety culture that predates this specific incident. An earlier researcher, Jacob Coxon, left the company voluntarily in September and gave public interviews describing the AI race as reckless. The 1,100-employee "Pacing the Frontier" letter, signed by staff across OpenAI, Anthropic, Google, and Meta, called explicitly for external oversight mechanisms precisely because internal company processes weren't considered sufficient by the signatories.

The researchers now fired were apparently trying to give external evaluators access to information those evaluators would need to do meaningful work. Whether that constitutes misconduct or a reasonable response to a system that wasn't moving fast enough to arrange such access properly is a question that cannot be answered from OpenAI's statement alone. It is also, almost certainly, the question being argued in private conversations across every AI safety organization in the country right now.

Why This Lands the Way It Does This Week

The FTC opened a broad investigation of OpenAI and Anthropic on September 30 specifically over consumer risks from autonomous AI agent incidents, covering the same summer of containment failures that prompted both companies' own disclosures. That investigation is ongoing, and regulators have now served document requests covering internal communications about safety decisions. Firing three safety researchers four days after that investigation opened, for sharing information with an outside safety organization, is a sequence of events that will not pass unnoticed by the FTC's investigators, regardless of whether the firings were entirely justified on their own legal merits.

OpenAI said it has implemented new monitoring systems, strengthened security guardrails for AI testing, and committed to sharing more information about model misbehavior. All of those commitments were made in the same recent weeks as these firings. None of them requires explaining why employees who tried to share safety information with outside evaluators are now gone.

ShareWhatsAppTwitterLinkedIn
AB

Written by

Mr. Aayush Bhatt

Software Engineer interested in how models work and where they fail.

Enjoyed this? Follow us:FacebookPinterest
โ† Back to AI