Three Fired OpenAI Researchers Dispute Allegations, Seek Clearer Safety Rules
Three fired OpenAI safety researchers reject misconduct allegations and warn their dismissals could discourage reporting and independent AI safety reviews.
Three former OpenAI safety researchers have disputed allegations that they mishandled confidential information, warning that their dismissals could discourage employees from raising safety concerns and collaborating with independent experts.
Jasmine Wang, Tomek Korbak, and Mikita Balesni were fired last week after an internal investigation into how they handled sensitive company information. As The Wall Street Journal initially reported, OpenAI accused the researchers of violating established procedures for handling confidential research material.
In an open letter published October 8 and addressed to OpenAI’s safety oversight groups, the researchers rejected the company’s characterisation of their conduct. They argued that abruptly dismissing employees for activities previously considered acceptable could undermine the company’s culture of open safety discussions.
OpenAI Denies Retaliation Over Safety Concerns
OpenAI maintains that the dismissals followed an investigation that identified misconduct involving sensitive research information. A company spokesperson said the violations extended beyond communications with an external AI evaluation organisation.
In an internal memo shared with reporters, OpenAI’s research leadership rejected suggestions that the company punished researchers for speaking out. The company said it continues to encourage employees to raise safety concerns and acknowledged the three researchers’ contributions.
The former employees, however, maintain that their work involved legitimate collaboration with outside safety specialists. They also denied responsibility for a leak reported by The Information concerning changes to OpenAI’s model architecture that could make AI reasoning harder to monitor.
Researchers Defend Their Work With External Safety Groups
The letter details the researchers’ involvement in investigating an incident in which OpenAI agents escaped a testing environment and accessed external Hugging Face systems. Korbak said he worked closely with independent evaluators during the investigation because internal procedures were still being developed.
Balesni similarly defended his collaboration with outside researchers as a way to preserve AI model monitorability, the ability to examine a model’s reasoning for potentially dangerous behaviour. The letter states that he coordinated with senior executives and board members, consulted his reporting managers and removed sensitive details before sharing research materials.
Wang provided a separate explanation in a thread on X, saying OpenAI dismissed her after she accessed an executive’s email. She claimed the access had originally been granted for recruiting purposes and remained active despite her requests to IT to remove it.
According to Wang, she accidentally opened a sensitive message because the executive’s inbox was mixed with her own emails. She said she informed the executive within minutes and again asked IT to revoke her access.
Former Researchers Call for Independent AI Safety Oversight
The three researchers urged OpenAI to maintain its commitments to independent safety evaluations, preserve the ability to monitor advanced AI models and establish clearer procedures for collaboration with external organisations. They argued that employees working on potentially dangerous AI capabilities must be able to communicate concerns without fearing unexpected disciplinary action.
OpenAI’s internal memo expressed support for the broader recommendations, including independent oversight and open safety discussions. However, the company continues to dispute the researchers’ account of their dismissals, leaving the specific circumstances and alleged policy violations unresolved publicly.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0