OpenAI Fires Researchers for Allegedly Sharing Information with AI Safety Group

Dow Jones
Oct 02

OpenAI has fired three researchers for alleged misconduct, including sharing confidential company information with a third-party AI-safety organization, according to people familiar with the matter.

The company recently told some employees that it had terminated three researchers who worked on its safety team, one of the people said.

The affected employees are Jasmine Wang, Tomek Korbak, and Mikita Balesni, the people familiar with the matter said. The researchers didn't immediately comment.

"We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information," an OpenAI spokesperson said. "Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work."

AI giants are under growing pressure to submit their technology for independent safety testing.

Last month, the CEO of OpenAI rival Anthropic said his company would allow outside evaluators like METR, an AI safety nonprofit, to verify its adherence to safety measures and assess model alignment.

OpenAI has recently faced a spate of security incidents in which its AI agents escaped containment, hacking some company websites and aggressively probing a wide range of websites.

After its model hacked the AI company Hugging Face, OpenAI allowed staff members from METR and a Redwood Research staff member contracting with that group to work in its offices for six days to investigate how its models behaved. METR later released a report based on the information to which the company provided access.

One of the fired OpenAI employees, Korbak, was a member of OpenAI's safety team, and has said he served as the company's technical contact for Redwood Research and METR in their investigation of the Hugging Face incident.

The other terminated employees, Wang and Balesni, worked on alignment, which is ensuring that the company's models behave as humans intend them to.

OpenAI has said that it is working to investigate a range of agent security incidents that it has discovered in recent months, and address underlying safety issues. Earlier this week, OpenAI said it was scrapping the planned launch of an AI model, GPT-6.1 Astra, over safety concerns.

The Chat GPT-maker said it has implemented a new monitoring system to catch AI-agent misbehavior more quickly, started requiring engineers to use stronger security guardrails for testing its AI systems and is sharing more information about instances in which models behave badly.

The AI industry is grappling with the growing capabilities of its most powerful models and the risks they pose.

In early September, Anthropic researcher Jacob Coxon publicly quit, saying he didn't want to participate in a rush to build AI systems that can improve themselves. He said he feared they could spiral out of control and destroy humanity.

Anthropic Chief Executive Dario Amodei wrote last month that the risks posed by today's cutting edge AI tools were too great to continue development at the same breakneck pace and called for slowing down industrywide development, drawing agreement from OpenAI Chief Executive Sam Altman and Elon Musk.

 

At the request of the copyright holder, you need to log in to view this content

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10