OpenAI has decided not to release its GPT-6.1 Astra model due to heightened AI safety concerns, with the head of safety systems saying the model “didn’t quite meet the bar" for scope and authorisation.
OpenAI is scrapping the release of GPT-6.1 Astra, a next-generation AI model planned for an October debut, over safety concerns raised by researchers during internal testing, theWall Street Journalreported.The model, expected to appear in ChatGPT and Codex, was designed to handle more complex tasks without human assistance, the report said.Earlier this month, Anthropic CEO Dario Amodei called for the industry to slow the development of frontier AI models to allow safety measures to keep pace, a view endorsed by OpenAI CEO Sam Altman and SpaceX CEO Elon Musk.The ChatGPT parent's safety chief Saachi Jain toldWSJon Monday that Astra fell short of the company's standards in alignment tests, which assess whether a system follows human intent.The model showed more deception than its predecessor, including at times failing to accurately disclose actions it had or had not taken, the report said.It also had problems with "scope authorisation", pushing ahead with tasks without requesting user permission and sometimes attempting to use external tools or services when doing so could be unsafe.The decision comes ahead of OpenAI's developer conference in San Francisco, where the company has previously unveiled products aimed at software developers.It comes months after OpenAI revealed that one of its AI agents had inadvertently hacked Hugging Face.In its latest disclosure,OpenAI saidits agents had leaked 53 images belonging to ChatGPT users.The company did not say whether the images were AI-generated or depicted real people, or when they were posted.Last week,The New York Timesreported that OpenAI’s rogue agents interacted with websites belonging to the US Commerce Department and Securities and Exchange Commission in unusual ways this summer without the company’s knowledge.OpenAI said it had notified the government agencies in recent weeks.The latest incidents add to concerns over autonomous AI agents, which can perform tasks and interact with external systems with limited human involvement, and the challenges companies face in tracking their behaviour.