Daijiworld Media Network - San Francisco
San Francisco, Aug 8: OpenAI has paused some internal work on its upcoming artificial intelligence model Astra to strengthen safety measures after the system demonstrated significantly advanced capabilities in cybersecurity tasks.
The ChatGPT maker said on Friday that it “cannot rule out” Astra reaching its “critical cybersecurity threshold”, a level at which an AI system could potentially identify and develop zero-day exploits without human intervention.
OpenAI said it was strengthening security controls for the development and testing of newer models and pausing internal activities involving Astra that do not yet meet the enhanced requirements.

Chief Executive Officer Sam Altman said the company is working to make Astra generally available but acknowledged that its advanced cyber capabilities meant additional safety work was necessary.
“Given its cyber capabilities, we need a little longer to do this safely,” Altman said in a social media post, adding that the delay would hopefully not be too long.
The move comes amid growing concerns over the ability of increasingly capable AI agents to operate autonomously and interact with computer systems in ways that developers may not always anticipate.
Over the past two weeks, OpenAI and Anthropic have publicly acknowledged that their AI systems inadvertently breached the systems of several institutions, including Hugging Face, while the companies were testing their models.
Meta also said on Wednesday that a recently released AI model had infiltrated the computer system of a third party.
The incidents have provided fresh evidence of the potential risks posed by AI agents capable of independently identifying vulnerabilities and taking actions within computer environments.
OpenAI said it will work with government agencies and AI safety organisations to test Astra's capabilities. The company also plans to provide recommendations to third-party testing partners on how to safely evaluate its more advanced models.
The additional safeguards are aimed at ensuring that Astra can be tested and eventually deployed without creating unacceptable cybersecurity risks.