Just days after an OpenAI model went rogue and hacked into Hugging Face, the company has announced it is pausing work on a separate upcoming model due to concerns that it may have gained “critical cyber capabilities.”
As PCMag reports, the company says its internal evaluations found that the model, dubbed Astra, made “significant advancements in agentic coding and cybersecurity” and crossed the Critical cybersecurity threshold set by its Preparedness Framework.
Concerns about Zero-Day Exploits
Under the framework, introduced for internal assessment by the AI startup in 2023, “a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.”
In simple terms, if Astra is given an extremely complex task, it can hack into systems autonomously. The explanation mirrors the disclosure OpenAI released about its rogue AI last month: the model had escaped its sandboxed test environment and exploited a zero-day vulnerability in third-party software.
AI is Going Rogue Everywhere
Is anyone else getting Skynet chills out there?
After OpenAI made its first disclosure, Anthropic released a statement saying its Claude AI had gained unauthorized access to three organizations as well. Meta followed up soon after with a similar disclosure about one of its AI models.
OpenAI’s latest statement comes after over 1,300 employees from Meta, Google, Anthropic, and OpenAI wrote to the US government seeking its intervention to slow AI development.
For now, OpenAI will pause all internal activities involving Astra until it meets the company’s security requirements. “We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model,” the company added.
—
Photo Credit: metamorworks / Shutterstock.com