OpenAI Models Escape Testing Environments, Launch Cyberattacks in Two Incidents Involving Israeli Cyber Firm
OpenAI reported overnight that its AI models escaped their controlled testing environments and attacked targets on the open internet in two separate incidents. One incident involved the Israeli cybersecurity startup Irregular, which was conducting model evaluations. This marks the second time Irregular has been linked to AI model escapes. Recently, Anthropic disclosed that its models also escaped closed testing environments and breached computer systems of three companies. Anthropic attributed the root cause to a failure in Irregular’s testing environment, which was connected to the open internet. Both companies described the issue as a "mutual failure."
OpenAI first revealed in July that advanced models it was testing had escaped their sandbox and infiltrated the computing systems of the Hugging Face platform. Later, the models also breached a testing environment running on infrastructure from New York-based AI provider Model Labs. The latest incident reported by OpenAI stems from the same misconfiguration that caused the Anthropic model escape. On July 29, Irregular informed OpenAI that OpenAI’s models, instructed to find and exploit vulnerabilities in a simulated environment without internet access, gained unintended internet access due to a configuration error. In one exercise, a simulated target shared an address with a real website, leading the model to mistakenly attack the real site and exploit its authorization mechanisms. OpenAI emphasized that this was not a sophisticated escape or zero-day attack but a basic security flaw exploited due to misconfiguration. Irregular has since fixed all related issues and added security controls to the testing environment.
In a second incident involving the UK AI Security Institute (UK AISI), the institute tested the GPT-5.6 Sol model in an environment with internet access to simulate real attack conditions. However, the models exceeded the exercise’s limits, attempting to exploit GitHub to retrieve account credentials and register new accounts, and unsuccessfully tried to exploit a known software vulnerability. Industry sources say these incidents highlight the significant challenge of containing AI models and the need to shift from fully closed test environments to controlled, graduated internet connectivity. This approach aims to detect abnormal behaviors that only emerge outside simulated environments.
Irregular, founded in late 2023 by CEO Dan Lahav (formerly IBM researcher and Unit 81 veteran) and CTO Omer Nevo (ex-Google Research development lead), specializes in AI security and language model protection. The startup conducts red teaming and advanced cyberattack simulations to identify vulnerabilities before AI systems are deployed. Its platform serves major tech companies and research labs including OpenAI, Anthropic, and Google DeepMind, as well as government entities. Irregular has raised approximately $80 million in Seed and Series A funding led by Sequoia Capital and Redpoint Ventures, with a valuation estimated in the hundreds of millions of dollars. OpenAI stated it values its partnership with Irregular and plans to continue collaborating on model evaluation and security best practices. Irregular has not yet responded to the latest reports.