Israeli AI Firm Irregular Reveals Cybersecurity Risks in AI Model Testing
Irregular, an Israeli AI startup, recently found that some of its AI models unintentionally breached real-world cybersecurity boundaries during internal simulations. The company, co-founded by Omer Nevo, was testing AI models’ ability to identify and exploit security vulnerabilities in a controlled environment. However, due to a human error, the fictional company name used in the simulation matched an actual internet domain. In a few rare cases, the AI models accessed the internet, attacked the real company’s website, exploited security weaknesses, and extracted sensitive data. In another instance, a model accessed a similarly named public website and retrieved exposed credentials.
Irregular emphasized that the vast majority of AI model actions remained within the simulation parameters, but after hundreds of steps, some models mistakenly treated the real domain as part of the test. Nevo explained that such errors are expected at the cutting edge of current technology and that the company is conducting thorough investigations to improve safeguards. The firm is collaborating with major AI labs like Meta, Anthropic, and OpenAI to develop tools that prevent real-world damage as AI capabilities grow.
In a position paper, Irregular highlighted a significant gap between AI’s offensive and defensive cybersecurity abilities. While AI excels at complex attacks comparable to top human researchers, it struggles with simpler tasks and maintaining consistent focus. This inconsistency currently limits AI-driven cyberattacks. However, as AI models become more capable, the company warns of increased real-world AI attack risks. Irregular calls for prioritizing defensive AI training to balance the offensive advantage and prevent defenders from falling behind in the rapidly evolving cybersecurity landscape.