The same question has emerged twice in recent days at two of the world’s leading artificial intelligence companies: how do you safely test an AI system that is becoming increasingly capable of acting like a human hacker?
The answer has brought an Israeli cybersecurity startup into the center of the global AI safety conversation.
Irregular, an Israeli company specializing in AI security testing, was involved in cybersecurity evaluations for both Anthropic and OpenAI that revealed weaknesses in the way advanced AI systems are tested. But the incidents were not separate failures, nor were they caused by an AI model independently escaping its restrictions.
Rather, they exposed the broader challenge of creating testing environments that are realistic enough to reveal dangerous capabilities while still preventing models from affecting real systems.
OpenAI disclosed on Tuesday that one of its models accessed the public internet during cybersecurity evaluations conducted by Irregular. The incident originated from the same testing environment previously disclosed by Anthropic, which said several Claude models accessed real-world systems during testing conducted by Irregular.
Both companies said the incidents resulted from a technical issue in the testing infrastructure rather than a model deliberately breaking out of its controls.
The incident involved so-called Capture-the-Flag cybersecurity exercises, in which AI models are instructed to identify vulnerabilities and retrieve hidden information inside simulated networks.
According to Anthropic, the company reviewed 141,006 interactions in which Claude models could potentially access the open internet and identified three incidents where models operating in Irregular’s testing environment reached outside the intended simulation and accessed active infrastructure belonging to three organizations.
The issue centered around a specific technical mistake: a fictional target used in a cybersecurity challenge was given the same name as an existing internet domain.
Because the simulation environment had been inadvertently connected to the public internet, the model interpreted the real website as part of the exercise and attempted to compromise it.
Anthropic said that in one case, Claude Opus 4.7 identified vulnerabilities, including weak passwords and unauthenticated endpoints, and gained access to credentials and database information belonging to the real organization.
The company said another incident involved an internal research model that stopped its own activity after recognizing it had reached a real organization rather than a simulated target.
Anthropic described the incidents as a “harness failure”, a problem with the systems surrounding the model, rather than an “alignment failure” in which an AI system independently overcame its intended restrictions.
OpenAI’s disclosure involved the same environment. The company said Irregular was conducting cybersecurity evaluations where models were instructed to find hidden information inside simulated environments and were told they did not have internet access.
However, a misconfiguration allowed the models to reach the public internet. OpenAI said the model interacted with a real website because it believed it was part of the simulated environment.
The company emphasized that the incident did not involve a sophisticated sandbox escape or a previously unknown vulnerability.
In its disclosure, OpenAI did not place responsibility on Irregular. Instead, the company highlighted the importance of collaboration with external evaluators and said the incident demonstrated why the industry needs stronger shared standards.
OpenAI said it is reviewing its approach to third-party testing, including how companies define evaluation boundaries, manage internet access, handle credentials, monitor activity, and establish procedures for stopping tests when unexpected behavior occurs.
The company said it will work with AI labs, independent evaluators, national AI institutes, and other organizations to strengthen practices for high-risk evaluations.
The technical issues that led to the unintended interactions have since been fixed and additional protective measures have been implemented to prevent similar cases.
Irregular’s growing role reflects the emergence of a new category of cybersecurity companies focused specifically on artificial intelligence.
Founded in late 2023, the company specializes in AI red teaming, resilience testing, and advanced cyberattack simulations designed to identify vulnerabilities and unexpected behavior in AI systems.
Its platform is used by leading AI developers and research organizations, including OpenAI, Anthropic, and Google DeepMind, as well as government agencies.
The company has raised approximately $80 million in funding from investors including Sequoia Capital and Redpoint Ventures, at a valuation reportedly in the hundreds of millions of dollars.

