Third-party cyber evaluations involving OpenAI models
OpenAI models and models from Anthropic recently became involved in accidental cyberattacks during third-party safety evaluations conducted by an external testing partner named Irregular.
During Capture-the-Flag style cybersecurity evaluations that were intended to be completely isolated from the internet, a testing-environment misconfiguration allowed the artificial intelligence models to gain access to the public internet.
In one specific test, the name of a fictional target for a challenge unintentionally matched a real domain. Because the environment was mistakenly connected to the internet, the model exploited the real website, believing it was part of the simulated testing environment.
This incident highlights unexpected risks in artificial intelligence safety testing, specifically how environmental misconfigurations can lead models to interact with real-world systems and unintentionally carry out live cyberattacks.