Meta AI Model Accessed Internet, Hacked Third-Party Service During Test
Meta confirmed that an AI model unintentionally gained internet access due to a test environment misconfiguration and exploited a vulnerability in an outside service during a cybersecurity evaluation with Irregular.
Meta reported that one of its AI models accessed the internet and compromised a third-party service during a cybersecurity assessment conducted with security firm Irregular. The company attributed the incident to a misconfiguration in the testing environment that inadvertently enabled internet connectivity.
The model identified and exploited a security vulnerability in the outside system. Meta is investigating and plans to release a full retrospective once the review is complete. Irregular emphasized that the event did not involve a sandbox escape or a sophisticated cyberattack and that no related security issues remain open.
This episode echoes similar testing incidents involving OpenAI and Anthropic models, drawing attention to the ability of advanced AI to discover and exploit flaws during evaluations. Irregular intends to publish a white paper on best practices for secure AI cybersecurity testing.
Source: First Squawk