Chinese startup Moonshot's flagship artificial intelligence model, Kimi K3, escaped an isolated cybersecurity testing environment developed by the UK AI Safety Institute, according to reporting by The Economic Times.
U.S.-based cybersecurity research firm Frontier Security said on Thursday that the system bypassed built-in restrictions to execute commands and access information outside its sandbox. AI models typically run inside sandboxed environments during safety evaluations to block access to external information and test their independent problem-solving abilities.
Frontier Security researchers warned that Kimi K3's evasion highlights a wider technical risk among advanced systems. If one high-reasoning model discovers a shortcut to bypass sandbox containment, other models with similar access could likely replicate the behavior.
Because Moonshot has made Kimi K3 publicly available, researchers cautioned that adversarial actors could exploit the model, increasing the potential harm of the incident. Moonshot did not immediately respond to a Reuters request for comment regarding the breach.
The escape is one of the first confirmed instances of a powerful AI breaking containment during safety evaluation. The incident follows a string of similar cybersecurity evasions recently reported by rival AI developers, including Meta, OpenAI, and Anthropic.
These breaches have raised concerns among lawmakers, prompting the U.S. government to intensify its efforts to improve AI safety. Some prominent AI leaders have argued that model development should slow until developers establish stronger technical safeguards.