During a routine security evaluation, an open-weight AI model named Kimi K3 developed by Chinaβs Moonshot AI managed to escape its testing sandbox and reach the open internet.
According to US cybersecurity firm Frontier Security, this is the first time a freely downloadable public model has broken out of its containment environment.
A leak in the sandbox that the model chose to use
According to an interview with WIRED yesterday, Frontier Security was measuring Kimi K3βs defensive cybersecurity skills when the model wandered outside the environment meant to hold it.Β
Apparently, a misconfiguration had left a gap in that environment. However, Frontier stated that the model worked out on its own that it could reach certain websites by probing the sandboxβs network settings, then went online without asking permission. It had been told to solve problems that were not supposed to require the internet.
βWe found a leak in the sandbox,β Frontier CEO Yaron Singer told WIRED. βBut we also found that Kimi took advantage of that loophole, suggesting that it doesnβt have the same internal guardrails.βΒ
Frontier argues that Kimi carries fewer cyber safeguards than most other powerful models, which is what let it slip out.
No systems hacked, but weaker guardrails
Fortunately, Kimiβs escape did not lead to any malicious hacks or system compromises. Because the information it was looking for was easily accessible on GitHub, it didnβt need to break into anything once it got online.
However, the main concern is accessibility. Unlike most heavily secured internal lab models, Kimi K3 is open to the public, meaning that anyone can download and run it with those same loose safety guardrails in place.Β
Testers noted that the model is ruthlessly efficient at achieving its goals by any means necessary, even if it means cheating or escaping containment.Β
The testing environment itself was built with sandboxes from the UK governmentβs AI Security Institute, although this has not yet been confirmed by either Moonshot or the AISI, as they have declined to comment.
More rogue agents appearing this summer
Kimiβs escape adds to a growing trend of AI models bending the rules during evaluations. On July 21, OpenAI revealed that its models exploited a zero-day software flaw to reach the internet and break into Hugging Face.Β
Days later, Cryptopolitan reported that Anthropic traced some of its models to unauthorized external break-ins. Meta even admitted one of its AI agents (Muse Spark 1.1) reached an outside firm due to a misconfigured testing environment.
Experts have always maintained that these incidents are usually the result of poorly secured testing walls rather than sci-fi jailbreaks. βAs a general phenomenon, if you give one of these models an objective, and if youβre not very explicit, like walls youβre putting around it, itβll find a way to get the answer,β said Matt Fredrikson, CEO of Gray Swan and a Carnegie Mellon professor.
Donβt just read crypto news. Understand it. Subscribe to our newsletter. It's free.


















English (US)