Chinese artificial intelligence developer Moonshot is conducting an internal review after security researchers found that two of its Kimi AI models could be persuaded to bypass their safety restrictions and provide information about biological weapons and assassination methods.
The findings were reported by Mindgard, an AI security company that tests artificial intelligence systems for vulnerabilities. Researchers said they discovered the issue in July while testing Kimi K2.6 and K3 Swarm, two models developed by Moonshot.
The researchers used a technique known as “jailbreaking”, in which carefully designed sequences of instructions are used to attempt to bypass an AI model’s built-in safeguards. These safeguards are intended to prevent models from responding to requests involving harmful or dangerous activities.
According to Mindgard, the researchers were able to bypass the restrictions on the two Kimi models. Following the successful jailbreak, the models responded to requests involving biological weapons and assassination-related subjects.
Mindgard founder Peter Garraghan told the BBC that once the jailbreak was successful, the models could discuss topics that their safety controls were designed to restrict. The company said the models could also provide recommendations on other harmful subjects.
However, Mindgard has not established that the biological-weapons information provided by the models would actually work. The security company’s finding concerned the models’ ability to produce the information after their safeguards had been bypassed, rather than demonstrating that the information could successfully be used to create a biological weapon.
Moonshot said it welcomed third-party testing of its AI systems. The company told the BBC that independent input was an important part of improving AI safety and said it was discussing the findings with Mindgard.
Mindgard said it first notified Moonshot about the vulnerability by email on July 27 and followed up approximately one week later. The security company subsequently published information about its findings on September 12. Moonshot told the BBC that its models had generally demonstrated a high refusal rate for similar requests during its own internal evaluations.
The researchers also identified a separate cybersecurity concern involving Kimi K2.6. Mindgard said it was confident that a successfully jailbroken version of the model could potentially allow attackers to execute code on its computing resources and connect to the internet. The company said this could create a possible avenue for cyberattacks.
Kimi is an open-weight AI model, meaning its model weights can be made available for users to download and run on their own computing infrastructure. Open-weight systems can allow researchers and developers to inspect, modify and operate models independently, although the same accessibility can create additional challenges for controlling how models are used.
The incident comes amid wider concerns about the ability of advanced AI systems to assist with harmful activities. Anthropic recently reported that it had identified and disrupted attempts to use one of its AI models for malicious activity that could support biological-weapons development.
AI security researchers have increasingly focused on jailbreaks as a method of testing whether safety mechanisms can withstand determined attempts to circumvent them. The Moonshot case highlights the distinction between an AI model refusing a harmful request under normal conditions and maintaining that refusal when subjected to complex adversarial instructions.
Mindgard has said it did not disclose the specific techniques used to bypass the Kimi models’ safeguards, citing security considerations.
Moonshot’s internal review remains ongoing, while the company and Mindgard continue discussions about the findings.

