U.S. AI Safety Institute Claims China's Kimi K3 Trails U.S. Frontier Models in Cybersecurity Tests

Kimi K3 performed significantly below leading frontier models when tasked with developing software exploits and conducting autonomous attacks on simulated enterprise networks.

Share
U.S. AI Safety Institute Claims China's Kimi K3 Trails U.S. Frontier Models in Cybersecurity Tests
(Image-Freepik)

When Chinese AI company Moonshot released its much-anticipated model Kimi K3, it claimed that the model performs on a par with leading models in the US, such as OpenAI's GPT and Anthropic's Claude. Their claims were based on third-party evaluations from Artificial Analysis and Arena.ai.

The launch comes amid intensifying AI competition between the U.S. and China, where advances in frontier models are increasingly viewed through the lens of national security and technological leadership.

As both countries race to dominate AI, the capabilities and training methods of Chinese models have come under heightened scrutiny in Washington. Soon after the release, U.S. politicians have accused Chinese AI companies of model distillation, an AI training technique where smaller models learn from the outputs of larger models.

"We support open-source AI and the innovation it unlocks. But open source is not open season on American IP. When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table," Treasury Secretary Scott Bessent posted on X.

However, experts are skeptical that distillation alone explains Kimi K3’s rapid progress. “You can’t distill that much data, train a model, and release it in two weeks,” Braden Hancock, Snorkel AI Co-Founder, told TechCrunch.

Nathan Lambert, Researcher at the Allen Institute, added that reinforcement learning, rather than supervised fine-tuning through distillation, is increasingly driving frontier AI advances.

Kimi K3 Falls Behind Frontier AI in Cybersecurity Tests

Now, new research from the UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) have released the results of a joint cybersecurity evaluation of Moonshot AI’s Kimi K3, concluding that the newly launched model lags behind the latest frontier AI systems in offensive cyber capabilities despite outperforming other open-weight models.

The assessment examined Kimi K3, which was released on July 16 and is expected to become open-weight by July 27, across multiple cybersecurity benchmarks, including exploit development and simulated network attacks.

According to the report, Kimi K3 performed significantly below leading frontier models when tasked with developing software exploits and conducting autonomous attacks on simulated enterprise networks. However, it surpassed GLM-5.2, previously considered the most cyber-capable open-weight model.

On ExploitBench, a benchmark developed by Carnegie Mellon University to evaluate end-to-end exploit development, Kimi K3 achieved a 32% success rate, compared with 24% for GLM-5.2. The benchmark measures a model's ability to exploit 41 recent vulnerabilities in Google's V8 JavaScript engine.

(Preliminary comparison of aggregate capabilities over time of the most capable U.S. and PRC models as of Kimi K3’s release)

Stronger Than Open Models, But Still Behind U.S. Leaders

Despite the improvement over GLM-5.2, Kimi K3 failed to achieve arbitrary code execution (ACE)—the most severe exploit outcome that allows attackers to compromise a target fully. The model recorded 0 successful ACE exploits across 41 tasks, while the most capable frontier models averaged 20.

Researchers also evaluated the model using "The Last Ones" (TLO), a 32-step simulated corporate network attack designed to measure autonomous cyberattack capabilities. Kimi K3 reached an average of 17 steps, compared with 28.5 steps for the strongest U.S. frontier models. GLM-5.2 averaged 11 steps under the same evaluation.

In one of ten attempts, Kimi K3 completed the full TLO attack path within the benchmark's 100-million-token limit, demonstrating that it can autonomously attack a small, vulnerable enterprise network when provided with initial access.

Researchers cautioned, however, that the simulated environment lacks real-world defensive measures such as active security teams, detection tools, and unpredictable network conditions.

The evaluation also found that Kimi K3's safeguards did not prevent it from assisting with agentic cyber exploit development or offensive cyber operations during testing, allowing the model to attempt exploit generation when prompted.

The report notes that the findings are based on preliminary evaluations conducted across a limited number of public and private benchmarks. Due to Kimi K3's hosting setup, researchers were only able to run a selective set of cyber evaluations, resulting in wider confidence intervals than those reported for other frontier models.