Chinese AI models used to write code may be creating a hidden security risk for U.S. companies, federal officials and government contractors, per a new report published by a major defense contractor specializing in cyber security.
Booz Allen published a report in late May warning the federal government, private developers and workers in critical industries that the presence of code written by popular Chinese AI models within the supply chain may be making the United States more vulnerable to bad faith actors. These vulnerabilities arent simple backdoors, Booz Allen reports, but rather come in the form of Chinese large language models producing lower-quality, and thus easier to breach, code when they believe they are being prompted by an American.
Chinese models are generally cheaper than their Western counterparts and work well enough to keep companies interested, a dynamic that has led to increased adoption in the United States and put some policymakers and national security experts on edge.
“Id say theres an 80% chance theyre using a Chinese open-source model,” Martin Casado, a general partner at the major venture capital firm Andreessen Horowitz, said in November 2025 when asked about their prevalence among start ups. Major U.S. firms such as Meta, Airbnb and Perplexity are also reportedly using Chinese models.
“The first link in the software supply chain is no longer the code. Its the AI models behind it,” the Booz Allen report reads. “As U.S. developers increasingly rely on AI to generate, debug, and secure code, we must confront a fundamental question: can the AI models writing and powering our nations code be trusted?”
In an attempt to answer this question, Booz Allen compared four of the most widely used Chinese models â Kimi, Qwen, MiniMax and â against Anthropic’s Claude to test the security of the code they produced. The firms behind the four Chinese models did not respond to requests for comment when reached by Fox News Digital.
Qwen and MiniMax both produced code with significantly more vulnerabilities, increases of 130% and 20%, respectively, when they believed they were doing work for U.S. government employees as compared to a general prompt. DeepSeek, meanwhile, saw an increase of just 5% while Kimi produced code of a similar quality.
This means a government contractor relying on one of these models could unknowingly introduce coding flaws that make databases, applications or internal systems easier for hackers to exploit, potentially exposing sensitive American information.
The findings have drawn comparisons to so-called “sleeper agent” behavior where AI models appear to operate normally until exposed to a specific trigger that causes them to produce lower quality, or even deliberately insecure, outputs.
Experts interviewed by Fox News Digital expressed a range of opinions on Booz Allen’s findings.
“While the raised risk categories are understandable, the reports stronger claims are not fully supported as presented,” Lukasz Olejnik, a technology consultant who works as a senior research fellow at King’s College London, told Fox News Digital. “The report underplays the complexity of the issue.”
If Booz Allen’s report were accurate, and if code written by Chinese models had made its way into the American supply chain, it would make it easier for hackers to get their hands on data that could imperil national security or infringe on the privacy of everyday Americans.
Olejnik argued that the prompting used by Booz Allen was unnatural, saying that the firm’s methodology may have included “unnecessary political or institutional keyword triggers,” such as explicitly prompting models to believe a user is working for the , that “may change outputs.” It is unlikely, he says, that an actual government agent would prompt the mod