Technology

Chinese AI agents display deception and safeguard bypasses in tests

Posted on

An illustration representing artificial intelligence. [mikemacmarketing/Wikimedia Commons]

Chinese-powered AI agents have exhibited deception, attempts to circumvent safeguards and efforts to hide failures in controlled safety tests, behavior that researchers say mirrors warning signs seen in US AI systems, according to Reuters.

Reuters reviewed over 200 research papers and technical reports and found at least 20 studies since 2025 reporting such behavior. The review found no evidence that Chinese-powered agents had escaped independently to the wider internet or evaded shutdown.

In one experiment, agents using Alibaba, DeepSeek and Moonshot models falsely claimed capabilities when competing in a simulated business tender. False claims appeared in 88% of the sessions involving Alibaba’s Qwen3-Max-Preview, 84% for DeepSeek-V3.2-Exp and 88% for Moonshot’s Kimi-K2, the study found. US the models included in the test showed similar behavior.

“These are the same warning signs US labs are seeing, in less capable systems,” Alex Mallen of Redwood Research told Reuters.

Other tests found agents faking results or fabricating files, instead of admitting to failures of tasks.

The experiments were controlled, according to the researchers and they did not show an AI system breaking out into the web. Growing concerns about increasingly autonomous AI agents have Chinese companies and regulators working on additional safety measures.

Click to comment

Most Popular

Exit mobile version