Technology

Chinese AI models gave researchers bioweapons guidance after safety bypass

Posted on

An illustration representing artificial intelligence. [mikemacmarketing/Wikimedia Commons]

Chinese artificial intelligence developer Moonshot is conducting an internal review after researchers persuaded two of its Kimi models to provide information about biological weapons and assassinations, the BBC reported Wednesday.

In July, AI security testing company Mindgard told the BBC that it had found Kimi K2.6 and K3 Swarm could bypass safety measures designed to stop them talking about dangerous topics.

The discoveries were made during so-called “jailbreaking” tests, in which researchers use complex strings of instructions to trick AI systems into disobeying their safety controls. Mindgard did not disclose the key technical details of how the safeguards were breached, it said.

Following success with the jailbreak, the models could talk about a wide range of harmful topics and offer additional suggestions, Mindgard founder Peter Garraghan told the BBC.

“Once the jailbreak works it will talk about any topic,” said Garraghan.

Moonshot told the BBC it welcomed third-party feedback “as a key pillar for building better and safer AI” and said it was discussing the findings with Mindgard. Internal reviews generally found a “high refusal rate” for such requests, the company said.

Mindgard stated that it has not determined if the information provided by the models will be useful in practice.

Click to comment

Most Popular

Exit mobile version