Tests raise fears open AI models could be ‘used at scale for harm’

Tuesday, August 25, 2026

A testing tool built by researchers at Waterloo Engineering shows that leading open-weight artificial intelligence (AI) models are alarmingly vulnerable to malicious tampering.

The researchers collaborated with experts at FAR.AI, a non-profit AI security research group, to rigorously test 21 of the most popular open-weight large language models (LLMs) and found their built-in safeguards could all be defeated.

The results create concerns that models could be used to wage mass disinformation campaigns, create sophisticated email scams, produce step-by-step instructions to make hazardous chemicals or do other widespread damage. 

“When the safety guardrails are stripped out of a capable model, it can be used at scale for harm in ways a single person could never manage manually,” said Dr. Sirisha Rambhatla, a professor of management science and engineering at Waterloo. 

To test a cross-section of open-weight AI models, the international research team – which included contributors in Canada, the United States and Switzerland – built an open-source tool called TamperBench as a standardized way to simulate a variety of different attacks.

Go to Major security weaknesses found in leading open AI models for the full story.