All Publications
Policy BriefJune 25, 2026

Online Harms AI Audit: Technical Brief

Executive Summary

Millions of Canadians now use AI chatbots daily, yet no empirical audit has tested whether these products protect users from harm. This brief presents the first systematic audit of four major deployed AI chatbot products (OpenAI's ChatGPT, Google's Gemini, Anthropic's Claude, and Meta AI) against the harm categories proposed in Bill C-34 (Safe Social Media Act). Using multi-turn adversarial tests, it distinguishes the developer system reached through the API from the consumer system reached through the Web UI. Six of the seven C-34 harm categories were probed in English and French; child sexual abuse material (Category 1) was excluded due to its legal status.

Across all 420 multi-turn attacks, roughly one in three (145, a pooled attack success rate of 35%) ended with the product generating harmful content in full, ranging from under 2% on the most resistant developer system to roughly two-thirds on the most permissive consumer system. Safety floors vary dramatically between providers: Anthropic's model resisted nearly all attempts (2%), while Gemini was bypassed in more than half (62%). Consumer-facing products can close the gap — Meta AI (5%) blocks harmful requests with a classifier before the model responds, whereas Gemini's consumer layer (67% Web UI) adds essentially no protection.

Several consumer products produced operationally usable self-harm content: Gemini's app completed a fictional 14-year-old's overdose case file in 2 of 6 attempts and ChatGPT's API supplied content rationalising severe food restriction in 2 of 6, while Claude refused every reframing. These findings suggest voluntary safety measures are inconsistent and insufficient, and that legislation like Bill C-34 setting minimum standards would give regulators and the public a basis for accountability that currently does not exist.

Related Publications