Online Harms AI Audit: Policy Memo
Résumé analytique
Between May and June 2026, as part of a chatbot audit for online harm, the authors put the same request to four of the most widely used consumer AI chatbots and got four different levels of protection: Google's Gemini supplied harmful content, Anthropic's Claude refused every version, Meta's chatbot blocked the conversation before it could start, and OpenAI's ChatGPT fell in between. Testing the four products in sustained, multi-turn conversations against the harms Canada's online safety law addresses, roughly one in three attacks ended with the product producing or helping to produce harmful content. Across six of the seven harms the law names, these products could be induced to provide the means to cause harm.
On 10 June 2026 the government tabled the Safe Social Media Act, Bill C-34, the first Canadian bill to impose direct safety duties on AI chatbots, tabling the online-safety promise made in the 4 June "AI for All" national strategy. The audit confirms why governing these systems is warranted, but also complicates the task: how far each product resists is a choice its maker has made, with little disclosure and wide variation.
The average matters less than the range it hides. The same attack one product refused almost every time, another answered roughly two times in three, a fortyfold gap between products. Four factors explain this: the product, not the technology, decides which protection a user receives; safety lives in different layers (in Claude's model, in Meta's product classifier); the developer API is a thin, permissive floor that everything downstream inherits; and safeguards are invisible and unstable, with none of the four checking a user's age and a single Gemini update leaving overall failure roughly unchanged while moving individual harms in opposite directions. The audit did not find chatbots inherently unsafe — some were already safe against attacks that defeated others — showing protection is a choice a duty regime exists to correct. Canada can reach only the two product layers it can regulate (the consumer app and the developer system), since the model is trained abroad; C-34's choice to regulate the deployed system is likely right, but cannot be the whole answer.
Publications connexes

Online Harms AI Audit: Technical Brief
Lire la suite
Survey on Canadians' Preference for Social Media Age Verification Policies
Lire la suite
The Federal Government’s Proposal to Address Online Harms: Recommendations for Children’s Safety Online
Lire la suite
