All Publications
Academic PaperMay 19, 2023

Governing Human-Derived Data

Executive Summary

Platforms collect vast amounts of data on users for commercial and operational purposes, including targeted advertising, and this data is often shared with third parties or scraped by researchers, civil society, commercial competitors, and data brokers. A notorious example is Clearview AI, which scraped online photographs to build its facial-recognition database. Questions about the legitimacy of collecting, using, or sharing such data usually turn on whether it concerns an identifiable person (personal data) or has been anonymized. This essay argues that the personal-versus-anonymized binary is no longer sufficient to govern data practices.

Data protection and privacy laws generally regulate information about identifiable persons, resting on individual control over one's own data; anonymized data is often placed outside their scope, as in the EU's GDPR, Canada's Bill C-27 (Consumer Privacy Protection Act), and Ontario's health privacy law. This creates a governance gap. Reidentification remains possible given current data volumes and analytic tools, group-level privacy interests go unrecognized, and even anonymized data can produce biased AI outputs and collective harms.

The author proposes a distinct category, "human-derived data" — data derived from humans or their activities that is not personal data (for example, COVID-19 wastewater surveillance). Anonymized data is only a subset of it. Its governance should rest on fundamental human rights rather than individual privacy, incorporating non-discrimination, ethics, transparency, public engagement, open access to results, and benefit-sharing to communities.

Related Publications