Web scraping and GDPR: the compliance guide
The GDPR does not ban web scraping: it governs how personal data is processed. Here is how to collect public data while complying with EU rules — legal basis, personal data, robots.txt and compliance measures.
Is web scraping allowed under the GDPR?
The GDPR does not regulate the collection technique itself, but the processing of personal data. Scraping web pages is therefore not unlawful per se: compliance depends on what you collect and what you do with it.
As soon as scraping involves personal data, it becomes a "processing" activity subject to GDPR obligations: a legal basis, informing individuals, respecting their rights and security.
Personal data or anonymous data?
Personal data is any information relating to an identified or identifiable natural person: a name, an email, a phone number, but also an IP address or a cookie identifier.
Data that is truly anonymised — irreversibly — falls outside the GDPR. Pseudonymisation, however, remains subject to the regulation, because the individual can still be re-identified.
- Name, email, phone number, address
- Online identifiers, IP address, cookies
- Location and browsing data
- Sensitive data (health, opinions, orientation): stricter rules
What legal basis applies to scraping?
Every processing of personal data must rely on a legal basis. For scraping, legitimate interest is the most common: it requires a balancing test between your interest and individuals’ rights.
Consent is rarely workable for automated collection; legal obligation or public interest apply only in specific cases. The CNIL has published a dedicated sheet on legitimate interest for web scraping.
- Legitimate interest: the most common (monitoring, research, security)
- Consent: hard to obtain for automated scraping
- Legal obligation or public interest: specific cases
- Mandatory balancing test against individuals’ rights
robots.txt, terms of use and contract law
The robots.txt file and a site’s terms of use express the publisher’s intent. Ignoring them may trigger contractual liability, independently of the GDPR.
Respecting these signals reduces legal and technical risk (blocking) — though they never replace a legal basis under the GDPR.
Compliance measures to implement
A compliant scrape applies the GDPR principles: collect only what is necessary, limit the purpose, set a retention period and secure the data.
The CNIL further recommends limiting collection to freely accessible data, informing individuals and providing a right to object.
- Minimisation: only the data you need
- Limited, defined purpose
- Defined and justified retention period
- Security: encryption, restricted access
- Informing individuals and right to object (opt-out)
DPIA, records and data subject rights
Some large-scale processing (profiling, monitoring) requires a data protection impact assessment (DPIA) and a record of processing activities.
Individuals keep their rights: access, rectification, erasure and objection. A compliant scrape must be able to honour them, which means tracking the sources of the data collected.
Case law, sanctions and best practices
European authorities sanction collection without a legal basis: in 2022, the CNIL fined Clearview AI €20 million for collecting and using data without a legal basis.
Best practices: document the legitimate-interest balancing test, limit collection to freely accessible content, anonymise or pseudonymise at collection time and keep a record of sources.
Frequently asked questions
Is web scraping legal?
Scraping is not unlawful in itself: it depends on the data collected and its use. As long as the collection involves non-personal data, or personal data processed with a legal basis and safeguards, it can comply with the GDPR.
Can you scrape public personal data?
Yes, under conditions. Publicly accessible data may be processed on the basis of legitimate interest, provided you balance individuals’ rights and take measures: minimisation, information, right to object and anonymisation at collection time.
Do you have to respect robots.txt?
robots.txt reflects the site publisher’s intent. Respecting it is not a GDPR obligation, but it reduces the risk of contractual dispute and technical blocking. It does not remove the need for a legal basis.
What are the penalties for non-compliant scraping?
The GDPR allows fines of up to €20 million or 4% of global annual turnover. The CNIL, for example, fined Clearview AI €20 million in 2022 for collection without a legal basis.