Web scraping and GDPR: the compliance guide

The GDPR does not ban web scraping: it governs how personal data is processed. Here is how to collect public data while complying with EU rules — legal basis, personal data, robots.txt and compliance measures.

  • Name, email, phone number, address
  • Online identifiers, IP address, cookies
  • Location and browsing data
  • Sensitive data (health, opinions, orientation): stricter rules
  • Legitimate interest: the most common (monitoring, research, security)
  • Consent: hard to obtain for automated scraping
  • Legal obligation or public interest: specific cases
  • Mandatory balancing test against individuals’ rights
  • Minimisation: only the data you need
  • Limited, defined purpose
  • Defined and justified retention period
  • Security: encryption, restricted access
  • Informing individuals and right to object (opt-out)
Is web scraping legal?

Scraping is not unlawful in itself: it depends on the data collected and its use. As long as the collection involves non-personal data, or personal data processed with a legal basis and safeguards, it can comply with the GDPR.

Can you scrape public personal data?

Yes, under conditions. Publicly accessible data may be processed on the basis of legitimate interest, provided you balance individuals’ rights and take measures: minimisation, information, right to object and anonymisation at collection time.

Do you have to respect robots.txt?

robots.txt reflects the site publisher’s intent. Respecting it is not a GDPR obligation, but it reduces the risk of contractual dispute and technical blocking. It does not remove the need for a legal basis.

What are the penalties for non-compliant scraping?

The GDPR allows fines of up to €20 million or 4% of global annual turnover. The CNIL, for example, fined Clearview AI €20 million in 2022 for collection without a legal basis.

Ready to scrape a website?

Run your first scrape in seconds, for free.

Start a scrape