Bypassing anti-bot protection: the complete guide

Fingerprinting, JavaScript challenges, CAPTCHA, rate limiting: here is how a site detects bots, how to recognize a block and which techniques reduce false positives without breaking the law.

  • Browser fingerprinting: Canvas, WebGL, fonts, resolution, time zone
  • Network fingerprint: TLS signature (JA3), order and consistency of HTTP headers
  • JavaScript challenges: a session cookie or token required before the response
  • CAPTCHA and interactive challenges: reCAPTCHA, hCaptcha, Turnstile
  • Rate limiting and honeypots: request caps, invisible trap links
  • 403 (forbidden) or 429 (too many requests) codes
  • A challenge page (“Checking your browser…”) instead of the expected content
  • Truncated HTML missing the data you see in the browser
  • Redirect loops or growing delays
  • Headless browser with a realistic fingerprint (Playwright + stealth engine)
  • Consistent headers and user-agent, with Accept-Language matching the locale
  • Human-like pacing: randomized delays and limited concurrency
  • Reusing cookies and sessions to keep the challenge token
  • Running JavaScript and waiting for the right selector (not just the load event)
  • Residential or mobile proxies when facing geo-blocks or banned IPs
Why does my scrape return a 403 or 429 error?

A 403 signals a block (fingerprint or headers seen as suspicious); a 429 an exceeded request limit. Slow down, use a realistic-fingerprint browser and check your headers.

Is it legal to bypass anti-bot protection?

It depends on the context. Bypassing a technical measure may violate the terms of use, and some protections are legally protected. Stick to public data and check your obligations (robots.txt, GDPR) before collecting.

Are Playwright or Puppeteer enough to avoid detection?

Not as-is. A default headless browser exposes detectable signals. You must tune the fingerprint, headers and pacing — hence the value of a stealth engine or a specialized service.

Do I need proxies to scrape a protected site?

Not always. Proxies become useful when your IP is blocked, shared or geo-restricted. A residential or mobile proxy pool improves reliability, but increases cost.

Ready to scrape a website?

Run your first scrape in seconds, for free.

Start a scrape