Cloudflare: Turnstile, error 1020 and TLS fingerprint
Understand and get past Cloudflare protections: Turnstile, error 1020, TLS fingerprint and headers. Causes, diagnosis and the ScraperFlow approach (realistic browser).
Read the pageSites protect their content against bots: browser fingerprint, CAPTCHA, rate limiting. Here is how these defenses work and how ScraperFlow approaches them, with a realistic browser and a controlled pace.
Anti-bot protections aim to tell human traffic apart from automated traffic. Their goal is not to stop you reading a page, but to contain load, limit mass scraping and protect sensitive data.
They often trigger on thresholds: a spike of requests, a too-regular behaviour or an unusual browser fingerprint is enough to flip a session to the “bot” side.
There are several families of defenses, often combined: request analysis (headers, TLS fingerprint), client-side challenge execution (JavaScript, CAPTCHA) and behavioural analysis (pace, navigation).
Some solutions are generic (application firewall, bot-management service), others are vendor-specific — Cloudflare and DataDome are two examples we detail in the child pages.
A real browser sends a coherent set of information: HTTP headers, request order, JavaScript capabilities, TLS fingerprint. A minimal HTTP client shows a different signature, easy to spot.
Reproducing that coherence requires a real (headless) browser, plausible headers and a standard TLS negotiation — the basis for stable access to modern sites.
When in doubt, many sites show a challenge: an image CAPTCHA or Turnstile (Cloudflare’s often-invisible alternative). The goal is to force a proof of humanity the client must solve.
These challenges live in a JavaScript page: it must run for the token to be computed and returned. A plain `curl` does not get past this step.
Beyond a single request, protections watch behaviour over time: requests per minute, clockwork regularity, absence of pauses. A “machine” pace is a strong signal.
Adapting the pace, spreading the load and respecting realistic pauses reduce the risk of blocking — and spare the source site.
ScraperFlow runs a realistic headless browser, sends coherent headers and waits for JavaScript rendering: content is reached as a visitor would.
The pace is controlled and the load spread. We do not promise a universal bypass: defenses evolve, and some sites remain hard. Our goal is reliable access, respectful of the site and its legal framework.
Accessing public content is not illegal in itself, but automation is framed: respect the site’s robots.txt and terms of use, limit the load, and process personal data under the GDPR. Bypassing protections for sensitive or protected data may, however, be unlawful.
No, and we do not claim it. Anti-bot defenses change constantly and some sites remain hard. ScraperFlow aims for reliable access through a realistic browser and a controlled pace; results vary by site and configuration.
A minimal HTTP request has an easy-to-spot fingerprint (headers, TLS) and does not run JavaScript: it gets past neither fingerprinting nor a client-side challenge. A realistic browser produces a coherent fingerprint and runs the page, which is enough for many sites.
Sometimes. When a site blocks by IP (reputation, geo-restriction, request threshold), varying addresses lowers the risk. On other sites, a realistic browser and a proper pace are enough. The need depends on the defense in front.
Understand and get past Cloudflare protections: Turnstile, error 1020, TLS fingerprint and headers. Causes, diagnosis and the ScraperFlow approach (realistic browser).
Read the pageUnderstand and handle DataDome blocks: 403 responses, behavioural detection, CAPTCHA. Causes, diagnosis and the ScraperFlow approach (realistic browser, controlled pace).
Read the page