How to scrape a website: the complete guide
Scraping a website means automatically collecting its content (HTML, data, assets). Here is the step-by-step method, from a simple request to AI-ready extraction.
What is website scraping?
Scraping (or crawling) means automatically downloading the content of a web page and extracting the useful data: text, prices, images, links or structure.
A single-page scrape differs from a crawl (several linked pages) and from full site archiving (the whole site, assets included).
Step 1 — Define the target and scope
Start from a seed URL and delimit the scope: one page, one section, or the whole site.
Set a page limit so you never archive an entire site by accident.
- Single page: one-off extraction
- A section: category, blog, catalog
- Whole site: full archiving + assets
Step 2 — Fetch the HTML (server-side or browser)
Many modern sites build their content with JavaScript, so a plain HTTP request returns an almost empty page.
A headless browser — Playwright, Puppeteer — runs the JavaScript and waits for load to obtain the real HTML.
Step 3 — Extract data and artifacts
Once you have the HTML, clean the DOM and convert content into usable formats: Markdown for reading, JSON for interactive elements, a text snapshot for language models.
ScraperFlow produces these three artifacts automatically and keeps the site structure and assets.
Step 4 — Handle anti-bot and respect the site
Many sites detect bots (fingerprinting, CAPTCHA, rate limiting). A stealth engine and a reasonable pace reduce blocking.
Always respect the site terms of use, robots.txt and GDPR: collect only public and necessary data.
Frequently asked questions
Can you scrape any website?
Technically, almost any. Legally, you must respect the terms of use, robots.txt and GDPR. Scraping public data is generally allowed; personal data requires a legal basis.
Do you need to code to scrape a website?
No. A tool like ScraperFlow is enough: enter a URL and start the scrape. The REST API lets you automate when needed.
Why does my scrape return an empty page?
Most often because the content is generated in JavaScript. You then need a headless browser that runs the JS before extracting the HTML.