Web scraping · Open source

Scrape any website.

Capture the source code, full structure and assets of any site — bypassing anti-bot protections — then export everything as Markdown, JSON or an LLM snapshot.

Open source Built-in anti-detection Export Markdown · JSON · LLM
run_8f3k2a · finished 2 min ago
exemple.com/
  assets/
    style.css
    app.js
  pages/
    index.html
    contact.html
  index.html
index.mdMarkdown
llm-snapshot.txtLLM snapshot
elements.jsonInteractive elements
How it works

From website to usable folder in three steps.

No complex setup. A clean, reusable result in minutes.

01

Configure

Enter the target site URL and the destination folder. Optionally set a page limit.

02

Scrape & bypass

The Playwright + Crawlee engine crawls the site while dodging detection: stealth, cookie banners, wait strategies.

03

Export

Download index.md, llm-snapshot.txt and elements.json — or save the full structure to the folder of your choice.

Capabilities

One complete engine, one tool.

Four building blocks: scraping, shield, extraction and API. All without proxies or scripts to maintain.

LIVE

Full scraping

Structure, assets, URL rewriting: a faithful, navigable copy of the site.

example.com/ · 42 pages · 3.2 MB
LIVE

Stealth Shield

Detection evasion, cookie banner dismissal and wait strategies to fly under the radar.

applyStealth(context) · cookie-dismisser · wait-strategies
LIVE

AI Extraction

DOM purification, interactive element tagging, serialization ready for language models.

index.md · llm-snapshot.txt · elements.json
LIVE

REST API

Launch scrapes in the background and track their status through a simple, robust API.

POST /api/scrape · GET /api/jobs · GET /api/health
Developers & AI

Built for developers and AI agents.

Automate your scrapes via the REST API, or feed your LLMs with the extraction artifacts.

REST API

Launch an async scrape and track its status. A single endpoint, JSON artifacts, a bounded queue and built-in anti-SSRF.

POST /api/scrape
curl -X POST http://localhost:2027/api/scrape \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://exemple.com",
    "folderName": "mon-site",
    "maxPages": 50
  }'
202 response · jobId + pollUrl Anti-SSRF · anti path-traversal

LLM artifacts

Each scrape produces files ready to be consumed by an AI agent or a data pipeline.

generated artifacts
mon-site/
├── index.md          # Markdown content
├── llm-snapshot.txt  # LLM snapshot
├── elements.json     # interactive elements
└── assets/           # static files
Clean Markdown Structured JSON LLM snapshot
Pricing

Pay for what you scrape.

Start with a free trial, then scale in volume, speed and support.

Free

Discovery trial.

€0/ month

100 tokens/month · 1 site · no card.

  • 1 site, a single trial scrape
  • 100 tokens free every month (25 pages max per site)
  • Markdown · JSON · LLM extraction
  • Results kept 7 days
  • Online documentation (no support)
Try for free

Scale

For teams.

€79/ month

6,000 tokens/month · 3 users.

  • 6,000 tokens / month
  • Unlimited sites · 3 users
  • Faster extraction
  • Results kept 1 year
  • Priority support (reply < 24h)
Choose Scale

Free trial, no card required · volume, API and support from Starter — see full pricing. — See pricing.

FAQ

Questions, answers.

What is ScraperFlow?

ScraperFlow scrapes an entire website — source code, structure and assets — while bypassing anti-bot protections. It produces reusable artifacts: Markdown, LLM snapshot and interactive elements as JSON.

Do I need technical skills?

No. The web interface is enough: enter a URL and a destination folder, then click Download. The REST API is there if you want to automate.

Which artifacts are generated?

Three files: index.md (Markdown content), llm-snapshot.txt (LLM-optimized snapshot) and elements.json (structured interactive elements). The structure and assets are preserved.

How does the anti-detection shield work?

It combines Playwright stealth, automatic cookie banner dismissal and framework-aware wait strategies to reduce the risk of being blocked.

Is it legal?

Scraping public content is generally legal, but always respect the target site's terms of service and applicable regulations (GDPR, copyright).

Is it free?

The Free plan is a discovery trial: 1 site, 3 pages, no credit card. Beyond that, paid plans unlock monthly volume, the API, longer retention and support.

The web has the data.
You have the source code.

Run your first scrape in seconds.