AI extraction: turn the web into Markdown and JSON
To feed an LLM or a RAG index you need clean content — not raw HTML. ScraperFlow fetches the real content, cleans it and delivers Markdown, structured JSON and a text snapshot, ready for your pipeline.
From raw HTML to AI-ready data
A page’s HTML drowns the useful content in markup, navigation, scripts and styles. A language model does not need that noise: it needs the text, its structure and its metadata.
ScraperFlow does the conversion upstream and delivers directly consumable artifacts — Markdown, structured JSON and a text snapshot — with no manual cleanup step.
- Native Markdown (headings, lists, tables, links)
- JSON of the extracted data
- Text snapshot for RAG indexing
Why Markdown instead of HTML
Markdown keeps the semantic structure (heading hierarchy, lists, tables, links) while removing presentation tags. The content stays readable for a human and a machine alike.
The result: less noise, an explicit structure and shorter documents — which helps models target the information and cuts the token cost.
- Explicit structure: headings, lists, tables
- Links and images as absolute URLs
- A single format, readable and diffable
Step 1 — Fetch the real content
Many sites generate their content with JavaScript: a plain HTTP request returns only an empty shell. ScraperFlow runs a headless browser and waits for rendering to get the real content.
Anti-bot protections (fingerprinting, CAPTCHA, rate limiting) are handled by a “stealth” engine and a controlled pace, with no side effect on the source site.
Step 2 — Clean the HTML (remove noise)
Raw HTML contains navigation, footers, cookie banners, ads and scripts. Converting it as-is would produce polluted Markdown.
ScraperFlow isolates the main content (article, product page, documentation) and strips these intrusive elements before conversion.
- Main-content extraction (Readability-style)
- Removal of menus, scripts and banners
- Relative links resolved to absolute URLs
Step 3 — Extract schema-based JSON
When you need precise fields (price, title, date, author), free text is not enough: you need structured data. ScraperFlow lets you describe the expected schema and returns a conforming JSON.
You then validate that JSON in code with your ecosystem’s tools (Pydantic in Python, Zod in TypeScript) for a typed, safe integration.
- Schema you declare yourself (fields and types)
- JSON output ready to use
- Client-side validation: Pydantic, Zod, JSON Schema
RAG index: structure, chunking, consistency
A good RAG index relies on clean, consistent documents correctly linked to their source. Markdown per page makes chunking into controlled-size segments easier.
ScraperFlow preserves the URL tree and keeps one artifact per page: you keep the source → segment traceability, essential to cite or verify an answer.
Wire the API into your pipeline
ScraperFlow’s REST API fits any pipeline: you provide a start URL and get the artifacts back (Markdown, JSON, snapshot), ready to index or process.
Billing is per page, with no multiplier: a page with JavaScript rendering costs the same token as a static page, keeping your RAG pipeline predictable in cost.
Frequently asked questions
Does ScraperFlow produce LLM-ready Markdown?
Yes. Each page can be delivered as native Markdown (headings, lists, tables, links as absolute URLs), with noise removed (navigation, scripts, cookies) for a document directly usable by an LLM or a RAG index.
Can I extract JSON according to my own schema?
Yes. You describe the expected fields (price, title, date, author…) and ScraperFlow returns structured JSON, which you then validate with Pydantic, Zod or a JSON Schema in your code.
Do I need an LLM to use ScraperFlow?
No. ScraperFlow fetches and structures the data without a language model: the Markdown and JSON are produced by deterministic processing. An LLM is useful downstream (summaries, Q&A, RAG), not for the extraction itself.
How do I cut the token cost of my RAG pipeline?
By working on cleaned Markdown rather than raw HTML: the useful content is kept, the noise disappears. The order of magnitude of the reduction depends on the pages (navigation and scripts often weigh a lot) — measure it on your own content before sizing.