Web scraping for LLMs: preparing the web for your models
Feeding an LLM with the web requires clean, structured, up-to-date content. Here is how to go from raw HTML to documents ready for a language model: fetching, cleaning, Markdown, token estimation and a RAG pipeline.
Why prepare the web for an LLM
A language model does not “see” a web page: it sees the text you give it. If that text is raw HTML full of menus, scripts and cookies, the model spends tokens on noise and answers less well.
Preparing the web for an LLM means turning pages into clean documents: useful content, explicit structure and metadata. It is the step that determines the quality of a RAG pipeline as much as of a training corpus.
- Useful content isolated from noise
- Explicit structure (headings, lists, tables)
- Short, consistent documents
The raw HTML problem
A typical HTML page buries the real content in navigation, footers, cookie banners, ads and tracking scripts. This markup makes up a large share of the document.
Feeding that HTML to a model means paying for noise: every tag, attribute and script burns tokens without adding useful information. So you must clean before indexing.
Step 1 — Fetch the real content
Many sites render their content client-side: a plain HTTP request returns only an empty shell. You need a headless browser (Playwright, Puppeteer) and must wait for rendering to get the real HTML.
Protected sites add fingerprinting, CAPTCHA and rate limiting. A “stealth” engine and a controlled pace let you reach the content with no side effect on the source site.
Step 2 — Convert to clean Markdown
First remove intrusive elements (navigation, scripts, cookies, ads) to keep only the main content, ideally via Readability-style extraction. Then convert the remaining HTML to Markdown.
Markdown keeps the semantic structure (headings, lists, tables, links) while removing presentation tags. Clean Markdown is readable by a human and a machine alike.
- One level-1 heading per document
- Tables and code blocks clearly delimited
- Links and images as absolute URLs
Estimate the token reduction
Tokens are counted on the text given to the model. Because a large share of an HTML page is markup and noise, replacing HTML with cleaned Markdown mechanically reduces the token count for the same useful content.
The order of magnitude varies widely by page: a page heavy on navigation, scripts and ads shrinks far more than an already minimal one. The common idea of a ~10× reduction matches those heavy pages — but it must be measured, not assumed.
- Method: count the tokens of the raw HTML, then of the cleaned Markdown
- Measure on YOUR pages: the ratio depends on the noise, not a universal number
- Compare at equal useful content (same page, same scope)
Extract structured JSON when you need fields
Markdown suits reading and RAG, but sometimes you need precise fields (price, date, author). Free text is then not enough: you need structured data.
Describe the expected schema and get conforming JSON, which you validate in code (Pydantic, Zod, JSON Schema). That way you combine a readable Markdown document with typed fields for automation.
Build a reliable RAG pipeline
A reliable RAG pipeline relies on clean documents, controlled chunking and source → segment traceability. Markdown per page makes chunking and source citation easier.
ScraperFlow delivers one artifact per page and preserves the URL tree: you keep the link between each segment and its origin, essential to verify an answer or cite the source.
Frequently asked questions
Can you feed an LLM with scraped content?
Yes, as long as you prepare the content: fetch the real HTML, remove the noise and convert to Markdown or JSON. Feeding raw HTML degrades quality and wastes tokens. Also check the legal framework (robots.txt, GDPR) of each source.
Which format should I prefer: HTML, Markdown or JSON?
Markdown for reading and RAG (structure kept, noise removed); JSON when you need precise, typed fields; raw HTML only if you must process the layout. In practice you combine Markdown for the text and JSON for the fields.
How do I estimate the token count of a page?
You count the tokens of the text given to the model. Convert the page to cleaned Markdown, then count the tokens of that Markdown with the target model’s tokenizer. Compare with the count on the raw HTML to get the real reduction on your pages.
Is it legal to scrape the web for an LLM?
Scraping is not forbidden per se, but it is framed: respect robots.txt and terms of use, minimise personal data (GDPR) and keep a legitimate purpose. For commercial use or training, check the rights on the content you fetch.