Web scraping for LLMs: preparing the web for your models

Feeding an LLM with the web requires clean, structured, up-to-date content. Here is how to go from raw HTML to documents ready for a language model: fetching, cleaning, Markdown, token estimation and a RAG pipeline.

  • Useful content isolated from noise
  • Explicit structure (headings, lists, tables)
  • Short, consistent documents
  • One level-1 heading per document
  • Tables and code blocks clearly delimited
  • Links and images as absolute URLs
  • Method: count the tokens of the raw HTML, then of the cleaned Markdown
  • Measure on YOUR pages: the ratio depends on the noise, not a universal number
  • Compare at equal useful content (same page, same scope)
Can you feed an LLM with scraped content?

Yes, as long as you prepare the content: fetch the real HTML, remove the noise and convert to Markdown or JSON. Feeding raw HTML degrades quality and wastes tokens. Also check the legal framework (robots.txt, GDPR) of each source.

Which format should I prefer: HTML, Markdown or JSON?

Markdown for reading and RAG (structure kept, noise removed); JSON when you need precise, typed fields; raw HTML only if you must process the layout. In practice you combine Markdown for the text and JSON for the fields.

How do I estimate the token count of a page?

You count the tokens of the text given to the model. Convert the page to cleaned Markdown, then count the tokens of that Markdown with the target model’s tokenizer. Compare with the count on the raw HTML to get the real reduction on your pages.

Is it legal to scrape the web for an LLM?

Scraping is not forbidden per se, but it is framed: respect robots.txt and terms of use, minimise personal data (GDPR) and keep a legitimate purpose. For commercial use or training, check the rights on the content you fetch.

Ready to scrape a website?

Run your first scrape in seconds, for free.

Start a scrape