Convert a website to Markdown: the complete guide
Markdown turns raw HTML into structured, lightweight text that language models read easily. Here is how to convert a website to clean Markdown: fetching the HTML, cleaning the DOM, conversion tools and automation at scale.
What is converting a website to Markdown?
Converting a website to Markdown means turning the HTML of one or more pages into structured Markdown text: headings, lists, tables, links and code blocks, without presentation tags.
Markdown is a lightweight text format that stays readable as-is while preserving the content hierarchy — exactly what a language model needs.
Why convert to Markdown for LLMs?
HTML carries a lot of noise (menus, scripts, styles, ads) that consumes tokens and blurs meaning. Markdown keeps only the useful structure.
The result: fewer tokens for the same content, an explicit hierarchy (headings, lists) and better use in RAG or for training.
- Less noise and fewer tokens than full HTML
- Explicit structure: headings, lists, tables, links
- Readable by humans and machines alike
- Ideal for RAG, documentation and archiving
Step 1 — Fetch the page HTML
You do not convert a site directly: first fetch the HTML of each page. A plain HTTP request is enough if the content is present in the response.
Many sites generate their content with JavaScript: a headless browser (Playwright, Puppeteer) then runs the JS and waits for load to obtain the real HTML.
Step 2 — Clean the DOM before conversion
Raw HTML includes navigation, footers, cookie banners, ads and scripts. Converting them would produce polluted Markdown.
So you remove these elements and keep only the main content (article, product page, documentation), ideally via a Readability-style extraction.
Step 3 — Convert HTML to Markdown
Conversion maps each HTML tag to its Markdown equivalent: `h1`–`h6` become `#`, `ul`/`ol` become lists, `table` a table, `a` a link, `pre`/`code` a code block.
Proven tools automate the job: Turndown (JavaScript), html2text (Python) or Pandoc (multi-format). ScraperFlow outputs Markdown directly, with no manual step.
Step 4 — Pitfalls of badly formed Markdown
Hastily converted Markdown keeps artifacts: duplicate headings, broken tables, stray line breaks, unresolved relative links.
For LLMs, these flaws hurt quality: clean Markdown, a consistent hierarchy and absolute links beat raw volume.
- Consistent heading hierarchy (one `#` per document)
- Well-delimited tables and code blocks
- Links and images as absolute URLs
- Metadata (title, URL, date) in a header when useful
Automating conversion at scale
Converting one page is easy; converting a whole site means following links (crawl), respecting a page limit and keeping the URL structure.
ScraperFlow automates this crawl and delivers one Markdown file per page, ready to feed a RAG index or a knowledge base.
Frequently asked questions
Can any website be converted to Markdown?
Technically, almost any. You first fetch the HTML (with a headless browser if the content depends on JavaScript), then respect robots.txt and GDPR. Markdown reproduces structure, not animations or complex layouts.
Does Markdown keep images and formatting?
Markdown keeps the structure (headings, lists, tables, links) and images as links, but not the visual style (colors, fonts, layout). Complex tables and interactive components may lose fidelity.
Do you need to code to convert a site to Markdown?
No. A tool like ScraperFlow is enough: you provide a URL and get Markdown back. For custom needs, libraries such as Turndown or Pandoc plug into a script.
Does Markdown really improve LLM results?
Generally yes: clean Markdown reduces noise and token count while preserving the content hierarchy, which helps models target information better — especially in RAG.