Most HTML-to-Markdown tools handle one file at a time. You paste some HTML, get Markdown back, repeat. That works for a quick snippet but not when you have 200+ pages from a help center export sitting in a folder.
I needed exactly that. I had a full site mirror (grabbed with wget --mirror
) and wanted clean Markdown I could feed into an LLM knowledge base. Nothing I found could handle it without up files to a server or converting one by one.
So I built HTML to Markdown AI.
No server involved. Your files never leave your machine.
The heavy lifting happens in Go/WASM. The pipeline:
Intentional decision. Down HTML from someone else's site has legal implications depending on jurisdiction and terms of service. I don't want to be in that business.
Down is also the easy part:
wget -r -l 0 -np -k -E -p -e robots=off \
--reject-regex '\.(png|jpe?g|gif|svg|webp|woff2?|ttf|css|js|zip|pdf)$' \
-w 0.5 --random-wait \
https://docs.example.com/
That gives you a local folder with all the HTML. The hard and annoying part is turning that into clean, usable Markdown. That's what this tool solves.
https://www.html-to-markdown-ai.com
Use cases I've tested it with:
If you work with LLMs and regularly need to get web content into a format they can digest, this might save you some time.
Feedback welcome. What would make this more useful for your workflow?