cd /news/large-language-models/i-tried-to-clean-commoncrawl · home topics large-language-models article
[ARTICLE · art-97025] src=discuss.huggingface.co ↗ pub= topic=large-language-models verified=true sentiment=· neutral

I tried to clean Commoncrawl

Adhavan has released CC-FilteredCorpus, a dataset on Hugging Face that filters Common Crawl and converts web content into clean Markdown, preserving document structure such as headings, lists, links, code blocks, and tables, intended for LLM pretraining and NLP research.

read1 min views1 publishedAug 14, 2026
I tried to clean Commoncrawl
Image: Discuss (auto-discovered)

I created a dataset called CC-FilteredCorpus where I’m trying to filter Common Crawl and extract the content as clean Markdown.

The goal is to preserve useful document structure—headings, lists, links, code blocks, tables, etc.—instead of flattening everything into plain text.

I’m hoping this can be useful for LLM pretraining and other NLP research where preserving the structure of web documents matters.

The dataset is at helloadhavan/CC-FilteredCorpus

── more in #large-language-models 4 stories · sorted by recency
── more on @adhavan 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-tried-to-clean-com…] indexed:0 read:1min 2026-08-14 ·