{"slug": "i-tried-to-clean-commoncrawl", "title": "I tried to clean Commoncrawl", "summary": "Adhavan has released CC-FilteredCorpus, a dataset on Hugging Face that filters Common Crawl and converts web content into clean Markdown, preserving document structure such as headings, lists, links, code blocks, and tables, intended for LLM pretraining and NLP research.", "body_md": "I created a dataset called **CC-FilteredCorpus** where I’m trying to filter Common Crawl and extract the content as clean **Markdown**.\n\nThe goal is to preserve useful document structure—headings, lists, links, code blocks, tables, etc.—instead of flattening everything into plain text.\n\nI’m hoping this can be useful for LLM pretraining and other NLP research where preserving the structure of web documents matters.\n\nThe dataset is at [helloadhavan/CC-FilteredCorpus](https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus)", "url": "https://wpnews.pro/news/i-tried-to-clean-commoncrawl", "canonical_source": "https://discuss.huggingface.co/t/i-tried-to-clean-commoncrawl/178660#post_1", "published_at": "2026-08-14 16:11:44+00:00", "updated_at": "2026-08-14 16:18:08.296448+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "ai-tools"], "entities": ["Adhavan", "CC-FilteredCorpus", "Common Crawl", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/i-tried-to-clean-commoncrawl", "markdown": "https://wpnews.pro/news/i-tried-to-clean-commoncrawl.md", "text": "https://wpnews.pro/news/i-tried-to-clean-commoncrawl.txt", "jsonld": "https://wpnews.pro/news/i-tried-to-clean-commoncrawl.jsonld"}}