{"slug": "im-building-a-search-engine-part-2-content-extraction", "title": "I’m Building a Search Engine (Part 2): Content Extraction", "summary": "SEO consultant Sara Taher published Part 2 of her newsletter series on building a search engine, covering the main content extraction step of indexation. Taher used Claude Code to extract text from only the 541 pages that returned HTTP 200, skipping JavaScript-dependent pages, and applied separate extraction rules for Substack, her own website, and YouTube plus the open-source trafilatura tool for other pages. Taher said the goal is a \"good enough\" search engine rather than a perfect one, and that normalization, tokenization, and keyword-to-page assignment remain for the next stage.", "body_md": "If you find value in my newsletter and want to support the work that goes into it consider —> [☕[Buy Me a Coffee](https://buymeacoffee.com/sarataher)☕]💡 Want more from me? [[SEO Strategy Course](https://sarataher.podia.com/how-to-create-an-seo-strategy)] | [[Build Your Own SEO Tools (with Python & Agentic AI)](https://sarataher.podia.com/introduction-of-python-for-marketers)] | [[Premium SEO Mastermind](https://sarataher.podia.com/in-house-seo-enterprise-meetup)]\n\nAlternatively you can become a paid subscriber:\n\n## **Intro**\n\nI shared the other week my new project, building a search engine. You can read about part 1 of this journey [here](https://www.seoriddler.com/p/im-building-a-search-engine-part). I’ve also started other projects since then **😄** but that’s a chat for another day! Today I’m sharing the next step in building the search engine:\n\n**Main Content Extraction**\n\nThe term “main content” has come to light again recently with Google’s update to their helpful content documentation. You see, your content quality is assessed based on the “main content” of the page. Things like top navigation, footer, sidebars, popups, etc… are not part of your main content, and would not influence your content quality score.\n\nSo the next step after crawling and collection the URLs, is extracting the main content. This is part of the indexation process. The first step basically.\n\nBecause in indexation, what you want to do is build a table mapping each page and a set of keywords. You’re basically saying for these keywords, this page provides a relevant answer.\n\nThe table will be used to find and serve pages when a user searches.\n\nYa… so let’s dive in!\n\n## **Main Content Extraction**\n\nUsing my dear friend Claude Code, I decided to extract text from pages in the following way:\n\n- I decided to ignore all non 200 pages. [Only 541 pages returned 200, and those are the ones worth extracting.]\n- Also, I decided to ignore all pages that require JS for their content, to simplify my process.\n- I build a different set of extraction rules. One is tailored specifically for my substack pages, one for my website pages (and body did I find space for technical optimizations there **😄)** , one for YouTube and an open source pre-existing tool for other pages called[trafilatur](https://trafilatura.readthedocs.io/en/latest/) .\n\nHere’s a summary of the extraction methods:\n\n- As a general rule, make sure to keep headings and paragraphs, as simple Markdown (# headings, paragraphs, lists), not one flat block of text.\n- Finally I generated a report to review the extraction process:\n\nlooks good right?\n\n## **Limitations**\n\nFew things I had to set aside at this stage to focus on delivering vs perfecting:\n\n- I would’ve loved to build a better extraction mechanism for YouTube that pulls the transcript.\n\nThe goal is not a perfect search engine, but rather a good enough one. I think I was able to do that at this point.\n\n## **Next Steps**\n\nClearly indexation is not done yet. I still need to normalize, and tokenize, and decide on the mechanism to assign keywords to pages. I like what I’m doing and that all that matters. Will keep you updated!\n\n## **And That’s a Wrap (Almost 😄)**\n\nBuilding tools is great, but building a search engine is something else. It’s definitely a passion project. I guess that’s why we do SEO **😄**\n\nI hope you found this helpful and inspirational.\n\n**That’s that for today folks and see you in the next newsletter!**\n\n## **Support the Riddler!**\n\n- Sign up for my newsletter if you’re not already. (Pssst, you can also become a paid subscriber)\n- Share the newsletter and invite your friends to signup. Help me reach 2k signups on Substack by end of 2026 please 🙂\n- Provide feedback on how I can make this newsletter better!!!\n- [Buy me coffee](https://buymeacoffee.com/sarataher?ref=sara-taher.com) .\n- If you’re an SEO tool or an SEO service provider, consider sponsoring my newsletter. I’m also open to other partnership ideas as well.\n\n*Disclaimer: LLMs were used to assist in wording and phrasing this blog.*", "url": "https://wpnews.pro/news/im-building-a-search-engine-part-2-content-extraction", "canonical_source": "https://www.seoriddler.com/p/im-building-a-search-engine-part-561", "published_at": "2026-10-06 07:28:56+00:00", "updated_at": "2026-10-06 07:48:37.466740+00:00", "lang": "en", "topics": ["ai-tools", "ai-agents", "ai-search"], "entities": ["Sara Taher", "Claude Code", "Substack", "YouTube", "trafilatura", "Google", "SEORiddler"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/im-building-a-search-engine-part-2-content-extraction", "markdown": "https://wpnews.pro/news/im-building-a-search-engine-part-2-content-extraction.md", "text": "https://wpnews.pro/news/im-building-a-search-engine-part-2-content-extraction.txt", "jsonld": "https://wpnews.pro/news/im-building-a-search-engine-part-2-content-extraction.jsonld"}}