{"slug": "building-the-wordpress-connector-for-cognee", "title": "Building the WordPress Connector for Cognee", "summary": "A developer built an official WordPress connector for Cognee (PR #221) during a hackathon, letting the AI memory engine ingest blog content as an interconnected knowledge graph rather than isolated vector snippets. The connector uses WordPress Application Passwords for read-only access, strips Gutenberg HTML and encoded entities into clean plain text, performs incremental syncs based on post modification timestamps, and detects deleted or unpublished posts so Cognee purges their nodes and vectors. It exposes a wordpress_source() function that feeds WordPress REST API v2 content into Cognee's cognify pipeline for graph-based search and Q&A.", "body_md": "**How I Built a WordPress Connector for Cognee: Giving AI Real Memory for Your Blog**\n\nIf you’ve ever tried building an AI chatbot or search tool for a WordPress site, you've probably noticed that standard search (and basic RAG) is pretty clunky:\n\nIt chops your articles into disconnected snippets, losing track of how posts, authors, and comments connect.\n\nIf you update a single typo in a blog post, traditional tools often re-read the entire site from scratch.\n\nIf you delete an article, the AI often keeps \"remembering\" it and making things up.\n\nFor the Cognee hackathon, I set out to fix this by building the official WordPress connector for Cognee (PR #221).\n\nHere’s a quick breakdown of how it works in plain English and how you can use it.\n\n**What Does Cognee Do Differently?**\n\nMost AI search tools just dump your text into a vector database. Cognee works more like human memory: it organizes your content into an interconnected knowledge graph alongside vectors.\n\nThat means when someone asks a question, Cognee doesn't just match keywords—it traces relationships between concepts, writers, categories, and discussions.\n\n**4 Practical Problems I Had to Solve**\n\nConnecting WordPress to an AI memory engine sounds simple until you actually look at the data coming out of the WordPress API. Here’s what I built to handle the real-world messiness:\n\nSafe, Hassle-Free Login\n\nNobody wants to hand their main admin password to an AI script. The connector uses WordPress Application Passwords (built into WordPress 5.6+). You generate a unique passcode right inside your WP Profile, pass it to the connector, and you can revoke it anytime with one click. It only makes read-only requests (GET), so it will never change or break anything on your site.\n\nStripping the HTML Clutter\n\nWordPress doesn't store plain text—it stores Gutenberg block comments, \n\ntags, and encoded symbols like & or –. Feeding raw HTML into an AI wastes tokens and confuses the model. I built a lightweight cleaner that strips the tags, converts entities back to normal characters, and hands clean plain text to Cognee.\n\nSyncing Only What’s New (Incremental Sync)\n\nIf you have 500 articles and publish one new tutorial, you shouldn't have to re-process all 500. The connector remembers the timestamp of the last sync and tells WordPress: \"Only send me posts modified after this date.\" This makes updates super fast.\n\nForgetting What You Deleted (\"Forget-on-Delete\")\n\nThis was one of the biggest requirements. If you delete or unpublish an article, the AI shouldn't keep answering questions with outdated information.\n\nEach sync, the connector does a quick, lightweight check of active post IDs. If a post that was there yesterday is gone today, it flags it as deleted. Cognee then automatically purges those nodes and vectors from its memory graph.\n\nHow to Use It in 5 Minutes\n\nHere is how simple it is to pull your WordPress content into Cognee and ask questions:\n\n**Architecture Diagram**\n\n```\nflowchart TD\n    subgraph WP [\"1. WordPress Site (REST API v2)\"]\n        A[\"Posts, Pages, Comments & Custom Types\"]\n    end\n\n    A -->|\"Application Passwords (HTTP Basic Auth)\"| B\n\n    subgraph CONNECTOR [\"2. Cognee WordPress Connector\"]\n        B[\"wordpress_source()\"]\n        B --> C[\"Lightweight ID Sweep\\n(Detects deleted posts)\"]\n        B --> D[\"Incremental Fetch\\n(Only gets modified posts)\"]\n        C --> E[\"Tombstone Emitter\\n(_deleted = True)\"]\n        D --> F[\"HTML Sanitizer\\n(Strips tags & unescapes)\"]\n    end\n\n    E -->|\"dlt merge table\"| G\n    F -->|\"dlt merge table\"| G\n\n    subgraph COGNEE [\"3. Cognee Memory Engine\"]\n        G[\"Document-Mode Routing\"]\n        G --> H[\"Cognify Pipeline\\n(Chunking & Entity Extraction)\"]\n        H --> I[\"Knowledge Graph + Vector Store\\n(Relationships & Embeddings)\"]\n    end\n\n    I --> J[\"4. Search & Q&A\\ncognee.search(..., GRAPH_COMPLETION)\"]\n```\n\npython\n\n``` python\nimport asyncio\nimport os\nimport cognee\nfrom cognee_community_connector_wordpress import wordpress_source\nasync def main():\n    # 1. Connect to your site (posts, pages, and comments)\n    source = wordpress_source(\n        base_url=\"https://your-site.com\",\n        username=os.getenv(\"WORDPRESS_USERNAME\"),\n        app_password=os.getenv(\"WORDPRESS_APP_PASSWORD\"),\n        content_types=[\"posts\", \"pages\", \"comments\"],\n    )\n    # 2. Ingest into Cognee memory\n    await cognee.remember(\n        source,\n        dataset_name=\"my_blog\",\n        primary_key=\"id\",\n        write_disposition=\"merge\",\n        max_rows_per_table=0,\n    )\n    # 3. Ask your blog anything!\n    answer = await cognee.search(\n        query_text=\"What are the main topics and recent tutorials covered on my site?\",\n        query_type=cognee.SearchType.GRAPH_COMPLETION,\n        datasets=[\"my_blog\"],\n    )\n    print(\"\\nAnswer from Memory:\\n\", answer)\nif __name__ == \"__main__\":\n    asyncio.run(main())\n```\n\n**Testing & Quality**\n\nTo make sure this works reliably in production, I wrote a test suite of 11 unit tests covering everything from URL formatting to transient error retries and deletion flags. Everything runs cleanly offline and passes all linter checks (pytest + ruff).\n\n**Wrapping Up**\n\nBuilding this connector showed me how powerful Cognee's cognitive graph approach is compared to standard vector RAG. It bridges the gap between the world's most popular CMS and modern AI memory.\n\nCheck out the code and discussion on GitHub:\n\nPull Request: topoteretes/cognee-community#221\n\nIssue: topoteretes/cognee#4789", "url": "https://wpnews.pro/news/building-the-wordpress-connector-for-cognee", "canonical_source": "https://dev.to/byterecon_264/building-the-wordpress-connector-for-cognee-2ken", "published_at": "2026-10-05 11:05:59+00:00", "updated_at": "2026-10-05 11:19:33.942789+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-agents", "developer-tools", "ai-search"], "entities": ["WordPress", "Cognee", "WordPress REST API v2", "Gutenberg", "WordPress Application Passwords", "dlt"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-the-wordpress-connector-for-cognee", "markdown": "https://wpnews.pro/news/building-the-wordpress-connector-for-cognee.md", "text": "https://wpnews.pro/news/building-the-wordpress-connector-for-cognee.txt", "jsonld": "https://wpnews.pro/news/building-the-wordpress-connector-for-cognee.jsonld"}}