{"slug": "why-pgvector", "title": "Why pgvector?", "summary": "A developer building Second-Memory, a service that stores user-written memories over time, chose PostgreSQL with the pgvector extension over Pinecone to add semantic search to the system's V1 architecture. The engineer converted text into embeddings and used vector similarity search inside the Memory Service's existing database, keeping memory data, embeddings and retrieval within a single data boundary rather than operating a separate vector database. The retrieval flow generates a query embedding, runs a pgvector similarity search, and passes the most relevant memories to the Ask Service as context for an LLM.", "body_md": "Second-Memory stores things a user writes over time.\n\nIf a user later asks:\n\n“What did I write about performance problems?”\n\nA normal keyword search isn't always enough.\n\nThe memory might say:\n\n“The database becomes slow when thousands of users query it at the same time.”\n\nThere may be no exact keyword match between the question and the memory.\n\nThis is where **semantic search** becomes useful.\n\nI can convert a piece of text into an **embedding** — a vector representation of its meaning.\n\nFor example:\n\n```\n\"The database becomes slow when thousands of users query it at the same time.\"\n         ↓\n[Embedding model]\n         ↓\n[0.021, -0.183, ...]\n```\n\nThe actual vector contains many dimensions, so it isn't meaningful to look at the individual numbers.\n\nWhat matters is the relationship between vectors.\n\nTexts with similar meanings tend to have vectors that are closer together.\n\nSo when a user asks a question, I can:\n\nQuestion\n\nEmbedding\n\nVector similarity search\n\nRelevant memories\n\nThis gives Second-Memory a way to retrieve memories based on **meaning**, rather than just matching words.\n\nOnce I decided to use embeddings, I needed somewhere to store and search them.\n\nI considered:\n\n|  | pgvector | Pinecone | \n|---|---|---|\n| Vector search | Yes | Yes | \n| Relational data | Yes | No | \n| Existing PostgreSQL | Yes | No | \n| Separate infrastructure | No | Yes | \n| Operational complexity | Lower | Higher | \n| Good fit for V1 | Yes | Yes | \n\nI chose **pgvector**.\n\nThe main reason wasn't that pgvector is necessarily better than Pinecone.\n\nIt was that Second-Memory already had a natural place for the vectors:\n\n**the Memory Service's database.**\n\nWith pgvector, I could keep the memory and its embedding together.\n\n```\nerDiagram\n    \"Memory Service\" ||--|| \"PostgreSQL + pgvector\" : utilizes\n    \"PostgreSQL + pgvector\" ||--|{ Memory : contains\n\n    Memory {\n        uuid user_id\n        text content\n        vector embedding\n    }\n```\n\nThat kept the architecture simple.\n\nA dedicated vector database could make sense at larger scale.\n\nBut introducing one also creates another system to operate and another boundary to manage.\n\nFor V1, I didn't see enough benefit to justify that complexity.\n\nThe Memory Service could own:\n\nmemory data\n\nembeddings\n\nvector search\n\nall within the same data boundary.\n\nThat also reinforced one of the architectural principles from the previous post:\n\n**The service that owns the data should own access to it.**\n\nThe Ask Service doesn't need to know whether semantic search is implemented with pgvector, Pinecone, or something else.\n\nIt simply asks the Memory Service for relevant memories.\n\nWith pgvector, the basic retrieval flow becomes:\n\n```\nUser question\n      ↓\nGenerate query embedding\n      ↓\nMemory Service\n      ↓\npgvector similarity search\n      ↓\nRelevant memories\n      ↓\nAsk Service\n      ↓\n     LLM\n```\n\nThis is the first important piece of the AI architecture.\n\nThe LLM doesn't need to know everything the user has ever written.\n\nInstead, the system retrieves the memories that are most relevant to the current question and uses those as context.\n\nFor Second-Memory V1, I chose:\n\n**PostgreSQL + pgvector**\n\nbecause it gave me semantic search without introducing another database.\n\nIt was a pragmatic choice:\n\nrelational data\n\none data boundary\n\nless infrastructure\n\nsimpler development", "url": "https://wpnews.pro/news/why-pgvector", "canonical_source": "https://dev.to/joungpark/-why-pgvector-2deb", "published_at": "2026-10-07 03:34:17+00:00", "updated_at": "2026-10-07 03:47:46.196298+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-agents", "mlops"], "entities": ["Second-Memory", "pgvector", "PostgreSQL", "Pinecone"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/why-pgvector", "markdown": "https://wpnews.pro/news/why-pgvector.md", "text": "https://wpnews.pro/news/why-pgvector.txt", "jsonld": "https://wpnews.pro/news/why-pgvector.jsonld"}}