Second-Memory stores things a user writes over time.
If a user later asks:
βWhat did I write about performance problems?β
A normal keyword search isn't always enough.
The memory might say:
βThe database becomes slow when thousands of users query it at the same time.β
There may be no exact keyword match between the question and the memory.
This is where semantic search becomes useful.
I can convert a piece of text into an embedding β a vector representation of its meaning.
For example:
"The database becomes slow when thousands of users query it at the same time."
β
[Embedding model]
β
[0.021, -0.183, ...]
The actual vector contains many dimensions, so it isn't meaningful to look at the individual numbers.
What matters is the relationship between vectors.
Texts with similar meanings tend to have vectors that are closer together.
So when a user asks a question, I can:
Question
Embedding
Vector similarity search
Relevant memories
This gives Second-Memory a way to retrieve memories based on meaning, rather than just matching words.
Once I decided to use embeddings, I needed somewhere to store and search them.
I considered:
| pgvector | Pinecone | |
|---|---|---|
| Vector search | Yes | Yes |
| Relational data | Yes | No |
| Existing PostgreSQL | Yes | No |
| Separate infrastructure | No | Yes |
| Operational complexity | Lower | Higher |
| Good fit for V1 | Yes | Yes |
I chose pgvector.
The main reason wasn't that pgvector is necessarily better than Pinecone.
It was that Second-Memory already had a natural place for the vectors:
the Memory Service's database.
With pgvector, I could keep the memory and its embedding together.
erDiagram
"Memory Service" ||--|| "PostgreSQL + pgvector" : utilizes
"PostgreSQL + pgvector" ||--|{ Memory : contains
Memory {
uuid user_id
text content
vector embedding
}
That kept the architecture simple.
A dedicated vector database could make sense at larger scale.
But introducing one also creates another system to operate and another boundary to manage.
For V1, I didn't see enough benefit to justify that complexity.
The Memory Service could own:
memory data
embeddings
vector search
all within the same data boundary.
That also reinforced one of the architectural principles from the previous post:
The service that owns the data should own access to it.
The Ask Service doesn't need to know whether semantic search is implemented with pgvector, Pinecone, or something else.
It simply asks the Memory Service for relevant memories.
With pgvector, the basic retrieval flow becomes:
User question
β
Generate query embedding
β
Memory Service
β
pgvector similarity search
β
Relevant memories
β
Ask Service
β
LLM
This is the first important piece of the AI architecture.
The LLM doesn't need to know everything the user has ever written.
Instead, the system retrieves the memories that are most relevant to the current question and uses those as context.
For Second-Memory V1, I chose:
PostgreSQL + pgvector
because it gave me semantic search without introducing another database.
It was a pragmatic choice:
relational data
one data boundary
less infrastructure
simpler development