The previous post covered why I chose pgvector for semantic search.
But there was another question:
When should a memory be embedded?
At first, it might seem simple:
Create memory
βΌ
Generate embedding
βΌ
Save embedding
But this makes memory creation dependent on the embedding process.
If the embedding provider is slow or unavailable, creating a memory could also fail or become slow.
I wanted to separate these two operations.
The architecture became:
Memory Service
βΌ
PostgreSQL
βΌ
Outbox Table
βΌ
Queue / Job Runner
βΌ
Embedding Worker
βΌ
pgvector
When a user creates a memory, the Memory Service writes the memory and an outbox event in the same database transaction.
For example:
Transaction
βββββββββββββββββββββββββββββββ
β INSERT memory β
β INSERT outbox event β
βββββββββββββββββββββββββββββββ
The important part is that they succeed or fail together.
This avoids a common problem with distributed systems:
Memory saved
βΌ
Embedding event lost
The outbox gives me a durable record of the work that needs to happen.
The outbox is not the queue itself.
It is a reliable bridge between the database transaction and asynchronous processing.
A separate process reads pending outbox events and puts jobs onto a queue.
Outbox
βΌ
Queue
βΌ
Embedding Worker
This gives the embedding process some useful properties:
asynchronous processing
retries
independent scaling
failure isolation
no need to block the memory API
The user can save a memory without waiting for the embedding provider.
The embedding worker has one main responsibility:
turn memory text into an embedding and store it in pgvector.
flowchart TD
A[Get Memory] --> B[Generate embedding]
B --> C[Store Vector]
This keeps embedding-specific logic out of the Memory Service's synchronous request path.
It also gives me a place to evolve the embedding pipeline later.
For example, I could change:
embedding models
chunking strategy
retry behaviour
batch processing
embedding dimensions
without changing the API used to create a memory.
There is an important consequence of this design:
A newly created memory may not be immediately searchable.
There can be a small delay between:
Memory created
β asynchronous processing
βΌ
Embedding created
βΌ
Memory available for semantic search
I accepted this trade-off.
For Second-Memory, immediate persistence is more important than making the embedding operation part of the user's request.
This is essentially eventual consistency between the memory record and its vector representation.
The asynchronous design also changes how failures work.
If the embedding provider temporarily fails:
Memory
βΌ
Outbox
βΌ
Queue
βΌ
Embedding Worker - X -> Retry
The memory itself has already been safely stored.
The embedding job can be retried without asking the user to submit the memory again.
That separation was important to me.
The final flow became:
graph TD
CM[Create Memory] --> PG
subgraph PG [PostgreSQL]
M[Memory]
OE[Outbox Event]
end
OE --> Q[Queue]
Q --> EW[Embedding Worker]
EW --> PV[(pgvector in PostgreSQL)]
Each part has a different responsibility:
| Component | Responsibility |
|---|---|
| Memory Service | Store the memory |
| Outbox | Reliably record the event |
| Queue | Deliver asynchronous work |
| Embedding Worker | Generate embeddings |
| pgvector | Store and search vectors |
This added some complexity compared with simply generating the embedding inside the API request.
But it gave me something more important:
memory creation is no longer tightly coupled to the availability of the embedding pipeline.
That was the trade-off I wanted.