How I Built a Self-Improving AI Reply System for E-commerce Sellers with pgvector A developer building Netaliz, a profit analytics tool for Trendyol marketplace sellers, added an AI module that drafts customer-question replies and improves itself by storing every seller-approved answer as a pgvector embedding in Postgres, then retrieving the three closest approved answers per store as few-shot examples in the prompt. The system keeps human review as the default, with auto-send opt-in, so each approval becomes a new training example without fine-tuning. Marketplace sellers in Turkey get a constant stream of customer questions: "Is this table waterproof?", "Will it fit a small balcony?", "Why are the reviews so bad?". Every unanswered question is a lost sale, but writing thoughtful replies all day doesn't scale. While building Netaliz https://netaliz.com , a profit analytics tool for Trendyol sellers, I added an AI module that drafts replies to these questions. The interesting part isn't calling an LLM. It's making the system get better every time the seller approves an answer . Here's how it works. My first version was simple: send the product info and the question to an LLM, get an answer back. It worked, but: Sellers were editing almost every draft. That's not automation, that's extra work. The final system has four layers that get assembled into the prompt: Layer 4 is what makes it self-improving. Every time a seller approves a draft or edits and sends it , we store the question, the final answer and an embedding of the question: CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE approved answers id BIGSERIAL PRIMARY KEY, store id BIGINT NOT NULL, question TEXT NOT NULL, answer TEXT NOT NULL, embedding vector 1536 , created at TIMESTAMPTZ DEFAULT now ; CREATE INDEX ON approved answers USING hnsw embedding vector cosine ops ; Keeping this in Postgres instead of a separate vector database was a deliberate choice. The data already lives there, it's scoped per store with a simple WHERE , and it's one less service to run. When a new question comes in, we embed it and pull the closest approved answers from the same store: SELECT question, answer FROM approved answers WHERE store id = $1 ORDER BY embedding <= $2 LIMIT 3; Those examples go into the prompt as few-shot demonstrations. If a seller always answers sizing questions in a particular way, the model sees that pattern and follows it, without any fine-tuning. We also give the model a fixed three-step structure, which made answers noticeably more persuasive: A simplified version of the prompt assembly: function buildPrompt ctx: ReplyContext : string { return You are a customer support writer for a marketplace seller. , Tone: ${ctx.brandVoice}. Never use: ${ctx.bannedWords.join ", " }. , Rules:\n${ctx.templateRules.map r = - ${r} .join "\n" } , Product facts:\n${JSON.stringify ctx.product } , Structure: acknowledge the concern, give one concrete argument, close with trust. , Examples of approved answers:\n${ctx.examples .map e = Q: ${e.question}\nA: ${e.answer} .join "\n\n" } , Customer question: ${ctx.question} , .join "\n\n" ; } By default, nothing is sent automatically. The AI drafts, the seller reviews and clicks send. Auto-send is opt-in, and even then the brand voice, banned words and template rules still apply. This turned out to matter for two reasons: sellers trust the system more, and every human approval becomes a new training example. The review step is the learning loop. If you're curious about the product side, Netaliz has a demo store with sample data, no signup needed: netaliz.com/demo https://netaliz.com/demo . I'd love to hear how others are handling per-user style in LLM apps. Are you using retrieval, fine-tuning, or something else?