{"slug": "pgvector-semantic-search-in-postgresql-a-python-checklist", "title": "pgvector Semantic Search in PostgreSQL: A Python Checklist", "summary": "A developer published a checklist for running pgvector semantic search inside PostgreSQL, aimed at teams whose documents, permissions and metadata already live in the database. The writeup covers version verification, dimension and operator-class matching, HNSW and IVFFlat index tuning, and a psycopg 3 query pattern, warning that filters can silently defeat the index and that failures usually surface as plausible but worse results rather than errors.", "body_md": "*Originally published on [kuryzhev.cloud](https://kuryzhev.cloud/2026/10/03/pgvector-semantic-search-in-postgresql-a-python-checklist)*\n\nA product team has 40,000 support articles in PostgreSQL and a search box that only matches exact keywords. Someone suggests bolting on a separate vector database, but the articles, permissions, and metadata already live in Postgres. For many workloads, pgvector semantic search inside the existing database is the smaller and easier-to-operate choice, provided the setup is done deliberately.\n\nThis article is a checklist, not a tutorial. Work through it before you ship, and again when recall or latency looks wrong.\n\nSemantic search failures rarely raise errors. A query returns ten rows and they look plausible. Nobody notices that the index is ignored, the embedding model changed, or a filter silently removed the best matches. The system \"works\" while quietly returning worse results.\n\nThe typical failure modes fall into three groups:\n\nA checklist catches these because each one is a yes/no question you can verify.\n\nDetails here follow the documented behavior of recent pgvector releases. Some features depend on the extension version. For example, `halfvec` requires 0.7.0+ and iterative index scans require 0.8.0+. Verify with the [pgvector README](https://github.com/pgvector/pgvector) for the version your server runs. Managed services often lag behind the latest release.\n\n`SELECT extversion FROM pg_extension WHERE extname = 'vector';` and compare it to the features you plan to use. Managed PostgreSQL offerings document which versions they ship.`vector(N)` where N matches the model output. A mismatch fails on insert, which is the good outcome.`<=>`, L2 is `<->`, and negative inner product is `<#>`. The index operator class must match the operator you query with.`hnsw.ef_search`; for IVFFlat, `ivfflat.probes`.\nStart with the schema and index. This example assumes a 1536-dimension embedding model; change the number to match yours.\n\n```\nCREATE EXTENSION IF NOT EXISTS vector;\n\nCREATE TABLE docs (\n    id          bigserial PRIMARY KEY,\n    tenant_id   integer NOT NULL,\n    content     text NOT NULL,\n    embed_model text NOT NULL,          -- record which model produced the vector\n    embedding   vector(1536) NOT NULL   -- must equal the model's output dimension\n);\n\n-- HNSW with cosine ops; must match the <=> operator used in queries\nCREATE INDEX docs_embedding_hnsw\n    ON docs USING hnsw (embedding vector_cosine_ops)\n    WITH (m = 16, ef_construction = 64);\n\n-- Higher ef_search improves recall at the cost of latency (per session)\nSET hnsw.ef_search = 100;\n```\n\nNext, the Python side with psycopg 3 and the `pgvector` package. The `embed()` function is a placeholder for your embedding provider, such as the OpenAI API, Amazon Bedrock, or a local open-weight model.\n\n``` python\nimport numpy as np\nimport psycopg\nfrom pgvector.psycopg import register_vector\n\nDSN = \"postgresql://app:secret@localhost:5432/appdb\"\nMODEL = \"your-embedding-model\"  # keep identical for ingest and query\n\ndef embed(texts: list[str]) -> list[np.ndarray]:\n    \"\"\"Call your embedding provider here; return one vector per text.\"\"\"\n    raise NotImplementedError\n\ndef search(query: str, tenant_id: int, k: int = 5):\n    q = np.array(embed([query])[0], dtype=np.float32)\n    with psycopg.connect(DSN) as conn:\n        register_vector(conn)  # extension must already exist in this database\n        rows = conn.execute(\n            \"\"\"\n            SELECT id, content, embedding <=> %s AS distance\n            FROM docs\n            WHERE tenant_id = %s AND embed_model = %s\n            ORDER BY embedding <=> %s   -- ascending distance expression + LIMIT enables the index\n            LIMIT %s\n            \"\"\",\n            (q, tenant_id, MODEL, q, k),\n        ).fetchall()\n    return rows  # cosine similarity is 1 - distance\n```\n\n**Watch out for filters that defeat the index.** With an approximate index, the database retrieves candidate neighbors first and then applies `WHERE`. A selective filter, such as one small tenant, can leave you with fewer than `LIMIT` rows, or none. pgvector 0.8.0 and later offer iterative index scans to keep scanning until enough rows match. For HNSW, use `SET hnsw.iterative_scan = relaxed_order;` or `strict_order`. IVFFlat has `ivfflat.iterative_scan`. Alternatives include partial indexes per large tenant or partitioning. Check which options your version supports.\n\n**Watch out for mixing models.** If you re-embed with a newer model but leave old vectors in place, searches will still return results, just meaningless ones. Even models with the same dimension produce incompatible spaces. The `embed_model` column above exists so you can filter and migrate in batches.\n\nOther items that tend to slip through:\n\n`vector` indexes cap at 2,000 dimensions. Larger embeddings have three options:\n`halfvec` (pgvector 0.7.0+, indexable up to 4,000 dimensions).`maintenance_work_mem`. When it no longer fits, pgvector emits a NOTICE to the client session. Watch for it on large builds.`vector_l2_ops` is not used for a For general PostgreSQL behavior around extensions and index maintenance, the [official CREATE EXTENSION documentation](https://www.postgresql.org/docs/current/sql-createextension.html) covers the installation side that managed providers wrap.\n\nChecklists decay unless something enforces them. A few checks are cheap to put into CI or a scheduled job.\n\n**Plan assertion.** Run `EXPLAIN` on a representative query against a staging database with realistic row counts. Fail the pipeline if the plan contains a sequential scan on the docs table. Tiny test tables will usually choose a sequential scan. Either seed enough data for the planner to behave like production, or verify with `SET enable_seqscan = off` in that test only.\n\n**Recall regression test.** Keep a small set of labeled queries with expected documents. Compute exact nearest neighbors by running the same query without the index, for example in a rolled-back transaction with index scans disabled. Then compare the exact results to the approximate ones. Track the overlap over time. Rerun after any change to `m`, `ef_search`, the model, or the extension version. The following sketch shows the idea:\n\n``` python\ndef recall_at_k(conn, q, k=10):\n    # Approximate result: whatever the planner picks with the index enabled\n    approx = {r[0] for r in conn.execute(\n        \"SELECT id FROM docs ORDER BY embedding <=> %s LIMIT %s\", (q, k))}\n\n    # Exact result: disable index scans, then force a rollback so the\n    # SET LOCAL settings never leak into the rest of the session\n    with conn.transaction(force_rollback=True):\n        conn.execute(\"SET LOCAL enable_indexscan = off\")\n        conn.execute(\"SET LOCAL enable_bitmapscan = off\")\n        exact = {r[0] for r in conn.execute(\n            \"SELECT id FROM docs ORDER BY embedding <=> %s LIMIT %s\", (q, k))}\n\n    return len(approx & exact) / k  # alert if this drops below your threshold\n```\n\n**Model-drift guard.** Add a scheduled query that counts distinct `embed_model` values per table. Alert when more than one is present outside a planned migration. A one-line check like this helps catch the silent quality loss described above.\n\n**Migration scripts.** Keep extension creation, table DDL, and index parameters in versioned migrations, not in ad hoc notebooks. That way `ef_construction` and the opclass are reviewed like any other schema change.\n\nMore infrastructure-focused checklists for running data services live on [kuryzhev.cloud](https://kuryzhev.cloud/). For search, the pattern holds: pick the model, match the operator and index, set recall explicitly, and measure it on a schedule.", "url": "https://wpnews.pro/news/pgvector-semantic-search-in-postgresql-a-python-checklist", "canonical_source": "https://dev.to/oleksandr_kuryzhev_42873f/pgvector-semantic-search-in-postgresql-a-python-checklist-5d4i", "published_at": "2026-10-03 07:02:22+00:00", "updated_at": "2026-10-03 07:07:53.435303+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "mlops", "structured-data"], "entities": ["PostgreSQL", "pgvector", "psycopg", "OpenAI", "Amazon Bedrock", "kuryzhev.cloud"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pgvector-semantic-search-in-postgresql-a-python-checklist", "markdown": "https://wpnews.pro/news/pgvector-semantic-search-in-postgresql-a-python-checklist.md", "text": "https://wpnews.pro/news/pgvector-semantic-search-in-postgresql-a-python-checklist.txt", "jsonld": "https://wpnews.pro/news/pgvector-semantic-search-in-postgresql-a-python-checklist.jsonld"}}