{"slug": "bm25-is-only-as-good-as-your-tokens", "title": "BM25 is only as good as your tokens", "summary": "Neon released the lakebase_tokenizer extension for Lakebase Search, moving Postgres full-text search synonym and stop-word configuration out of server-side text files and into SQL tables so users can define entries with INSERT statements. The extension feeds those terms into GIN and lakebase_bm25 indexes as tsvector data, and because Lakebase branches, a database branch carries its tokenizer configuration, letting users test a new synonym set on a branch before shipping it. The change addresses a limitation of managed Postgres, where users connect to the database rather than the server and cannot add dictionary files, leaving them restricted to the lists that ship with Postgres.", "body_md": "[Lakebase Search](https://neon.com/docs/ai/lakebase-search) is the search primitive of the Neon backend. [One of the best things about it](https://neon.com/blog/lakebase-search-retrieval-agents) is that it ranks keyword results with BM25; that said, BM25 can only score the terms it's given. What a search can find is decided earlier, when the text is split into terms - a step called tokenization.\n\nPostgres lets you customize that step, but not from SQL. Its synonym and stop-word dictionaries read their entries from text files in a folder on the database server - teaching Postgres that \"k8s\" means \"kubernetes\" means putting a file on that machine. The limitation is that a managed Postgres service you connect to the database, not the server, so you can't add those files and you're limited to the lists that ship with Postgres.\n\nThe new `lakebase_tokenizer` extension, [now packaged with Lakebase Search](https://neon.com/docs/extensions/lakebase-tokenizer), moves that configuration into SQL tables. You define synonyms and stop words with `INSERT`, and they flow into GIN and `lakebase_bm25` indexes like any other `tsvector`.\n\nLike everything in the Neon universe, it [branches](https://neon.com/docs/introduction/branching). If you branch your database, all its Lakebase Search configuration comes along, tokenizer included - so for example, you can test a new synonym set on a branch before shipping it.\n\n## A quick refresher on Lakebase Search\n\n[Lakebase Search](https://neon.com/docs/ai/lakebase-search) adds vector, keyword, and hybrid search to Lakebase Postgres through Postgres extensions:\n\n- `lakebase_vector` adds the`lakebase_ann` index for vector similarity search. It uses the same`vector` types and operators as`pgvector` , so there's no migration.\n- `lakebase_text` adds the`lakebase_bm25` index for full text search. It works on standard`tsvector` columns and adds BM25 ranking and top-K pushdown, which GIN with`ts_rank` doesn't have.\n- `lakebase_tokenizer` is the newest piece we're discussing. It controls how text becomes the terms that`lakebase_bm25` (and GIN) index.\n\nLakebase Search introduces new indexes specifically designed for the [lakebase architecture](https://neon.com/docs/introduction/architecture-overview), where compute is ephemeral and storage is durable and shared. A Lakebase Search index lives in storage rather than in compute memory, it's available right after a cold start with no warmup, and it shows up on every new branch without a rebuild.\n\nThis post is about the full text search piece of Lakebase Search - specifically, about the step that runs before any ranking happens.\n\n## Your dictionaries shape your search results\n\nPostgres full-text search doesn't index raw text. It runs each document through a text-search configuration and stores the result as a `tsvector`: a sorted list of normalized tokens (Postgres calls them lexemes).\n\nHere is what the built-in `english` configuration produces for a short support ticket:\n\nA few things happened here:\n\n- Words were lowercased\n- \"in\" and \"the\" were dropped as stop words\n- The remaining words were stemmed (\"credentials\" became `credenti` )\n- The accent in \"Zürich\" was kept\n\nQueries go through the same process. A document matches when its tokens match the query's tokens, and BM25 then scores that overlap. So if a document and a query describe the same thing with different tokens, no ranking function can connect them. However good BM25 is, it never sees a match.\n\n## Adding your own dictionary to Postgres\n\nThe changes in the previous example (lowercasing, dropping stop words, stemming) don't come from the `tsvector` type itself - they come from Postgres' [dictionaries](https://www.postgresql.org/docs/current/textsearch-dictionaries.html). These built-in dictionaries handle general language well, but they can't know your application's vocabulary. E.g. your users might use \"k8s,\" \"kube,\" and \"Kubernetes\" for the same thing.\n\nOn a self-managed server, adding your own dictionary would be easy - you'd simply copy a file into that directory. On a hosted service like Neon, you connect to the database, not the server, so you can't add files to its disk.\n\nThe lakebase architecture makes this even stricter. In Lakebase Postgres, [compute nodes don't hold durable state - they scale to zero, restart, and get replaced.](https://neon.com/docs/introduction/architecture-overview) A file on one compute's disk would disappear with that compute and wouldn't exist on a new branch, so search configuration needs to live in the database itself.\n\n## How lakebase_tokenizer works\n\nWe built `lakebase_tokenizer` so you can add your own dictionaries to Lakebase Search and get more out of BM25.\n\nYour words live in two tables that the extension creates: `lakebase_tokenizer_synonyms` and `lakebase_tokenizer_stopwords`. Rows are grouped into named sets, and a set can hold up to 100,000 rows.\n\nAdding a dictionary takes four steps, all in SQL:\n\n1. **Add your words:**`INSERT` synonyms (`k8s` →`kubernetes` ) and stop words into a named set.\n2. **Create a dictionary:** Build it from the extension's`tokenizer_wholeword` template, and choose your sets plus options such as accent stripping and stemming.\n3. **Create a text-search configuration** that sends words to your dictionary\n4**Use it:** Pass the configuration to`to_tsvector` , and index the result with GIN or`lakebase_bm25` as usual\n\n## Example: searching through support tickets\n\nSay you run a developer tool, and you want your support engineers (or agents) to search past tickets. Your users write those tickets in their own words:\n\n| id | body | \n|---|---|\n| 1 | Kubernetes pods crash-looping after the node upgrade | \n| 2 | K8s ingress returns 502 errors on every deploy | \n| 3 | Kube scheduler keeps evicting pods, they crash on restart | \n| 4 | Rotating Postgres credentials in the Zürich region | \n| 5 | PostgreSQL connection pool exhausted during the nightly batch | \n| 10 | PG creds expired, cannot connect from CI | \n\nSo here's what happens when a support engineer searches using the built-in `english` configuration:\n\n| Query | Should find | Matches? | \n|---|---|---|\n| `k8s crash` | Ticket 1 | No | \n| `zurich postgres` | Ticket 4 | No | \n| `pg credentials` | Ticket 4 | No | \n\nLet's fix it.\n\n**1. Install the Lakebase Search extensions**\n\n**2. Add your vocabulary**\n\n**3. Create the dictionary and the configuration**\n\n**4. Use it in a column and index it**\n\nYou would now gnerate a `tsvector` column with the new configuration and index it with `lakebase_bm25`.\n\nThe Zürich ticket from earlier now produces these tokens:\n\n## Before and after, ranked by BM25\n\nTo compare the two approaches, we loaded all ten tickets (the six above plus four unrelated ones) and gave the table a second column that uses Postgres' built-in english configuration, with its own BM25 index:\n\nThen we ran the same query against both indexes. Here it is against the `support_cfg` index:\n\nWith the built-in `english` configuration, one ticket matches:\n\n| Rank | Ticket | BM25 score | \n|---|---|---|\n| 1 | PG creds expired, cannot connect from CI | -4.012 | \n\nWith `support_cfg`, three tickets match:\n\n| Rank | Ticket | BM25 score | \n|---|---|---|\n| 1 | Rotating Postgres credentials in the Zürich region | -5.180 | \n| 2 | PG creds expired, cannot connect from CI | -2.596 | \n| 3 | PostgreSQL connection pool exhausted during the nightly batch | -1.132 | \n\n#### Reading BM25 scores\n\nBM25 gives each document a positive relevance score: 0 means no query terms matched, and higher means more relevant. The `<@>` operator returns that score with a minus sign, so you can sort with `ORDER BY ... ASC`, the same way you sort by distance. A score of -5.180 is a stronger match than -1.132, and 0 means no match.\n\nThe query `k8s pods crashing` shows a subtler effect. Both configurations return the same Kubernetes tickets, but the scores change:\n\n- With the built-in configuration, tickets 1 and 3 tie at -2.477. They match only on \"pods\" and \"crash,\" because their \"Kubernetes\" and \"Kube\" don't match \"k8s.\"\n- With `support_cfg` , all three tickets share the token`kubernetes` . Tickets 1 and 3 rise to -3.727 and -3.331, because they now match on every query word.\n- Ticket 2 drops from -1.879 to -1.068. With the built-in configuration, `k8s` appeared in only one ticket, so BM25 treated it as a rare and valuable term. Once the three spellings merge into one token, that token appears in three tickets and counts for less.\n\n## Why dictionaries matter so much to BM25\n\nBM25 scores a document using three things:\n\n1. how often each query term appears in it (term frequency)\n2. how long the document is\n3. and how rare each term is across all your documents (inverse document frequency, or IDF)\n\n`lakebase_bm25` computes those statistics from the tokens in your `tsvector` column, so whatever your dictionary produces is what BM25 works with.\n\nThat shows up in three ways:\n\n### Spelling variants distort rarity\n\nWhen one concept is split across `kubernet`, `k8s`, and `kube`, each variant looks rarer than the concept really is. A query that uses one variant finds only a fraction of the relevant documents, and it gives them an inflated rarity score. Mapping the variants to one token gives BM25 one term with accurate statistics.\n\n### Filler words break match filters\n\nMany search setups pair BM25 ranking with an `@@` match filter, so a query with no real matches returns nothing instead of a ranked list of weak results. `plainto_tsquery` combines every term with AND, so one stray word sinks the query. With the built-in configuration, \"please help, kube pods crashing\" becomes `'pleas' & 'help' & 'kube' & 'pod' & 'crash'` and matches zero tickets. With `support_cfg`, it becomes `'kubernetes' & 'pod' & 'crash'` and matches two. This matters more now that many queries come from agents and chat interfaces, where people write in full sentences.\n\n### Accents block matches entirely\n\nA query for `zurich` and a ticket containing `zürich` share no token, so there's nothing for any ranking function to score.\n\n## Get started\n\nPoint your agent to the [Lakebase Search docs](https://neon.com/docs/ai/lakebase-search) and ask it to set up a dictionary for your app's vocabulary. If your agent uses the Neon [agent skill](https://neon.com/docs/ai/agent-skills), it already knows all about Lakebase Search:\n\nFor the details, see the [lakebase_tokenizer reference](https://neon.com/docs/extensions/lakebase-tokenizer).", "url": "https://wpnews.pro/news/bm25-is-only-as-good-as-your-tokens", "canonical_source": "https://neon.com/blog/bm25-is-only-as-good-as-your-tokens", "published_at": "2026-10-07 12:00:00+00:00", "updated_at": "2026-10-07 16:16:07.972107+00:00", "lang": "en", "topics": ["ai-search", "structured-data", "developer-tools", "ai-infrastructure"], "entities": ["Neon", "Lakebase Search", "lakebase_tokenizer", "Postgres", "lakebase_vector", "lakebase_text", "lakebase_bm25", "pgvector"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/bm25-is-only-as-good-as-your-tokens", "markdown": "https://wpnews.pro/news/bm25-is-only-as-good-as-your-tokens.md", "text": "https://wpnews.pro/news/bm25-is-only-as-good-as-your-tokens.txt", "jsonld": "https://wpnews.pro/news/bm25-is-only-as-good-as-your-tokens.jsonld"}}