{"slug": "i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no", "title": "I Tried to Make Natural Language a SQL Predicate. The Numbers Said No.", "summary": "A developer built SemPred, a prototype that adds natural-language predicates like SEM_PREDICT(text, 'customer is asking for a refund') to SQL queries using a frozen all-MiniLM-L6-v2 encoder with per-predicate logistic-regression heads and a three-valued TRUE/FALSE/UNKNOWN output. Benchmarked on Banking77, the 16-shot MiniLM setup reached 72.16% accuracy versus 60.25% for matched TF-IDF, but a full-data TF-IDF baseline hit 85.45%, so the developer shelved the model for missing a pre-set bar while keeping the approach.", "body_md": "A few weeks ago I got stuck on a question: what if part of a `WHERE` clause could be written in plain English?\n\n```\nSELECT *\nFROM support_tickets\nWHERE SEM_PREDICT(\n    ticket_text,\n    'customer is asking for a refund'\n);\n```\n\nI didn't want an LLM writing SQL for me, and I didn't want a chatbot. I also didn't want to send every row to an API. I wanted the natural-language condition to behave like any other predicate the engine can evaluate.\n\nI built a prototype called **SemPred**, benchmarked it, and stopped when it missed the bar I'd set beforehand. The idea holds up. My first model doesn't. This post covers both.\n\nCode: [github.com/Yudeeswaran/SemPred](https://github.com/Yudeeswaran/SemPred)\n\nDatabases are great at structured predicates. `amount > 1000 AND currency = 'USD'` is deterministic, cheap, indexable, and easy to reason about.\n\nPlenty of real data doesn't fit that shape. Take a ticket table:\n\n| id | ticket_text | \n|---|---|\n| 1 | The ATM kept my card | \n| 2 | I don't recognize this transaction | \n| 3 | Can I get my money back for this charge? | \n| 4 | My card hasn't arrived yet | \n\n\"Find the tickets where the customer wants a refund\" has no SQL operator. You can write keyword rules, pre-classify everything, or call an LLM per row. I wanted to see if there was a fourth option, where the semantic condition is just another function in the query.\n\nTwo functions:\n\n`SEM_SCORE(text, predicate)` returns a score.`SEM_PREDICT(text, predicate)` returns `TRUE`, `FALSE`, or `UNKNOWN`.\n\n```\nSELECT\n    ticket_text,\n    SEM_SCORE(ticket_text, 'customer is asking for a refund')   AS score,\n    SEM_PREDICT(ticket_text, 'customer is asking for a refund') AS decision\nFROM tickets;\n```\n\nThe third state matters most. In DuckDB, `UNKNOWN` maps to SQL `NULL`, so the engine can tell \"the model thinks this is false\" apart from \"the model has no idea\". I'd rather get a `NULL` than a confident wrong answer, and three-valued logic already gives SQL a place to put it.\n\nThe SQL stays SQL. The semantic part is one more predicate next to your joins and filters.\n\nThe constraints I set up front:\n\n**The model.** A frozen `all-MiniLM-L6-v2` encoder turns text into an embedding. Each predicate gets its own small logistic-regression head on top of it. The encoder isn't trained at all.\n\n**Embed once, score many.** If one ticket needs checking against 50 predicates, the encoder runs once and the 50 heads run on the same vector. Encoding is the expensive part, and the heads are nearly free by comparison.\n\n**A bounded embedding cache.** Repeated text skips the encoder. This is a small detail in a notebook, but over millions of rows it's an architecture decision, and it was the point where the project started to feel like a data-systems problem.\n\n**DuckDB integration.** Where PyArrow is available, the functions use DuckDB's Arrow UDF path, so values arrive in batches. The naive alternative (DuckDB → Python → model → DuckDB, once per row) looks fine at 100 rows and falls over at scale.\n\n**Abstention.** The decision isn't `p >= 0.5`. There's a margin around the threshold. With a threshold of 0.5 and a margin of 0.1:\n\n| score | result | \n|---|---|\n| < 0.4 | FALSE | \n| 0.4 to 0.6 | UNKNOWN | \n| ≥ 0.6 | TRUE | \n\nBoth numbers are configurable.\n\nI used Banking77: 10,003 training examples, 3,080 test examples, 77 intents, official split. I framed it as a few-shot predicate problem. Given a handful of labeled examples per predicate, can the system decide whether unseen text satisfies it?\n\nBefore any neural model, I ran TF-IDF. If a lexical baseline gets most of the way there, embeddings aren't earning their cost.\n\n| Examples per predicate | TF-IDF | MiniLM | \n|---|---|---|\n| 8 | 47.90% | 62.70% | \n| 16 | 60.25% | 72.16% | \n\nAt 16 examples, MiniLM beats the matched TF-IDF setup by about 11.9 points. That looked like a win until I trained TF-IDF on all the training data:\n\n| Setup | Accuracy | \n|---|---|\n| Full-data TF-IDF | 85.45% | \n| 16-shot MiniLM | 72.16% | \n\n\"Improves accuracy by 12 points\" was true and also misleading. The useful question is always \"compared to what?\"\n\nI swapped in MPNet with the same frozen-encoder setup. Its best 16-shot result was **74.93% ± 1.77 pp**. That's better than MiniLM, but below the 80% continuation threshold I'd defined before running it.\n\nIt was also much slower. Encoding 13,072 unique texts:\n\n| Encoder | Time | Throughput | \n|---|---|---|\n| MiniLM | 18.24 s | ~717 texts/s | \n| MPNet | 133.06 s | ~98 texts/s | \n\nThat's about 7× slower. A slower model can be worth it if quality jumps, but here it didn't.\n\nAccuracy alone wasn't the real problem. The real problem was how accuracy trades against coverage.\n\nAt an abstention margin of 0.10, the 16-shot system committed to a decision on about **44.1%** of ticket-predicate pairs. Of those committed decisions, about **96.3%** matched the binary labels. That looks strong until you notice it's answering less than half the questions.\n\nWidening the margin pushes accuracy up and coverage down. At one operating point I got roughly 90% ticket accuracy, but coverage fell to about 21%.\n\nBefore running anything, I'd written a gate:\n\n**≥ 90% committed-ticket accuracy AND ≥ 50% ticket coverage**\n\nNo margin I evaluated met both. You can make almost any model look accurate by letting it answer fewer questions, so every abstention number needs its coverage figure next to it.\n\nThis is where I could have kept trying encoders and margins until something looked good. I didn't, because the gate existed to prevent exactly that. The gate wasn't met, so the conclusion is that this architecture isn't production-ready.\n\nHaving a written stopping condition changed how I ran everything. \"Is this model better?\" became \"Did it meet the requirement?\", and the second question is much harder to fudge.\n\nAverages hide a lot. The *card swallowed* predicate did far better than the overall numbers, with the 16-shot model at the selected abstention setting:\n\nMatched low-shot TF-IDF scored around 35.7% F1 on the same predicate. Some predicates have distinctive language, while others overlap heavily with neighboring intents. A real system probably shouldn't treat every predicate as equally hard.\n\nIn the current design, the predicate is the classifier head. The text goes through the encoder, and a per-predicate model decides. But the task is really pairwise: *does this text satisfy this predicate?* The predicate's wording carries information the head never sees.\n\n```\nText:      \"My account was charged twice.\"\nPredicate: \"customer is reporting a duplicate charge\"\n```\n\nThe relationship between those two strings is the signal. So the next hypothesis is to make the predicate an input to the model, not the name of a head.\n\nThis is a hypothesis to test, not a result. I haven't claimed it works.\n\nAn LLM given `(ticket, predicate)` would likely reason better. It also brings cost, latency, throughput, hosting, batching, privacy, and determinism concerns. At a million rows, a million LLM calls is hard to justify.\n\nI see it more as a teacher or an escalation path. A compact model handles the bulk, and only the `UNKNOWN` cases go to a bigger model.\n\nThat's future work as well.\n\nThe systems side held up better than the quality side. In the benchmark environment, the DuckDB path scored **237,160 predicate pairs in 5.57 seconds**, about **42,582 pairs/sec**.\n\nThat's possible because unique texts are embedded once and the embeddings are reused across predicates. With `N` unique texts and `P` predicates, the cost is one encoder pass over `N` texts plus `N × P` cheap head evaluations, instead of `N × P` encoder runs.\n\nI started out thinking the hard part was classifying text. After the prototype I'm less sure. The harder part may be running uncertain predicates safely over large tables:\n\n`NOT`?\nNone of those are NLP questions.\n\nI'm not going to try five more embedding models, because that's benchmark roulette. The next experiment changes the architecture (pairwise scoring with the predicate as input), judged against the same gate and the same evaluation setup. After that, in rough order:\n\nNothing on that list counts as progress until it beats a benchmark defined in advance.\n\nAs a product, no. The model I evaluated misses its own quality and coverage bar.\n\nAs an experiment, yes. It turned \"can SQL take a natural-language condition?\" into a more specific question: what would it take for semantic reasoning to be a measurable, cacheable, safely executable primitive inside a data system? I don't have the full answer, but I now have a prototype, a benchmark, performance numbers, and a clear picture of where the design breaks.\n\nThe repo is open if you want to reproduce the benchmark or tell me where the architecture is wrong: [github.com/Yudeeswaran/SemPred](https://github.com/Yudeeswaran/SemPred)\n\nIf you work on semantic query execution, database-native ML, DuckDB extensions, or cheap inference at scale, I'd like to hear from you.\n\n*SemPred is a research prototype. It hasn't been validated as a general-purpose classifier or a production decision system, and the current benchmark results don't support using it for high-impact decisions.*", "url": "https://wpnews.pro/news/i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no", "canonical_source": "https://dev.to/yudeeswaran/i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no-16l3", "published_at": "2026-09-30 19:56:53+00:00", "updated_at": "2026-09-30 20:16:41.827087+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-tools", "developer-tools"], "entities": ["SemPred", "DuckDB", "all-MiniLM-L6-v2", "Banking77", "PyArrow", "Yudeeswaran"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no", "markdown": "https://wpnews.pro/news/i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no.md", "text": "https://wpnews.pro/news/i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no.txt", "jsonld": "https://wpnews.pro/news/i-tried-to-make-natural-language-a-sql-predicate-the-numbers-said-no.jsonld"}}