{"slug": "jev-for-analytics-turn-text-columns-into-gold-in-plain-sql", "title": "Jev for analytics: turn text columns into gold, in plain SQL", "summary": "TypeSafe's Jev classified 100,000 complaint narratives in 82.2 seconds using its prompt_jev() function, with no LLM generating the per-row answers, according to a launch benchmark that put Jev at $0.50 per 100,000 rows versus $1.58 for gpt-5-nano and $37.58 for GPT-5.6 Terra. On the same complaints, gpt-5-nano through prompt() needed 6 minutes 9 seconds for only 10,000 rows, where Jev took 12.5 seconds. Jev is described as a \"System One\" model that reads text once and scores a fixed list of allowed answers, returning a typed label, probability, or score with a confidence value instead of free text.", "body_md": "Every now and then, we all end up with a table that has a big text column. A comment, some customer feedback, or worse, an address nobody parsed upstream. There are insights in that column, but they're hard to get out. And sometimes it's plain extra compute you keep paying for: the address never got parsed into typed columns, so every query parses text again, and parsing text is $$$.\n\nLet's take a dataset with 100,000 complaints. Each row has a narrative (someone explaining what went wrong with a bank, a lender, or their credit report) and how the company responded. The column I actually want is missing: what was the person asking for?\n\nIt sounds like a job for an LLM. But running an LLM over every row just to pick one of seven known answers is like using a cannon to kill a fly. It's slow on a big table, so it scales badly, or it just burns your credit card.\n\nprompt_jev() classified the 100,000 rows in 82.2 seconds, with no LLM generating the per-row answers. On the same complaints, gpt-5-nano through prompt() needed 6 minutes 9 seconds for only 10,000 rows (Jev: 12.5 seconds). On cost, the launch benchmark put Jev at $0.50 per 100,000 rows, against $1.58 for gpt-5-nano and $37.58 for GPT-5.6 Terra (API prices). That's about 1% of the frontier model's bill.\n\nFaster and cheaper? Let's look at what Jev does, run it on the dataset, and see where I would (and wouldn't) trust the result.\n\nWhat Jev does\n\nTypeSafe calls Jev a \"System One\" model, borrowing from Kahneman's Thinking, Fast and Slow. System 1 is the quick, intuitive answer. System 2 is the slow, deliberate one. A classic LLM is System 2: it reasons and writes its answer token by token. Jev is System 1: it reads the text once and scores the answers you allowed.\n\nSystem 2: a classic LLM\n\nSystem 1: Jev\n\nHow it answers\n\nWrites text, token by token\n\nScores a fixed list of answers in one pass\n\nWhat you get back\n\nFree text (you parse it, or force a schema)\n\nA typed label, probability, or score\n\nCan invent a new answer\n\nYes\n\nNo, only the options you give it\n\nTells you how sure it is\n\nNot directly\n\nA probability per option, plus a confidence\n\n100k rows benchmark\n\n18 to 32 minutes, $1.58 to $37.58\n\n40 seconds, $0.50\n\nGood at\n\nNew taxonomies, summaries, explanations\n\nThe same bounded decision over many rows\n\nExample: customer request: \"Please refund the two fees you charged me.\"\n\nAnswer is something like \"The customer is asking for a refund of the two fees...\" (a sentence to parse)\n\nFor SQL, think of it as a smart decision function: you give it one row of text, a question, and the allowed answers. You get back a typed decision, plus a number that tells you how sure it is.\n\nThere are three question shapes. Here they are on the same real complaint (10306039, someone charged low balance fees on an account they had closed):\n\nShape\n\nYou ask\n\nYou get back\n\nReal answer\n\nSQL use\n\nChoice\n\n\"What outcome is the consumer asking for?\" + 7 labels\n\none label + confidence\n\nrefund_or_charge_reversal, 0.95\n\nGROUP BY\n\nNoul\n\n\"The complaint is about fees the bank charged.\"\n\na probability, 0 to 1\n\n0.97 (and 0.02 for \"reports identity theft\")\n\nWHERE\n\nScore\n\n\"How urgent is this?\" on low / medium / high\n\na position on the scale + confidence\n\n1.62 (0 = low, 2 = high), 0.43\n\nORDER BY\n\nThis walkthrough uses Choice, since each complaint should land in one request category.\n\nHow does that compare to what you already use?\n\nAn LLM is the right tool for a new taxonomy, a summary, or an explanation. A schema can hold it to your labels, but it still writes each one out token by token. Jev picks from your list, so you always get back one of your strings.\n\nCASE WHEN or a regex only works when the rule is literal. People write \"reverse the charge\" instead of \"refund\", and mentioning money isn't always asking for it.\n\nA trained classifier scales well but is painful to run (labelled examples, a model to maintain). With Jev, you describe the categories in the query and try them right away.\n\nParsing rules (JSON paths, address parsers) are static and break when the format changes. Jev works from what the text says, so messy input hurts it less.\n\nJev can only return one of the options you gave it, so poor options get you a neatly typed poor answer. Garbage in, garbage out.\n\nAbout the dataset: complaints from CFPB\n\nThe complaints come from the U.S. Consumer Financial Protection Bureau (CFPB's archive), the agency that collects complaints about banks, lenders, credit bureaus, and other financial services.\n\nMy table complaints_100k has (spoiler alert) 100,000 rows with complaint_id, narrative, company_response, and a few other fields. Here's what four rows look like:\n\ncomplaint_id\n\nproduct\n\nnarrative (trimmed)\n\ncompany_response\n\n9906201\n\ncredit reporting\n\n\"This account is incorrectly being reported as a charged off account with a balance due. Please provide proof of the charge off and update the balance to {$0.00}...\"\n\nClosed with explanation\n\n10306039\n\nbank account\n\n\"I closed my Chase checking account in XXXX by sending Chase a letter via mail. Chase conveniently ignored the account closure request so they can charge low balance fees...\"\n\nClosed with monetary relief\n\n6084737\n\ndebt collection\n\n\"I have no outstanding debts owed that I am aware of. Americollect has called numerous times almost daily...\"\n\nClosed with explanation\n\n2042216\n\ncredit card\n\n\"I paid my credit card with a money order. They did not apply the payment to my account...\"\n\nClosed with monetary relief\n\nLet's get the gold out of that narrative column.\n\nStep 1: define the labels, then classify\n\nBefore classifying 100,000 rows, we need the categories. I sampled 40 narratives and asked MotherDuck's prompt() function, running gpt-5-mini, to suggest seven requested-outcome categories. That took 7.6 seconds. How you sample matters (and you can take a bigger one), but that's a business question more than a technical one.\n\nThen Jev classifies every row against these categories, in less than 2 minutes:\n\nJev requires the question and the choices to be constants, so the simplest way is to paste the taxonomy straight into choice :=. Here's the classification step, with the same labels and descriptions as the measured run:\n\n```\nCREATE TABLE complaint_requests AS\nSELECT complaint_id, company_response,\n       prompt_jev(\n         narrative,\n         'What outcome is the consumer asking for?',\n         choice := [\n           {label: 'refund_or_charge_reversal',\n            description: 'Consumer requests money returned, reversal of fees/charges, or reimbursement.'},\n           {label: 'correct_or_remove_credit_report_entry',\n            description: 'Consumer asks to fix, update, or remove inaccurate items on their credit report.'},\n           {label: 'investigate_and_credit_transaction_error',\n            description: 'Consumer requests investigation and correction of a transaction or deposit error and credit to their account.'},\n           {label: 'stop_harassment_or_contact',\n            description: 'Consumer asks for communications or collection calls/texts/emails to stop or follow rules.'},\n           {label: 'address_identity_theft_or_fraud',\n            description: 'Consumer requests action to resolve identity theft, fraudulent accounts, or unauthorized inquiries.'},\n           {label: 'explain_decision_or_account_status',\n            description: 'Consumer asks for clarification or explanation of a decision, account action, or why charges occurred.'},\n           {label: 'other',\n            description: 'Any requested outcome that does not fit the categories above.'}\n         ]\n       ) AS request\nFROM complaints_100k;\n```\n\nrequest is a struct. request.choice is the picked label, and request.confidence and request.probabilities help you inspect the decision. Here's what came back for the Chase fees complaint (10306039):\n\n```\n{\n  choice: 'refund_or_charge_reversal',\n  confidence: 0.95,\n  probabilities: [\n    {value: 'refund_or_charge_reversal', probability: 0.96},\n    {value: 'other', probability: 0.03},\n    {value: 'explain_decision_or_account_status', probability: 0.01},\n    ... -- the other 4 labels at 0.00\n  ]\n}\n```\n\nBonus: keep the labels in a table\n\nPasting 7 labels into a query is fine once. If the taxonomy sticks around (versioned, shared across queries, edited by the business), you'd rather keep it in a table. A subquery in choice := doesn't work (\"choice\" parameter must be a constant value), but a DuckDB variable does, because it's resolved before the query runs:\n\n```\nSET VARIABLE request_labels = (\n  SELECT list({label: label, description: description} ORDER BY label)\n  FROM request_taxonomy\n);\n\nSELECT complaint_id,\n       prompt_jev(narrative, 'What outcome is the consumer asking for?',\n                  choice := getvariable('request_labels')) AS request\nFROM complaints_100k;\n```\n\nWhy bother: one source of truth for the labels, easy to version (add a taxonomy_version column), and the business can review the labels without reading SQL.\n\nStep 2: don't trust Jev blindly\n\nJev can't invent a label, but it can still pick the wrong one.\n\nThe first thing to look at is confidence. Jev scores every label, and confidence says how concentrated those scores are: 1.0 when one label takes everything, low when two or three labels share the probability.\nAcross the 100,000 rows, the mean confidence is 0.83, 65.8% of rows are at 0.8 or above, and 10.9% are below 0.5.\n\nSQL gives you the review queue:\n\n```\nSELECT complaint_id, request.choice, request.confidence\nFROM complaint_requests\nWHERE request.confidence < 0.8\nORDER BY request.confidence;\n```\n\n0.8 is an example, and a row above it can still be wrong. Send uncertain rows to a human or a bigger model, and keep the rest in the analytics table once you've validated that cut on your data.\n\nThe second check is to compare against another model on a sample. On 300 rows, I labelled them twice: with GPT-5 through prompt() (the newest model MotherDuck's prompt() offers today) and with Claude Opus 5.5, which read each narrative blind, without seeing any other label. Here's what came out:\n\nComparison on 300 rows\n\nAgreement\n\nGPT-5 vs Claude Opus 5.5\n\n72.7%\n\nJev vs GPT-5\n\n77.3%\n\nJev vs Claude Opus 5.5\n\n74.0%\n\nJev, on the 218 rows where GPT-5 and Opus agree\n\n90.4%\n\nJev, on those rows with confidence >= 0.8 (161 rows)\n\n96.9%\n\nThe two big LLMs agree with each other less often than Jev agrees with either of them. So the task itself is fuzzy: most disagreements sit on the boundaries we already saw (identity theft vs credit-report fix, \"other\" vs a specific label).\n\nStep 3: ask the question (the good boring part)\n\nOnce the requested outcome is a typed column, it's plain SQL:\n\n```\nSELECT request.choice AS requested_outcome,\n       count(*) AS complaints,\n       round(100.0 * avg(\n         (company_response = 'Closed with monetary relief')::INT\n       ), 1) AS pct_monetary_relief\nFROM complaint_requests\nGROUP BY 1\nORDER BY complaints DESC;\n```\n\nIt runs in 0.9 seconds on MotherDuck. The 82.2 seconds of classification happen once, and every question after that costs what any other GROUP BY costs.\n\nWhen I would use this pattern\n\nUse prompt_jev() when the same decision repeats across many rows and you can write the possible answers down: complaint intent, product family, \"did they ask for a refund?\", a severity rubric. Use an LLM when you need new categories, generated text, or an explanation a human will read. Or use both, like here: the LLM helps design the question, Jev applies it at scale.\n\nThe MotherDuck launch benchmark classified 100,000 short news articles in 40 seconds. My 82.2-second complaint run is a different task (longer narratives, seven described choices), so I wouldn't promise either number on your table. Test on a sample first :)\n\nprompt_jev() is available on MotherDuck paid plans. The text is sent to TypeSafe for inference, so check your data-handling requirements before passing sensitive customer data. The guide and function reference have the full syntax.\n\nDefine your dashboards in YAML with dbt Charts, then the same dbt build that runs your pipeline deploys them to MotherDuck Dives. Just 2 config changes! Your Dives will refresh live whenever your stakeholders take a look. See how we constructed the dbt package as well.\n\nCommon Crawl publishes petabytes of web crawl data on S3. With DuckDB and MotherDuck you can query the Common Crawl dataset directly, no cluster and no download, and measure how fast the vibe-coded web is growing.", "url": "https://wpnews.pro/news/jev-for-analytics-turn-text-columns-into-gold-in-plain-sql", "canonical_source": "https://motherduck.com/blog/jev-for-analytics", "published_at": "2026-09-29 00:00:00+00:00", "updated_at": "2026-09-29 20:50:12.210242+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["TypeSafe", "Jev", "prompt_jev()", "gpt-5-nano", "GPT-5.6 Terra", "Kahneman"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/jev-for-analytics-turn-text-columns-into-gold-in-plain-sql", "markdown": "https://wpnews.pro/news/jev-for-analytics-turn-text-columns-into-gold-in-plain-sql.md", "text": "https://wpnews.pro/news/jev-for-analytics-turn-text-columns-into-gold-in-plain-sql.txt", "jsonld": "https://wpnews.pro/news/jev-for-analytics-turn-text-columns-into-gold-in-plain-sql.jsonld"}}