{"slug": "not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions", "title": "Not every AI call needs to generate text: ej, an 11 MB model for typed decisions", "summary": "A developer has released ej, an 11 MB open model that answers typed decision questions — choice, yes/no, and ordered score — by returning one probability distribution per question in a single forward pass without generating any tokens. The model, installable via pip as ejai (version 0.0.1), runs on CPU and lets callers threshold on probabilities rather than parse generated JSON, positioning it as a lighter alternative to prompting an LLM or running per-pair NLI zero-shot classifiers.", "body_md": "A lot of \"AI\" in production pipelines isn't writing anything. It's deciding.\n\nWhich queue does this ticket go to? Does this message need a human? How urgent is this request: low, medium or high? The\n\nanswer is one item from a list you already know. Yet the usual way to get it is to send the text to a language model,\n\nask it to reply in JSON, parse the reply, retry when the JSON is broken, and hope the word \"high\" means the same thing\n\nevery time.\n\nThat works, but it's an odd fit. You pay for generation you don't need, you get text where you wanted a number, and you\n\nhave no clean way to say \"only automate this when the model is at least 90% sure.\"\n\nI built **ej** to do that one job differently. It's an open model that takes a piece of text plus a set of typed\n\nquestions, and returns **one probability distribution per question, in one forward pass, without generating a single token**. The whole model is one 11 MB file that runs on a CPU.\n\n`pip install ejai` (the import name is `ej`)\nThis is version 0.0.1. Below is what it does, how I measured it, and, just as important, where it is weak.\n\nYou give ej a `state` (free text, or a JSON object written as text) and any number of questions. Each question has one\n\nof three types:\n\n| Type | You give | You get | \n|---|---|---|\n| `choice` | instructions + a list of option texts, chosen when you call it | a probability for each option | \n| `noul` | a yes/no statement | P(false), P(true) | \n| `score` | instructions + ordered levels | a distribution over the levels | \n\nThe options are plain text you write at call time. There is no fixed label set baked in at training.\n\n``` python\nimport ej\n\nmodel = ej.load(\"5ak3t/ej\", revision=\"v0.0.1\")   # downloads model.ejpack and checks its SHA-256\n\nrecord = {\n    \"state\": '{\"customer_tier\": \"gold\", \"message\": \"The blender arrived with a cracked jug. Replace it before Friday.\"}',\n    \"questions\": {\n        \"route\": {\"type\": \"choice\", \"instructions\": \"Which team should handle this request?\",\n                  \"options\": [{\"key\": \"returns\", \"text\": \"Returns and replacements for damaged or wrong items\"},\n                              {\"key\": \"billing\", \"text\": \"Billing, invoices and payment problems\"}]},\n        \"needs_human\": {\"type\": \"noul\",\n                        \"instructions\": \"The customer is upset enough that a human agent should reply.\",\n                        \"options\": ej.NOUL_OPTIONS},\n        \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent is this request?\",\n                    \"options\": [{\"key\": \"0\", \"text\": \"Low: can wait a week\"},\n                                {\"key\": \"1\", \"text\": \"Medium: answer within two days\"},\n                                {\"key\": \"2\", \"text\": \"High: answer today\"}]},\n    },\n}\n\n(probs,) = model.predict([record])\nprint(probs[\"route\"])        # [p_returns, p_billing], sums to 1\nprint(probs[\"needs_human\"])  # [p_false, p_true]\nprint(probs[\"urgency\"])      # [p_low, p_medium, p_high]\n```\n\nAll three questions are answered in the same pass. Because the output is a distribution, the routing logic in your code\n\nis a threshold, not a string parser:\n\n```\nif max(probs[\"route\"]) >= threshold:\n    auto_route(record, probs[\"route\"])\nelse:\n    send_to_human(record)\n```\n\n(Pick that threshold on labelled examples of *your* workflow. More on why below.)\n\nThere are three common ways to make this kind of decision today. ej is a different trade-off from each, not a strict\n\nupgrade.\n\n**Prompting an LLM.** It's flexible and strong on new tasks. But it generates text you have to parse, it needs either an\n\nAPI call (cost, latency, your data leaving your servers) or a local model that is hundreds of MB to several GB, and its\n\nreply isn't a probability you can threshold. ej skips generation entirely and is small enough to ship inside a service.\n\n**NLI zero-shot classifiers** (the \"zero-shot-classification\" pipeline). These are the closest relatives. They usually\n\nrun one forward pass per (text, option) pair, so cost grows with the number of options, and each question type needs its\n\nown setup. ej reads the text once and answers choice, yes/no and score questions together. To be clear, though, a good\n\nNLI model beats ej on workflows ej has never seen (numbers below).\n\n**Fine-tuning your own classifier** (BERT, SetFit and friends). It's accurate on its own task, but the label set is fixed\n\nat training, and every change means new labelled data and a retrain. With ej the options are input. It also has an\n\n`adapt` method that takes a handful of labelled records:\n\n```\nadapted = model.adapt(examples=labelled[:8])   # per-option offsets; the weights don't change\nprobs = adapted.predict(new_records)\n```\n\nBriefly (a technical report with the full method is coming):\n\n`e5-small-v2`, quantised to `model.ejpack`: On the test box (a shared 4-core Xeon, one thread), a warm call takes roughly **200 to 360 ms per record**, depending on\n\nthe run and the inputs, and the process peaks at about **558 MB** of RAM. Most of that is torch itself: the model's own tensors peak at 128 MB, or 35 MB with\n\n`ej.load(..., low_memory=True)`, which is about 1.5x slower and gives identical predictions.\n\nI tried hard to make these numbers trustworthy rather than flattering:\n\n| Test suite | ej 0.0.1 accuracy [95% CI] | \n|---|---|\n| Typed Decisions, workflows seen in training | .721 [.695, .746] | \n| MASSIVE intents (leak-free split) | .832 [.808, .857] | \n| Support tickets (in-house) | .712 [.679, .746] | \n| Support tickets, new writing styles | .683 [.639, .726] | \n| **105 workflows never seen in training** | **.419 [.379, .458]** | \n| One held-out Typed Decisions workflow (also used during model selection) | .422 [.388, .452] | \n\nHow that compares with the six rivals:\n\nA few more things you should know before using it:\n\nej makes sense if all of these are true:\n\nIf your task is brand-new and you have no labelled data at all, a prompted LLM or an NLI zero-shot model will probably\n\nserve you better today. I'd rather say that here than have you find out in production.\n\n```\npip install ejai\npython\nimport ej\nmodel = ej.load(\"5ak3t/ej\", revision=\"v0.0.1\")\n```\n\nThe repository also contains the training code (`python -m ej.train`), the evaluator that produced every number above\n\n(`python -m ej.eval`), and the benchmark runner with the rival adapters, so you can check my numbers or train your own\n\npack.\n\nIf you run it on your own workflow, I'd love to hear what accuracy you get, good or bad. Results on real workflows are\n\nexactly what version 0.0.2 needs.", "url": "https://wpnews.pro/news/not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions", "canonical_source": "https://dev.to/5ak3t/not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions-3ko9", "published_at": "2026-10-09 08:42:51+00:00", "updated_at": "2026-10-09 08:51:28.216497+00:00", "lang": "en", "topics": ["ai-tools", "machine-learning", "natural-language-processing", "ai-products", "developer-tools"], "entities": ["ej", "ejai", "5ak3t/ej"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions", "markdown": "https://wpnews.pro/news/not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions.md", "text": "https://wpnews.pro/news/not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions.txt", "jsonld": "https://wpnews.pro/news/not-every-ai-call-needs-to-generate-text-ej-an-11-mb-model-for-typed-decisions.jsonld"}}