{"slug": "jev-vs-classical-ml-strong-on-sentiment-mixed-across-tasks", "title": "Jev vs. classical ML. Strong on sentiment: Mixed across tasks", "summary": "Jev 1.13.0's largest advantage over eleven classical classification pipelines across eight datasets came on IMDb, where raw zero-shot Jev reached 96.3% balanced accuracy versus 88.4% for logistic regression, a +7.9 percentage point lead that held at 96.1% versus 88.3% after threshold adjustment, according to a Protocol 3.0.1 bounded-budget comparison. Jev also led the best classical mean on SMS Spam, but classical pipelines beat Jev on Banking77, Bank Marketing, Online Shoppers, Breast Cancer, and Iris, and the report states the differences are descriptive rather than statistical-significance claims. The comparison used three training seeds on the same test cases and notes the experiment does not establish superiority over other language models or stronger modern text representations.", "body_md": "We compared Jev 1.13.0 with eleven classical classification pipelines across eight datasets. Its largest advantage was on IMDb. Across the rest of the suite, the picture was more nuanced.\n\nProtocol 3.0.1 · Bounded-budget comparison · Balanced accuracy · Not affiliated with OpenAI or Anthropic\n\n01 / Inspect the full comparison\n\nEvery model. Both decision rules.\n\nSwitch between raw and threshold-adjusted decisions. Jev’s columns appear first. Each cell shows mean ± sample standard deviation across three training seeds on the same test cases. Bold blue cells mark the best displayed mean in each row, including ties.\n\nRaw balanced accuracy: mean ± sample standard deviation, three seeds.\n\nDataset\n\nJev zero-shot\n\nJev few-shot\n\nLogistic regression\n\nSVM\n\nDecision tree\n\nRandom forest\n\nExtra trees\n\nk-NN\n\nNaive Bayes\n\nHist gradient boost\n\nXGBoost\n\nCatBoost\n\nVoting ensemble\n\nMajority baseline\n\nAG News\n\n87.5 ± 0.0\n\n86.3 ± 0.6\n\n87.4 ± 0.8\n\n88.4 ± 0.3\n\n67.5 ± 0.9\n\n71.6 ± 1.1\n\n75.1 ± 0.7\n\n78.4 ± 0.3\n\n87.4 ± 0.2\n\n80.3 ± 0.8\n\n82.6 ± 0.2\n\n82.0 ± 1.0\n\n87.0 ± 1.0\n\n25.0 ± 0.0\n\nBanking77\n\n78.9 ± 0.0\n\n81.9 ± 1.7\n\n89.4 ± 0.2\n\n89.7 ± 0.6\n\n62.2 ± 0.9\n\n68.2 ± 1.4\n\n69.8 ± 0.4\n\n58.4 ± 1.9\n\n86.0 ± 0.5\n\n62.2 ± 1.2\n\n71.9 ± 0.4\n\n68.1 ± 0.7\n\n86.2 ± 0.4\n\n1.3 ± 0.0\n\nSMS Spam\n\n96.1 ± 0.0\n\n95.6 ± 0.9\n\n86.4 ± 8.1\n\n93.7 ± 0.0\n\n85.7 ± 2.8\n\n89.0 ± 0.3\n\n88.7 ± 0.8\n\n87.8 ± 2.8\n\n95.0 ± 1.9\n\n92.1 ± 0.1\n\n88.3 ± 2.8\n\n90.7 ± 1.6\n\n89.0 ± 0.8\n\n50.0 ± 0.0\n\nIMDb\n\n96.3 ± 0.0\n\n95.9 ± 0.5\n\n88.4 ± 0.2\n\n87.8 ± 0.9\n\n70.7 ± 1.1\n\n81.2 ± 0.8\n\n83.3 ± 1.1\n\n79.9 ± 1.5\n\n86.2 ± 0.4\n\n83.2 ± 0.3\n\n83.6 ± 0.2\n\n84.0 ± 0.2\n\n87.3 ± 0.2\n\n50.0 ± 0.0\n\nBank Marketing\n\n53.4 ± 0.0\n\n55.3 ± 3.1\n\n58.0 ± 0.6\n\n71.8 ± 0.7\n\n66.0 ± 4.8\n\n59.3 ± 1.4\n\n60.2 ± 0.3\n\n55.5 ± 0.5\n\n71.0 ± 0.4\n\n63.9 ± 7.4\n\n64.1 ± 7.0\n\n62.3 ± 7.8\n\n61.4 ± 5.5\n\n50.0 ± 0.0\n\nOnline Shoppers\n\n51.4 ± 0.0\n\n54.7 ± 9.5\n\n63.8 ± 11.0\n\n69.1 ± 0.3\n\n60.3 ± 5.1\n\n56.8 ± 6.5\n\n53.3 ± 3.8\n\n51.2 ± 1.3\n\n59.5 ± 0.4\n\n51.8 ± 0.8\n\n65.1 ± 9.5\n\n63.7 ± 10.0\n\n63.5 ± 10.0\n\n50.0 ± 0.0\n\nBreast Cancer\n\n61.0 ± 0.0\n\n88.8 ± 5.5\n\n99.5 ± 0.4\n\n100.0 ± 0.0\n\n94.0 ± 2.2\n\n96.6 ± 1.9\n\n97.4 ± 1.2\n\n100.0 ± 0.0\n\n92.2 ± 0.7\n\n97.3 ± 0.6\n\n97.7 ± 0.4\n\n98.0 ± 0.6\n\n99.1 ± 0.4\n\n50.0 ± 0.0\n\nIris\n\n97.0 ± 0.0\n\n94.5 ± 4.8\n\n100.0 ± 0.0\n\n100.0 ± 0.0\n\n97.5 ± 2.1\n\n98.9 ± 1.9\n\n100.0 ± 0.0\n\n97.8 ± 3.9\n\n100.0 ± 0.0\n\n100.0 ± 0.0\n\n95.7 ± 2.1\n\n97.5 ± 2.1\n\n100.0 ± 0.0\n\n33.3 ± 0.0\n\nScroll horizontally to inspect all columns. Sample SD is not a confidence interval. Zero SD on cached zero-shot predictions is not independent evidence of repeatability.\n\n02 / Compare the results\n\nA different story for each task.\n\nCompare Jev with the highest-scoring classical pipeline on each dataset. The controls in the full table above also update this chart. Compare raw decisions with binary thresholds learned from separate labeled data.\n\nFigure 1. Mean balanced accuracy, with a common 0–100% scale. Values are rounded to one decimal. The classical comparator is the highest test mean among the eleven pipelines in that panel, selected retrospectively. It is not a deployment selection rule. Read uncertainty notes ↓\n\nReading the result\n\nRaw zero-shot Jev leads the best classical mean on IMDb and SMS Spam. The largest lead is IMDb: +7.9 percentage points. These are descriptive differences, not statistical-significance claims.\n\nThe raw chart and complete raw table are available without JavaScript. Enable JavaScript to switch panels, or download the adjusted CSV below.\n\n03 / What changes the interpretation\n\nThe headline is only part of the result.\n\nA\n\nSentiment is the standout.\n\nOn IMDb, raw zero-shot Jev reaches 96.3%, compared with 88.4% for logistic regression. The +7.9-point advantage remains almost unchanged after threshold adjustment: 96.1% versus 88.3%.\n\nThis is the strongest descriptive evidence in Jev’s favor here. The experiment compares a pretrained API model against these classical text pipelines; it does not establish superiority over other language models or stronger modern text representations.\n\nB\n\nThresholds change the SMS story.\n\nWith raw decisions, Jev zero-shot leads Naive Bayes 96.1% to 95.0%. After policy threshold selection, the comparison becomes 95.9% to 96.3%.\n\nThe apparent lead becomes a small deficit. A model’s default decision rule and its ability to separate classes are not the same question. Publish both panels; the adjusted Jev result uses labeled policy data.\n\nC\n\nBusiness tabular tasks remain difficult.\n\nOn Bank Marketing, adjusted zero-shot-prompt Jev reaches 59.7%, against the voting ensemble’s 73.3%. On Online Shoppers, even adjusted few-shot Jev reaches only 53.6%, against 71.2%.\n\nThreshold adjustment does not close these gaps. For context, a constant-class prediction scores 50% balanced accuracy on these binary tasks. This is evidence of weak performance on these particular datasets, not a claim about every tabular task.\n\nD\n\nExamples help selectively.\n\nOne example per class lifts raw Breast Cancer performance from 61.0% to 88.8%, and Banking77 from 78.9% to 81.9%. But it lowers mean performance on AG News, SMS Spam, IMDb, and Iris.\n\nBreast Cancer also improves to 88.4% using zero-shot prompts with a learned threshold alone. Its raw 61.0% score therefore does not tell the whole story. The holdout is small, and none of these Jev variants reaches the strongest classical result.\n\nOne metric, carefully read\n\nBalanced accuracy is the average recall across classes. It gives each class equal weight, even when most examples belong to one class. For 77-class Banking77, the constant-class baseline is about 1.3%; for a binary task it is 50%. These are different problems, so we avoid collapsing this suite into one overall score.\n\n04 / Experimental design\n\nOne holdout. Separate decisions.\n\nThe protocol separates model selection from decision-threshold selection. All models are evaluated on the same test cases for a dataset. The test set stays fixed across training seeds.\n\n01 / Learn\n\nTraining\n\nUp to 8,000 rows. Classical models fit here; Jev’s few-shot examples are drawn from here.\n\n02 / Select\n\nValidation\n\nUp to 1,000 rows. Four candidate configurations per classical family are compared.\n\n03 / Adjust\n\nPolicy\n\nUp to 500 labeled rows. Binary thresholds are selected independently of the test set.\n\n04 / Evaluate\n\nTest\n\nOne fixed holdout per dataset, shared by models and seeds. Small datasets retain fewer rows.\n\nThe classical side\n\nLogistic regression, SVM, decision tree, random forest, extra trees, k-NN, Naive Bayes, histogram gradient boosting, XGBoost, CatBoost, and a voting ensemble. A majority baseline is also reported.\n\nThe V3 speed preset caps trees at 150 and histogram iterations at 60. Text vocabulary and representation budgets are limited. CPU and GPU implementations are explicitly mixed; backend changes are not claimed to be numerically equivalent.\n\nThe Jev side\n\nThe requested model is jev-1.13.0. Zero-shot prompts provide task and class descriptions; few-shot prompts add one labeled training example per class. Tabular rows are passed as structured feature values.\n\nSuccessful identical API requests are cached. Exhausted request failures count as incorrect predictions. Binary adjusted results use labeled policy data, even when their prompts contain no examples.\n\nHeld-out test sizes\n\nDatasets\n\nTest rows\n\nClasses\n\nAG News\n\n1,000\n\n4\n\nBanking77\n\n1,500\n\n77\n\nSMS Spam, IMDb, Bank Marketing, Online Shoppers\n\n1,000 each\n\n2 each\n\nBreast Cancer\n\n114\n\n2\n\nIris\n\n30\n\n3\n\nTraining seeds: 2027, 2028, 2029. Holdout seed: 20260920. Kaggle configuration: two T4 GPUs. Source and protocol are frozen inside the completed notebook. Read the full protocol ↗\n\n05 / Scope of the evidence\n\nWhat this release can’t establish.\n\nStatistical significance\n\nThe displayed ± values measure variation across training seeds. They are not confidence intervals. Paired bootstrap intervals were computed by the reporting code, but their output was not supplied with the notebook. We make no significance claims from these tables.\n\nIndependent zero-shot repetitions\n\nIdentical successful requests are reused from cache across seeds. The three zero-shot rows therefore do not represent three independent API replications. Policy splits can still produce different adjusted thresholds.\n\nPure model quality on Banking77\n\nThe saved Jev run includes warnings about predictions outside the true label set. The adapter uses −1 for failed requests and counts them as incorrect. Without the run diagnostics, we cannot quantify how much of the score reflects API failures.\n\nGeneralization from tiny holdouts\n\nIris has only 30 test cases and Breast Cancer has 114. Perfect classical scores on these samples do not imply perfect performance on new data. Public-dataset pretraining exposure is neither established nor ruled out.\n\nBest possible ML performance\n\nFour candidates and restricted feature, training, and tree budgets are practical constraints. This suite does not cover all classical tuning strategies, pretrained embedding pipelines, fine-tuned transformers, or other language-model APIs.\n\nLatency, cost, or full reproducibility\n\nV3 latency summaries and the full Kaggle results archive are unavailable here. The run records 38,922 request attempts; its approximately $4.19 input-cost estimate is not an invoice. Exact snapshots, per-example predictions, and diagnostics are not included.\n\n06 / Open artifacts\n\nFollow the numbers back to the run.\n\nThe completed notebook is preserved byte for byte. Both CSVs are extracted from its saved final outputs, retaining their displayed rounding. The interactive figures use those same tables.", "url": "https://wpnews.pro/news/jev-vs-classical-ml-strong-on-sentiment-mixed-across-tasks", "canonical_source": "https://quicqdev.github.io/Jev-vs-ML/", "published_at": "2026-09-20 10:14:31+00:00", "updated_at": "2026-09-20 10:53:05.280332+00:00", "lang": "en", "topics": ["machine-learning", "natural-language-processing", "ai-research"], "entities": ["Jev", "Jev 1.13.0", "IMDb", "Banking77", "SMS Spam", "AG News", "OpenAI", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/jev-vs-classical-ml-strong-on-sentiment-mixed-across-tasks", "markdown": "https://wpnews.pro/news/jev-vs-classical-ml-strong-on-sentiment-mixed-across-tasks.md", "text": "https://wpnews.pro/news/jev-vs-classical-ml-strong-on-sentiment-mixed-across-tasks.txt", "jsonld": "https://wpnews.pro/news/jev-vs-classical-ml-strong-on-sentiment-mixed-across-tasks.jsonld"}}