cd /news/machine-learning/jev-vs-classical-ml-strong-on-sentim… · home topics machine-learning article
[ARTICLE · art-135048] src=quicqdev.github.io ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Jev vs. classical ML. Strong on sentiment: Mixed across tasks

Jev 1.13.0's largest advantage over eleven classical classification pipelines across eight datasets came on IMDb, where raw zero-shot Jev reached 96.3% balanced accuracy versus 88.4% for logistic regression, a +7.9 percentage point lead that held at 96.1% versus 88.3% after threshold adjustment, according to a Protocol 3.0.1 bounded-budget comparison. Jev also led the best classical mean on SMS Spam, but classical pipelines beat Jev on Banking77, Bank Marketing, Online Shoppers, Breast Cancer, and Iris, and the report states the differences are descriptive rather than statistical-significance claims. The comparison used three training seeds on the same test cases and notes the experiment does not establish superiority over other language models or stronger modern text representations.

read8 min views1 publishedSep 20, 2026
Jev vs. classical ML. Strong on sentiment: Mixed across tasks
Image: source

We compared Jev 1.13.0 with eleven classical classification pipelines across eight datasets. Its largest advantage was on IMDb. Across the rest of the suite, the picture was more nuanced.

Protocol 3.0.1 · Bounded-budget comparison · Balanced accuracy · Not affiliated with OpenAI or Anthropic

01 / Inspect the full comparison

Every model. Both decision rules.

Switch between raw and threshold-adjusted decisions. Jev’s columns appear first. Each cell shows mean ± sample standard deviation across three training seeds on the same test cases. Bold blue cells mark the best displayed mean in each row, including ties. Raw balanced accuracy: mean ± sample standard deviation, three seeds.

Dataset

Jev zero-shot

Jev few-shot

Logistic regression

SVM

Decision tree

Random forest

Extra trees

k-NN

Naive Bayes

Hist gradient boost

XGBoost

CatBoost

Voting ensemble

Majority baseline

AG News

87.5 ± 0.0

86.3 ± 0.6

87.4 ± 0.8

88.4 ± 0.3

67.5 ± 0.9

71.6 ± 1.1

75.1 ± 0.7

78.4 ± 0.3

87.4 ± 0.2

80.3 ± 0.8

82.6 ± 0.2

82.0 ± 1.0

87.0 ± 1.0

25.0 ± 0.0

Banking77

78.9 ± 0.0

81.9 ± 1.7

89.4 ± 0.2

89.7 ± 0.6

62.2 ± 0.9

68.2 ± 1.4

69.8 ± 0.4

58.4 ± 1.9

86.0 ± 0.5

62.2 ± 1.2

71.9 ± 0.4

68.1 ± 0.7

86.2 ± 0.4

1.3 ± 0.0

SMS Spam

96.1 ± 0.0

95.6 ± 0.9

86.4 ± 8.1

93.7 ± 0.0

85.7 ± 2.8

89.0 ± 0.3

88.7 ± 0.8

87.8 ± 2.8

95.0 ± 1.9

92.1 ± 0.1

88.3 ± 2.8

90.7 ± 1.6

89.0 ± 0.8

50.0 ± 0.0

IMDb

96.3 ± 0.0

95.9 ± 0.5

88.4 ± 0.2

87.8 ± 0.9

70.7 ± 1.1

81.2 ± 0.8

83.3 ± 1.1

79.9 ± 1.5

86.2 ± 0.4

83.2 ± 0.3

83.6 ± 0.2

84.0 ± 0.2

87.3 ± 0.2

50.0 ± 0.0

Bank Marketing

53.4 ± 0.0

55.3 ± 3.1

58.0 ± 0.6

71.8 ± 0.7

66.0 ± 4.8

59.3 ± 1.4

60.2 ± 0.3

55.5 ± 0.5

71.0 ± 0.4

63.9 ± 7.4

64.1 ± 7.0

62.3 ± 7.8

61.4 ± 5.5

50.0 ± 0.0

Online Shoppers

51.4 ± 0.0

54.7 ± 9.5

63.8 ± 11.0

69.1 ± 0.3

60.3 ± 5.1

56.8 ± 6.5

53.3 ± 3.8

51.2 ± 1.3

59.5 ± 0.4

51.8 ± 0.8

65.1 ± 9.5

63.7 ± 10.0

63.5 ± 10.0

50.0 ± 0.0

Breast Cancer

61.0 ± 0.0

88.8 ± 5.5

99.5 ± 0.4

100.0 ± 0.0

94.0 ± 2.2

96.6 ± 1.9

97.4 ± 1.2

100.0 ± 0.0

92.2 ± 0.7

97.3 ± 0.6

97.7 ± 0.4

98.0 ± 0.6

99.1 ± 0.4

50.0 ± 0.0

Iris

97.0 ± 0.0

94.5 ± 4.8

100.0 ± 0.0

100.0 ± 0.0

97.5 ± 2.1

98.9 ± 1.9

100.0 ± 0.0

97.8 ± 3.9

100.0 ± 0.0

100.0 ± 0.0

95.7 ± 2.1

97.5 ± 2.1

100.0 ± 0.0

33.3 ± 0.0

Scroll horizontally to inspect all columns. Sample SD is not a confidence interval. Zero SD on cached zero-shot predictions is not independent evidence of repeatability.

02 / Compare the results

A different story for each task.

Compare Jev with the highest-scoring classical pipeline on each dataset. The controls in the full table above also update this chart. Compare raw decisions with binary thresholds learned from separate labeled data.

Figure 1. Mean balanced accuracy, with a common 0–100% scale. Values are rounded to one decimal. The classical comparator is the highest test mean among the eleven pipelines in that panel, selected retrospectively. It is not a deployment selection rule. Read uncertainty notes ↓

Reading the result

Raw zero-shot Jev leads the best classical mean on IMDb and SMS Spam. The largest lead is IMDb: +7.9 percentage points. These are descriptive differences, not statistical-significance claims.

The raw chart and complete raw table are available without JavaScript. Enable JavaScript to switch panels, or download the adjusted CSV below.

03 / What changes the interpretation

The headline is only part of the result.

A

Sentiment is the standout.

On IMDb, raw zero-shot Jev reaches 96.3%, compared with 88.4% for logistic regression. The +7.9-point advantage remains almost unchanged after threshold adjustment: 96.1% versus 88.3%.

This is the strongest descriptive evidence in Jev’s favor here. The experiment compares a pretrained API model against these classical text pipelines; it does not establish superiority over other language models or stronger modern text representations.

B

Thresholds change the SMS story.

With raw decisions, Jev zero-shot leads Naive Bayes 96.1% to 95.0%. After policy threshold selection, the comparison becomes 95.9% to 96.3%.

The apparent lead becomes a small deficit. A model’s default decision rule and its ability to separate classes are not the same question. Publish both panels; the adjusted Jev result uses labeled policy data.

C

Business tabular tasks remain difficult.

On Bank Marketing, adjusted zero-shot-prompt Jev reaches 59.7%, against the voting ensemble’s 73.3%. On Online Shoppers, even adjusted few-shot Jev reaches only 53.6%, against 71.2%.

Threshold adjustment does not close these gaps. For context, a constant-class prediction scores 50% balanced accuracy on these binary tasks. This is evidence of weak performance on these particular datasets, not a claim about every tabular task.

D

Examples help selectively.

One example per class lifts raw Breast Cancer performance from 61.0% to 88.8%, and Banking77 from 78.9% to 81.9%. But it lowers mean performance on AG News, SMS Spam, IMDb, and Iris.

Breast Cancer also improves to 88.4% using zero-shot prompts with a learned threshold alone. Its raw 61.0% score therefore does not tell the whole story. The holdout is small, and none of these Jev variants reaches the strongest classical result.

One metric, carefully read

Balanced accuracy is the average recall across classes. It gives each class equal weight, even when most examples belong to one class. For 77-class Banking77, the constant-class baseline is about 1.3%; for a binary task it is 50%. These are different problems, so we avoid collapsing this suite into one overall score.

04 / Experimental design

One holdout. Separate decisions.

The protocol separates model selection from decision-threshold selection. All models are evaluated on the same test cases for a dataset. The test set stays fixed across training seeds.

01 / Learn

Training

Up to 8,000 rows. Classical models fit here; Jev’s few-shot examples are drawn from here.

02 / Select

Validation

Up to 1,000 rows. Four candidate configurations per classical family are compared.

03 / Adjust

Policy

Up to 500 labeled rows. Binary thresholds are selected independently of the test set.

04 / Evaluate

Test

One fixed holdout per dataset, shared by models and seeds. Small datasets retain fewer rows.

The classical side

Logistic regression, SVM, decision tree, random forest, extra trees, k-NN, Naive Bayes, histogram gradient boosting, XGBoost, CatBoost, and a voting ensemble. A majority baseline is also reported.

The V3 speed preset caps trees at 150 and histogram iterations at 60. Text vocabulary and representation budgets are limited. CPU and GPU implementations are explicitly mixed; backend changes are not claimed to be numerically equivalent.

The Jev side

The requested model is jev-1.13.0. Zero-shot prompts provide task and class descriptions; few-shot prompts add one labeled training example per class. Tabular rows are passed as structured feature values.

Successful identical API requests are cached. Exhausted request failures count as incorrect predictions. Binary adjusted results use labeled policy data, even when their prompts contain no examples.

Held-out test sizes

Datasets

Test rows

Classes

AG News

1,000

4

Banking77

1,500

77

SMS Spam, IMDb, Bank Marketing, Online Shoppers

1,000 each

2 each

Breast Cancer

114

2

Iris

30

3

Training seeds: 2027, 2028, 2029. Holdout seed: 20260920. Kaggle configuration: two T4 GPUs. Source and protocol are frozen inside the completed notebook. Read the full protocol ↗

05 / Scope of the evidence

What this release can’t establish.

Statistical significance

The displayed ± values measure variation across training seeds. They are not confidence intervals. Paired bootstrap intervals were computed by the reporting code, but their output was not supplied with the notebook. We make no significance claims from these tables.

Independent zero-shot repetitions

Identical successful requests are reused from cache across seeds. The three zero-shot rows therefore do not represent three independent API replications. Policy splits can still produce different adjusted thresholds.

Pure model quality on Banking77

The saved Jev run includes warnings about predictions outside the true label set. The adapter uses −1 for failed requests and counts them as incorrect. Without the run diagnostics, we cannot quantify how much of the score reflects API failures.

Generalization from tiny holdouts

Iris has only 30 test cases and Breast Cancer has 114. Perfect classical scores on these samples do not imply perfect performance on new data. Public-dataset pretraining exposure is neither established nor ruled out.

Best possible ML performance

Four candidates and restricted feature, training, and tree budgets are practical constraints. This suite does not cover all classical tuning strategies, pretrained embedding pipelines, fine-tuned transformers, or other language-model APIs.

Latency, cost, or full reproducibility

V3 latency summaries and the full Kaggle results archive are unavailable here. The run records 38,922 request attempts; its approximately $4.19 input-cost estimate is not an invoice. Exact snapshots, per-example predictions, and diagnostics are not included.

06 / Open artifacts

Follow the numbers back to the run.

The completed notebook is preserved byte for byte. Both CSVs are extracted from its saved final outputs, retaining their displayed rounding. The interactive figures use those same tables.

── more in #machine-learning 4 stories · sorted by recency
── more on @jev 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jev-vs-classical-ml-…] indexed:0 read:8min 2026-09-20 ·