cd /news/large-language-models/jev-vs-claude-as-blind-headline-judg… · home › topics › large-language-models › article
[ARTICLE · art-148853] src=alexvidovich.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Jev vs. Claude as blind headline judge: six gates before either touches a trade

Independent researchers Alex Vidovich and Zivana Zerjal published a 11-page working paper on SSRN on October 9, 2026 describing six tests for using a small commercial text model as a headline judge beside a live intraday and overnight trading system on one account. The paper reports the judge read one plain fact correctly against a real label, produced one small effect below its bar, and saw several effects vanish, with its margin over a simple headline count described as borderline; the authors make no claim about either strategy's returns. The six gates require the model to read a fact correctly, beat a simple count of its own headlines, show no memory of famous past events, give the same answer twice, always see the decision-time date, and check each forward verdict rule against random flags.

read2 min views2 publishedOct 10, 2026
Jev vs. Claude as blind headline judge: six gates before either touches a trade
Image: source

October 9, 2026 · 11 pages · Working paper · posted on SSRN

Abstract #

A text model can sit next to a live trading system the way a thermometer sits next to a patient: it reads, it never prescribes. We follow one such model into a lab running an intraday strategy and a second, overnight strategy on one account. It is a small commercial judge, pinned to one version, that answers only typed questions about headlines. We describe six tests for a text signal: read a plain fact correctly against a real label, beat a simple count of its own headlines, show no memory of famous past events, give the same answer twice, always see the decision-time date, and check each forward verdict rule against random flags. They were assembled over three days, most of them after a failure, and we present them as the checklist we would now apply up front. We compare the judge with a general chatbot on blind rows, list what it may touch (stamps written after the close, counted forward) and what it may never touch (any number in the trading system), and report the plumbing failures that corrupted its readings. Its readings so far are small: one fact read well, one small effect below its bar, several that vanished; the fact was read under one setup, and its margin over a simple headline count is borderline. None of the six tests is new on its own; what is new is using them together on a text instrument beside a live system, and reporting every study it was used in, failures included. This is a paper about checking data and the instrument, not about profit: we make no claim about either strategy's returns.

Keywords: large language models, text classification, news analytics, pre-registration, look-ahead bias, research infrastructure, trading systems

JEL classification: G14, G17, C45, C52, C81, C88

BibTeX #

@techreport{vidovich2026text,
  title   = {A Thermometer, Not an Analyst: Lawful Uses of a Text Judge in a Live Trading Lab},
  author  = {Alex Vidovich and Zivana Zerjal},
  year    = {2026},
  month   = {10},
  type    = {Working paper},
  institution = {alexvidovich.com},
  doi     = {10.2139/ssrn.7589261},
  url     = {https://alexvidovich.com/papers/text-judge},
}
── more in #large-language-models 4 stories · sorted by recency
── more on @alex vidovich 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/jev-vs-claude-as-bli…] indexed:0 read:2min 2026-10-10 · —