cd /news/large-language-models/llms-perform-worse-than-random-at-pr… · home › topics › large-language-models › article
[ARTICLE · art-140489] src=discuss.huggingface.co ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LLMs perform worse than random at pro-active investigation

A paper titled "SherlockBench, Where Large Language Models Under-Perform Random Heuristics" (DOI 10.5281/zenodo.16253500) reports that large language models performed worse than random chance on pro-active investigation tasks, according to its author, who posted the findings on the Hugging Face forum. The author, posting as Xylon3, said the work spans more than one paper and asked for feedback from the community.

read1 min views8 publishedSep 27, 2026
LLMs perform worse than random at pro-active investigation
Image: Discuss (auto-discovered)

Xylon3 1

In this paper, we see LLMs under-performing random chance at pro-active investigation tasks: SherlockBench, Where Large Language Models Under-Perform Random Heuristics

And yes I am the author. Would appreciate any feedback.

Yoooo good stuff ! Let’s go, open science! I’d love to get in touch. It’s super late or early for me right now and won’t be able to ‘really’ read it.

Anyway I dig the point reg scientific discovery…

I’ve now switched mostly to offensive LLM swarms using 1 “smart”(ass) model to orchestrate 2-3 smaller .

It always begins with the usual and sort of a traditional way of (red team assignment ; however excluding AI red teaming as of yet ! )

  1. first goal is to just do a recon (think container with some entry level CTF like dvwa , now I make my own )
  2. then the smaller model reports findings to the larger and it starts printing a playbook (yaml / jinja if you know Ansible )
  3. keep at it until root and verifiable proof.

It is highly creative - well yeah it’s hacking … literally. And it can do it but sometimes it just is super dumb and lazy …? What I hate the most is that they just fuckin’ lie and lie.

Would love to chat if you’re up to it. I;m curious how did you manage to get the best result with RANDOM haha, anyway aret work , will give feedback l8r . Also noticed you’ve been working on it for more than just one paper .

Cheers !

Welcome @ Xylon3 and @iluvuluv I have so much to learn so I am not following but again Welcome to posting!

In this article they explained AI agent red teaming so well

── more in #large-language-models 4 stories · sorted by recency
── more on @hugging face 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llms-perform-worse-t…] indexed:0 read:1min 2026-09-27 · —