# LLMs perform worse than random at pro-active investigation

> Source: <https://discuss.huggingface.co/t/llms-perform-worse-than-random-at-pro-active-investigation/164244#post_4>
> Published: 2026-09-27 14:24:50+00:00

[Xylon3](https://discuss.huggingface.co/u/Xylon3)
1
 
In this paper, we see LLMs under-performing random chance at pro-active investigation tasks: [SherlockBench, Where Large Language Models Under-Perform Random Heuristics](https://doi.org/10.5281/zenodo.16253500)

And yes I am the author. Would appreciate any feedback.

 
 
Yoooo good stuff ! Let’s go, open science! I’d love to get in touch. It’s super late or early for me right now and won’t be able to ‘really’ read it.

Anyway I dig the point reg scientific discovery…

I’ve now switched mostly to offensive LLM swarms using 1 “smart”(ass) model to orchestrate 2-3 smaller .

It always begins with the usual and sort of a traditional way of (red team assignment ; however excluding AI red teaming as of yet ! )

1. first goal is to just do a recon (think container with some entry level CTF like dvwa , now I make my own )
2. then the smaller model reports findings to the larger and it starts printing a playbook (yaml / jinja if you know Ansible )
3. keep at it until root and verifiable proof.

It is highly creative - well yeah it’s hacking  … literally. And it can do it but sometimes it just is super dumb and lazy …? What I hate the most is that they just fuckin’ lie and lie.

Would love to chat if you’re up to it. I;m curious how did you manage to get the best result with `RANDOM` haha, anyway aret work , will give feedback l8r . Also noticed you’ve been working on it for more than just one paper .

Cheers !

 
 
Welcome @ Xylon3 and [@iluvuluv](https://discuss.huggingface.co/u/iluvuluv)

I have so much to learn so I am not following but again Welcome to posting!

 
 
In this article they explained [AI agent red teaming](https://www.botgauge.com/blog/llm-red-teaming-vs-agent-red-teaming) so well
