LLMs perform worse than random at pro-active investigation A paper titled "SherlockBench, Where Large Language Models Under-Perform Random Heuristics" (DOI 10.5281/zenodo.16253500) reports that large language models performed worse than random chance on pro-active investigation tasks, according to its author, who posted the findings on the Hugging Face forum. The author, posting as Xylon3, said the work spans more than one paper and asked for feedback from the community. Xylon3 https://discuss.huggingface.co/u/Xylon3 1 In this paper, we see LLMs under-performing random chance at pro-active investigation tasks: SherlockBench, Where Large Language Models Under-Perform Random Heuristics https://doi.org/10.5281/zenodo.16253500 And yes I am the author. Would appreciate any feedback. Yoooo good stuff Let’s go, open science I’d love to get in touch. It’s super late or early for me right now and won’t be able to ‘really’ read it. Anyway I dig the point reg scientific discovery… I’ve now switched mostly to offensive LLM swarms using 1 “smart” ass model to orchestrate 2-3 smaller . It always begins with the usual and sort of a traditional way of red team assignment ; however excluding AI red teaming as of yet 1. first goal is to just do a recon think container with some entry level CTF like dvwa , now I make my own 2. then the smaller model reports findings to the larger and it starts printing a playbook yaml / jinja if you know Ansible 3. keep at it until root and verifiable proof. It is highly creative - well yeah it’s hacking … literally. And it can do it but sometimes it just is super dumb and lazy …? What I hate the most is that they just fuckin’ lie and lie. Would love to chat if you’re up to it. I;m curious how did you manage to get the best result with RANDOM haha, anyway aret work , will give feedback l8r . Also noticed you’ve been working on it for more than just one paper . Cheers Welcome @ Xylon3 and @iluvuluv https://discuss.huggingface.co/u/iluvuluv I have so much to learn so I am not following but again Welcome to posting In this article they explained AI agent red teaming https://www.botgauge.com/blog/llm-red-teaming-vs-agent-red-teaming so well