{"slug": "llms-perform-worse-than-random-at-pro-active-investigation", "title": "LLMs perform worse than random at pro-active investigation", "summary": "A paper titled \"SherlockBench, Where Large Language Models Under-Perform Random Heuristics\" (DOI 10.5281/zenodo.16253500) reports that large language models performed worse than random chance on pro-active investigation tasks, according to its author, who posted the findings on the Hugging Face forum. The author, posting as Xylon3, said the work spans more than one paper and asked for feedback from the community.", "body_md": "[Xylon3](https://discuss.huggingface.co/u/Xylon3)\n1\n \nIn this paper, we see LLMs under-performing random chance at pro-active investigation tasks: [SherlockBench, Where Large Language Models Under-Perform Random Heuristics](https://doi.org/10.5281/zenodo.16253500)\n\nAnd yes I am the author. Would appreciate any feedback.\n\n \n \nYoooo good stuff ! Let’s go, open science! I’d love to get in touch. It’s super late or early for me right now and won’t be able to ‘really’ read it.\n\nAnyway I dig the point reg scientific discovery…\n\nI’ve now switched mostly to offensive LLM swarms using 1 “smart”(ass) model to orchestrate 2-3 smaller .\n\nIt always begins with the usual and sort of a traditional way of (red team assignment ; however excluding AI red teaming as of yet ! )\n\n1. first goal is to just do a recon (think container with some entry level CTF like dvwa , now I make my own )\n2. then the smaller model reports findings to the larger and it starts printing a playbook (yaml / jinja if you know Ansible )\n3. keep at it until root and verifiable proof.\n\nIt is highly creative - well yeah it’s hacking  … literally. And it can do it but sometimes it just is super dumb and lazy …? What I hate the most is that they just fuckin’ lie and lie.\n\nWould love to chat if you’re up to it. I;m curious how did you manage to get the best result with `RANDOM` haha, anyway aret work , will give feedback l8r . Also noticed you’ve been working on it for more than just one paper .\n\nCheers !\n\n \n \nWelcome @ Xylon3 and [@iluvuluv](https://discuss.huggingface.co/u/iluvuluv)\n\nI have so much to learn so I am not following but again Welcome to posting!\n\n \n \nIn this article they explained [AI agent red teaming](https://www.botgauge.com/blog/llm-red-teaming-vs-agent-red-teaming) so well", "url": "https://wpnews.pro/news/llms-perform-worse-than-random-at-pro-active-investigation", "canonical_source": "https://discuss.huggingface.co/t/llms-perform-worse-than-random-at-pro-active-investigation/164244#post_4", "published_at": "2026-09-27 14:24:50+00:00", "updated_at": "2026-09-27 14:30:28.086745+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-agents"], "entities": ["Hugging Face", "Xylon3", "SherlockBench", "Zenodo"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/llms-perform-worse-than-random-at-pro-active-investigation", "markdown": "https://wpnews.pro/news/llms-perform-worse-than-random-at-pro-active-investigation.md", "text": "https://wpnews.pro/news/llms-perform-worse-than-random-at-pro-active-investigation.txt", "jsonld": "https://wpnews.pro/news/llms-perform-worse-than-random-at-pro-active-investigation.jsonld"}}