14:24
2026-09-27
discuss.huggingface.co
large-language-models
LLMs perform worse than random at pro-active investigation
A paper titled "SherlockBench, Where Large Language Models Under-Perform Random Heuristics" (DOI 10.5281/zenodo.16253500) reports that large language models performed worse than random chance on pro-a…