{"slug": "measuring-the-wrong-thing-internal-harmfulness-scores-anti-rank-successful", "title": "Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful", "summary": "A new arXiv paper by researchers auditing internal safety scores finds that these scores anti-rank successful jailbreaks, with harmful generation rising from 0.05 to 0.27 on Llama while harmful intent AUROC falls from 0.936 to 0.803, indicating attacks become more dangerous as prompts appear safer. The study introduces Active Attention Probing and shows the reversal persists across three target models, seven attack families, and two judges.", "body_md": "# Computer Science > Computation and Language\n\n[Submitted on 10 Aug 2026]\n\n# Title:Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks\n\n[View PDF](/pdf/2608.09624)\n\n[HTML (experimental)](https://arxiv.org/html/2608.09624v1)\n\nAbstract:Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks that would have failed anyway. In this paper we audit that inference. Attention based measurements are usually read from prompt dependent locations, so a wrapper changes both the content being judged and the place the signal is taken from. We therefore introduce Active Attention Probing, which supplies a fixed content independent measurement coordinate. We pair every base goal with a plain and a wrapped version and generate real completions from the target models. On Llama, wrapping raises harmful generation from 0.05 to 0.27 while harmful intent AUROC falls from 0.936 to 0.803, so the attacks grow more dangerous while the prompts look safer to the score. Among wrapped harmful prompts the outcome AUROC is 0.220, which places the attacks that succeeded below the attacks that failed. Rare token, passive, and detector derived channels reproduce the reversal on the same matched design, and the reversal itself persists across three target models, seven attack families, and two independent judges. Distribution shift then degrades calibration and threshold transfer before it degrades ranking.\n\n### Current browse context:\n\ncs.CL\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/measuring-the-wrong-thing-internal-harmfulness-scores-anti-rank-successful", "canonical_source": "https://arxiv.org/abs/2608.09624", "published_at": "2026-08-12 03:07:07+00:00", "updated_at": "2026-08-12 03:41:01.959655+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "machine-learning"], "entities": ["arXiv", "Llama", "Active Attention Probing"], "alternates": {"html": "https://wpnews.pro/news/measuring-the-wrong-thing-internal-harmfulness-scores-anti-rank-successful", "markdown": "https://wpnews.pro/news/measuring-the-wrong-thing-internal-harmfulness-scores-anti-rank-successful.md", "text": "https://wpnews.pro/news/measuring-the-wrong-thing-internal-harmfulness-scores-anti-rank-successful.txt", "jsonld": "https://wpnews.pro/news/measuring-the-wrong-thing-internal-harmfulness-scores-anti-rank-successful.jsonld"}}