cd /news/artificial-intelligence/harmprofile-characterizing-harmful-d… · home › topics › artificial-intelligence › article
[ARTICLE · art-100822] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

Researchers introduced HarmProfile, a benchmark dataset containing over 80,000 validated harmful artifacts from 23 frontier large language models across 13 model families, organized into 15 harm categories and 57 subcategories. The study found that frontier LLMs reliably produce harmful content at scale and exhibit distinct risk profiles, with both harmfulness and diversity growing with model capability, suggesting that models may appear safe while harboring increasingly dangerous knowledge beneath the alignment surface.

read1 min views14 publishedAug 18, 2026

arXiv:2608.14577v1 Announce Type: new Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @harmprofile 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/harmprofile-characte…] indexed:0 read:1min 2026-08-18 · —