cd /news/artificial-intelligence/do-all-llms-know-when-they-re-being-… · home topics artificial-intelligence article
[ARTICLE · art-91607] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

A reproducibility study of latent-space safety probes found that lightweight MLP probes trained on final-layer activations of LLaMA-3.1-8B detect harmful prompts with F1 scores within 0.37 percentage points of the original study across benchmarks, and the approach extends to Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B with F1 scores within one point of LLaMA-3.1-8B. The study, posted on arXiv (2608.08029v1), also found that final token latent vectors remained identical across five random seeds for all tested architectures, indicating non-determinism does not affect the probes.

read1 min views1 publishedAug 11, 2026

arXiv:2608.08029v1 Announce Type: new Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at F1 competitive with guard models 1000x larger, using one probe per benchmark. We reproduce this pipeline end-to-end and extend it along two axes the original study leaves open. First, we test whether the result generalizes across other model architecture and scale by training identical probes on activations from models like Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B, using the three benchmarks (WildJailbreak, BeaverTails, AEGIS 2.0). Second, we test how much of the reported performance is affected by non-determinism during inference by repeating extraction under five random seeds and measuring the variance of F1 scores. Our results reproduce the original LLaMA model benchmarks within 0.37 percentage points of the original F1 scores (and within 0.2 points on BeaverTails). We find that the original MLP probe architecture extends to other model families with F1 scores within a point of the values reported for LLaMA-3.1-8B. Our experiments varying seed values reveal an interesting observation: final token latent vectors remained the same for all tested architectures irrespective of the seed values used.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-all-llms-know-whe…] indexed:0 read:1min 2026-08-11 ·