cd /news/artificial-intelligence/enforcing-llm-safety-through-dmd-bas… · home topics artificial-intelligence article
[ARTICLE · art-105420] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

Researchers introduced a black-box method that classifies unsafe LLM outputs by fitting Koopman-based predictive models to prompt-response embedding dynamics, using a differential residual score to compare safe and unsafe regimes. Evaluated across three safety benchmarks with three embedding models, the method improved detection of interaction-dependent violations, especially with causal decoders like Llama-3, while response-only violations benefited from dense semantic embeddings. The work, posted on arXiv (2608.19579v1), suggests using dynamical systems to analyze AI systems.

read1 min views5 publishedAug 21, 2026

arXiv:2608.19579v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and responses into high-dimensional embedding spaces and fitting separate Koopman-based predictive models for safe and unsafe regimes, we classify new outputs using a new differential residual score that compares prediction errors of the safe and unsafe regimes. A key contribution is the incorporation of the prompt and response embedding dynamics, yielding fitted Koopman operators that capture crucial interaction patterns. We evaluate our black-box method across three safety benchmarks using three embedding models. Our results show that incorporating prompt embeddings yields consistent improvements, particularly for interaction-dependent violations when paired with causal decoders (e.g., in Llama-3), while response-only violations benefit more from dense semantic embedding representations. These findings opens the door for using dynamical systems to analyze AI systems rather than the dominant paradigm of using AI to model dynamical systems.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/enforcing-llm-safety…] indexed:0 read:1min 2026-08-21 ·