{"slug": "enforcing-llm-safety-through-dmd-based-classification-of-prompt-response", "title": "Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics", "summary": "Researchers introduced a black-box method that classifies unsafe LLM outputs by fitting Koopman-based predictive models to prompt-response embedding dynamics, using a differential residual score to compare safe and unsafe regimes. Evaluated across three safety benchmarks with three embedding models, the method improved detection of interaction-dependent violations, especially with causal decoders like Llama-3, while response-only violations benefited from dense semantic embeddings. The work, posted on arXiv (2608.19579v1), suggests using dynamical systems to analyze AI systems.", "body_md": "arXiv:2608.19579v1 Announce Type: new\nAbstract: Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and responses into high-dimensional embedding spaces and fitting separate Koopman-based predictive models for safe and unsafe regimes, we classify new outputs using a new differential residual score that compares prediction errors of the safe and unsafe regimes. A key contribution is the incorporation of the prompt and response embedding dynamics, yielding fitted Koopman operators that capture crucial interaction patterns. We evaluate our black-box method across three safety benchmarks using three embedding models. Our results show that incorporating prompt embeddings yields consistent improvements, particularly for interaction-dependent violations when paired with causal decoders (e.g., in Llama-3), while response-only violations benefit more from dense semantic embedding representations. These findings opens the door for using dynamical systems to analyze AI systems rather than the dominant paradigm of using AI to model dynamical systems.", "url": "https://wpnews.pro/news/enforcing-llm-safety-through-dmd-based-classification-of-prompt-response", "canonical_source": "https://arxiv.org/abs/2608.19579", "published_at": "2026-08-21 04:00:00+00:00", "updated_at": "2026-08-21 04:12:41.128704+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv", "Llama-3", "Koopman"], "alternates": {"html": "https://wpnews.pro/news/enforcing-llm-safety-through-dmd-based-classification-of-prompt-response", "markdown": "https://wpnews.pro/news/enforcing-llm-safety-through-dmd-based-classification-of-prompt-response.md", "text": "https://wpnews.pro/news/enforcing-llm-safety-through-dmd-based-classification-of-prompt-response.txt", "jsonld": "https://wpnews.pro/news/enforcing-llm-safety-through-dmd-based-classification-of-prompt-response.jsonld"}}