04:00
2026-08-21
arxiv.org
artificial-intelligence
Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics
Researchers introduced a black-box method that classifies unsafe LLM outputs by fitting Koopman-based predictive models to prompt-response embedding dynamics, using a differential residual score to co…