{"slug": "stateful-guardrails-for-multi-turn-llm-systems-a-conversational-risk-framework", "title": "Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework", "summary": "Researchers have introduced a Conversational Risk Accumulation (CRA) Framework to detect safety failures in multi-turn large language model (LLM) systems that arise from benign turns composing into harm over a dialogue. The framework tracks semantic drift, sensitivity-weighted information accumulation, and compliance-gradient signals, and includes CRA-Net DA, a learned trajectory model. To benchmark CRA, the team released CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families), CRA-Bench v0.2 (LLM-paraphrased variants), and an extended 5-family set (2,000 sessions).", "body_md": "arXiv:2607.19361v1 Announce Type: new\nAbstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk Accumulation (CRA): gradual intent drift, fragmented assembly of prohibited instructions, and sensitivity build-up from repeated disclosures. We propose a session-layer CRA Framework that tracks three trajectory signals: semantic drift from a session anchor, a sensitivity-weighted information accumulation graph over extracted entities, and a compliance-gradient signal capturing increasing willingness to comply. For scoring, we provide (i) an unsupervised convex fusion for attribution and ablations, and (ii) CRA-Net DA, a compact learned trajectory model trained with family-adversarial objectives to reduce length and topic-coverage confounds. To benchmark CRA, we release CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families with topic-matched benign twins), CRA-Bench v0.2 (LLM-paraphrased variants to reduce template artifacts), and an extended 5-family set (2,000 sessions adding persona priming and context stuffing). We introduce a trajectory-native evaluation protocol with session-level splits, mixed-set threshold calibration, Trajectory AUROC, turns-to-detection, calibrated false-positive metrics, bootstrap confidence intervals, leave-one-family-out diagnostic stress tests, and synthetic-to-human transfer checks. Claims focus on within-distribution session scoring on CRA-Bench and human-transfer subsets.", "url": "https://wpnews.pro/news/stateful-guardrails-for-multi-turn-llm-systems-a-conversational-risk-framework", "canonical_source": "https://www.machinebrief.com/news/stateful-guardrails-for-multi-turn-llm-systems-a-conversatio-i8q3", "published_at": "2026-07-23 04:00:00+00:00", "updated_at": "2026-07-23 04:03:55.415428+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv", "CRA-Net DA", "CRA-Bench v0.1", "CRA-Bench v0.2"], "alternates": {"html": "https://wpnews.pro/news/stateful-guardrails-for-multi-turn-llm-systems-a-conversational-risk-framework", "markdown": "https://wpnews.pro/news/stateful-guardrails-for-multi-turn-llm-systems-a-conversational-risk-framework.md", "text": "https://wpnews.pro/news/stateful-guardrails-for-multi-turn-llm-systems-a-conversational-risk-framework.txt", "jsonld": "https://wpnews.pro/news/stateful-guardrails-for-multi-turn-llm-systems-a-conversational-risk-framework.jsonld"}}