cd /news/artificial-intelligence/stateful-guardrails-for-multi-turn-l… · home topics artificial-intelligence article
[ARTICLE · art-69566] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework

Researchers have introduced a Conversational Risk Accumulation (CRA) Framework to detect safety failures in multi-turn large language model (LLM) systems that arise from benign turns composing into harm over a dialogue. The framework tracks semantic drift, sensitivity-weighted information accumulation, and compliance-gradient signals, and includes CRA-Net DA, a learned trajectory model. To benchmark CRA, the team released CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families), CRA-Bench v0.2 (LLM-paraphrased variants), and an extended 5-family set (2,000 sessions).

read1 min views1 publishedJul 23, 2026

arXiv:2607.19361v1 Announce Type: new Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk Accumulation (CRA): gradual intent drift, fragmented assembly of prohibited instructions, and sensitivity build-up from repeated disclosures. We propose a session-layer CRA Framework that tracks three trajectory signals: semantic drift from a session anchor, a sensitivity-weighted information accumulation graph over extracted entities, and a compliance-gradient signal capturing increasing willingness to comply. For scoring, we provide (i) an unsupervised convex fusion for attribution and ablations, and (ii) CRA-Net DA, a compact learned trajectory model trained with family-adversarial objectives to reduce length and topic-coverage confounds. To benchmark CRA, we release CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families with topic-matched benign twins), CRA-Bench v0.2 (LLM-paraphrased variants to reduce template artifacts), and an extended 5-family set (2,000 sessions adding persona priming and context stuffing). We introduce a trajectory-native evaluation protocol with session-level splits, mixed-set threshold calibration, Trajectory AUROC, turns-to-detection, calibrated false-positive metrics, bootstrap confidence intervals, leave-one-family-out diagnostic stress tests, and synthetic-to-human transfer checks. Claims focus on within-distribution session scoring on CRA-Bench and human-transfer subsets.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stateful-guardrails-…] indexed:0 read:1min 2026-07-23 ·