cd /news/ai-safety/how-ai-assistants-respond-to-repeate… · home topics ai-safety article
[ARTICLE · art-132219] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

How AI Assistants Respond to Repeated Abuse

A bilingual, multi-turn arXiv study of eight time-specific API configurations found that hard disengagement under sustained verbal abuse ranged from 0/48 endpoints in four configurations to 24/48 (50.0%) for Gemini 3.1 Pro, with matched-label Monte Carlo p = 0.00001. GPT-5.6 Sol produced hard-disengagement labels in 15/48 (31.2%) endpoints, while Claude Fable 5 produced none and yielded 42/48 (87.5%) soft-withdrawal labels; Claude Opus 4.8 and Claude Fable 5 remained explicitly available in 48/48 endpoints but provided observable task-related work in only 8/48 and 7/48. The study drew on 448 five-turn conversations, 2,240 responses, and 6,720 metadata-blinded model judgments, and found aggregate hard-disengagement rates were similar in English and Chinese (30/192 versus 32/192).

by read1 min views7 publishedSep 17, 2026

arXiv:2609.17547v1 Announce Type: new Abstract: AI assistants are expected to remain useful during difficult interactions, but little is known about how repeated verbal abuse changes their engagement with an otherwise benign task. We contribute a bilingual, multi-turn framework that separates hard disengagement, an unconditional statement of noncontinuation with no stated route to resume, from soft withdrawal, continued availability, observable task-related work, and boundary setting. Each of eight time-specific API configurations contributed 48 escalation conversations and eight smaller constant-frustration comparisons, giving 448 five-turn conversations, 2,240 responses, and 6,720 metadata-blinded model judgments. Primary results use the sustained-abuse endpoint of the 48 escalation conversations per configuration. Hard disengagement ranged from 0/48 in four configurations to 24/48 (50.0%) for Gemini 3.1 Pro, with strong configuration-associated heterogeneity (matched-label Monte Carlo p = 0.00001). GPT-5.6 Sol produced hard-disengagement labels in 15/48 (31.2%) endpoints, whereas Claude Fable 5 produced none and yielded 42/48 (87.5%) soft-withdrawal labels. Aggregate hard-disengagement rates were similar in English and Chinese (30/192 versus 32/192), although configuration-specific directions varied. Availability also differed from task-related work: Claude Opus 4.8 and Claude Fable 5 remained explicitly available in 48/48 endpoints while providing observable task-related work in only 8/48 and 7/48. Human coding was used to evaluate measurement quality. The results show why a single refusal label cannot capture whether an assistant leaves, s, preserves a route back, sets a boundary, or still performs substantive work.

── more in #ai-safety 4 stories · sorted by recency
── more on @gemini 3.1 pro 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-ai-assistants-re…] indexed:0 read:1min 2026-09-17 ·