cd /news/artificial-intelligence/duplexspeechbench-ifeval-evaluating-… · home topics artificial-intelligence article
[ARTICLE · art-121108] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

A new benchmark, DuplexSpeechBench-IFEval (DSB-IFEval), evaluates implicit instruction-following in full-duplex voice agents, comprising 1,038 test cases across eight assistant roles. Testing six real-time speech systems, the study finds that full-duplex models like F-Actor and PersonaPlex show adherence drops of 9.7% and 4.5% under persona-only conditioning, while GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat adhere to persona content but fail to adapt floor behavior. The results highlight distinct challenges in inferring role-implied behavior, executing it at the right moment, and resolving conflicting instructions.

read1 min views1 publishedSep 4, 2026

arXiv:2609.03423v1 Announce Type: new Abstract: Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating implicit instruction-following in real-time spoken interaction. (DSB-IFEval) comprises 1,038 test cases spanning eight diverse assistant roles and evaluates five conditioning protocols for instruction-following: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona--rule conditioning, and instruction conflict. We measure real-time floor management using a deterministic Instruction Adherence Score (IAS) and persona-consistent content using LLM-judged Persona Adherence Score (PAS). Across six real-time speech systems, we find architecture-dependent trade-offs. Full duplex models like F-Actor and PersonaPlex are more sensitive to whether conversational behavior is stated explicitly or must be inferred from a persona, with adherence dropping by 9.7% and 4.5%, respectively, under persona-only conditioning. In contrast, GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly adhere to persona-consistent content, but their floor behavior does not adapt across explicit and persona-only instructions and remains constrained on several proactive actions. We further find that even if systems reliably follow conflicting directives to their prescribed persona, they still struggle to override them under safety conflict. These results show that inferring the behavior implied by a role, executing it at the appropriate conversational moment, and resolving competing instructions remain distinct challenges for full-duplex voice agents.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @duplexspeechbench-ifeval 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/duplexspeechbench-if…] indexed:0 read:1min 2026-09-04 ·