# Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents

> Source: <https://huggingface.co/papers/2609.13117>
> Published: 2026-09-14 14:10:02+00:00

[Collection Weekly signals in speech-to-speech and full-duplex voice AI. Latest: 2026-W35, Aug 17 - Aug 23, 2026. Archive: fullduplex.ai/signals • 24 items • Updated](/collections/otoearth/fullduplex-signals)  

[Papers](/papers)

# Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents

## Abstract

Duplex Cue evaluates how voice agents adapt to listener contributions during ongoing turns, revealing that current models adapt far less often than humans.

[thinkingmachines/Inkling-Small](/thinkingmachines/Inkling-Small)

[Full-duplex](/papers?q=Full-duplex) evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this [in-turn adaptation](/papers?q=in-turn%20adaptation) in [full-duplex](/papers?q=full-duplex) voice agents. Duplex Cue separates listener intent ([backchannel](/papers?q=backchannel), [collaboration](/papers?q=collaboration), or [interruption](/papers?q=interruption)) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with [PersonaPlex](/papers?q=PersonaPlex) continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2\% of cases, compared with 34.8\% for [PersonaPlex](/papers?q=PersonaPlex). The model otherwise continues unchanged (42.4\%) or yields (22.7\%). These findings show why evaluating natural voice interaction requires measuring how an agent responds to a listener's contribution as well as whether it keeps speaking.

Get this paper in your agent:

`hf papers read 2609.13117` ## Don't have the latest CLI?

`curl -LsSf https://hf.co/cli/install.sh | bash` ## Models citing this paper 0

No model linking this paper

## Datasets citing this paper 0

No dataset linking this paper

### Spaces citing this paper 0

No Space linking this paper
