{"slug": "beyond-surface-style-aligning-multi-turn-user-simulators-with-behavioral", "title": "Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency", "summary": "Researchers introduced TRACER, a multi-turn user simulator that models evolving user intent and aligns simulated behavior with real interaction trajectories, with TRACER-7B surpassing the strongest baseline by 11.4 conversion F1 on real customer-service sessions organized into reference cohorts. TRACER is trained in two stages — supervised fine-tuning on real user dialogues followed by multi-turn reinforcement learning that combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation — and also achieves the lowest group-level conversion-rate error and semantic trajectory distance while generalizing to out-of-distribution scenarios. The team additionally released the Dynamic Marketing Benchmark, which evaluates LLM persuasion effectiveness and response quality through simulated interactions and finds that higher response quality does not necessarily correspond to higher conversion rates.", "body_md": "arXiv:2609.28690v1 Announce Type: new \nAbstract: Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.", "url": "https://wpnews.pro/news/beyond-surface-style-aligning-multi-turn-user-simulators-with-behavioral", "canonical_source": "https://arxiv.org/abs/2609.28690", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 04:29:27.084445+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "machine-learning", "ai-agents"], "entities": ["TRACER", "TRACER-7B", "Dynamic Marketing Benchmark", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/beyond-surface-style-aligning-multi-turn-user-simulators-with-behavioral", "markdown": "https://wpnews.pro/news/beyond-surface-style-aligning-multi-turn-user-simulators-with-behavioral.md", "text": "https://wpnews.pro/news/beyond-surface-style-aligning-multi-turn-user-simulators-with-behavioral.txt", "jsonld": "https://wpnews.pro/news/beyond-surface-style-aligning-multi-turn-user-simulators-with-behavioral.jsonld"}}