{"slug": "moba-vl-event-localized-multi-turn-reinforcement-learning-for-real-time-moba", "title": "MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary", "summary": "Researchers introduced MOBA-VL, a 9B-parameter vision-language model trained with event-localized multi-turn reinforcement learning on game telemetry to deliver real-time commentary for Multiplayer Online Battle Arena esports. On the held-out MOBACast-Bench, MOBA-VL scored 63.25 Overall on full matches versus 55.12 for StreamingVLM and 63.45 on clips versus 56.22 for DeepSeek-V4.1-Flash, while event-localized credit raised event recall from 34.5 to 42.1 over supervised fine-tuning. The team also released MOBACast, 860 professional matches totaling about 460 hours across three MOBA games with word-level timestamped commentary, with code and data to be released.", "body_md": "arXiv:2609.38428v1 Announce Type: new \nAbstract: Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at https://moba-vl.github.io.", "url": "https://wpnews.pro/news/moba-vl-event-localized-multi-turn-reinforcement-learning-for-real-time-moba", "canonical_source": "https://arxiv.org/abs/2609.38428", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:20:12.084459+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "natural-language-processing", "computer-vision"], "entities": ["MOBA-VL", "MOBACast", "MOBACast-Bench", "StreamingVLM", "DeepSeek-V4.1-Flash", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/moba-vl-event-localized-multi-turn-reinforcement-learning-for-real-time-moba", "markdown": "https://wpnews.pro/news/moba-vl-event-localized-multi-turn-reinforcement-learning-for-real-time-moba.md", "text": "https://wpnews.pro/news/moba-vl-event-localized-multi-turn-reinforcement-learning-for-real-time-moba.txt", "jsonld": "https://wpnews.pro/news/moba-vl-event-localized-multi-turn-reinforcement-learning-for-real-time-moba.jsonld"}}