MOBA-VL: Event-Localized Multi-Turn RL Delivers Improved Real-Time MOBA Commentary
MOBA-VL is a new 9B-parameter vision-language model for real-time commentary in MOBA esports. It uses event-localized multi-turn reinforcement learning with supervision from game telemetry, improving accuracy on key moments. MOBA-VL outperforms prior streaming VLMs on new benchmarks, and code, data, and demos are available.
MOBA-VL introduces a 9B-parameter vision-language model designed for real-time commentary in Multiplayer Online Battle Arena (MOBA) esports. Unlike existing models, which often miss critical gameplay events despite sounding fluent, MOBA-VL leverages precise game telemetry. This data, which includes event timings, is used as supervision to train the model using event-localized multi-turn reinforcement learning.

Key Improvements
- Trained with direct supervision from game event timings, improving event description accuracy.
- Uses reinforcement learning where rewards are localized to commentary turns that describe key game events.
- Introduces the MOBACast dataset (860 matches, 460 hours) and the MOBACast-Bench benchmark for evaluation.
- Delivers higher Overall score than previous state-of-the-art models on full matches (63.25 vs. 55.12 for StreamingVLM) and on clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash).
- Raises event recall from 34.5 to 42.1 compared to traditional supervised fine-tuning.
What Developers Should Do
- Evaluate the demo and explore how event-localized RL could boost output accuracy for any system requiring real-time event tracking and narration.
- Review the MOBACast dataset and MOBACast-Bench benchmark for benchmarking or transfer learning tasks related to esports commentary.
