StereoBind: Enabling Dynamic Spatial Sound for Immersive Audio-Visual Generation
StereoBind introduces dynamic spatial correspondence between moving visual sources and stereo audio, enabling immersive experiences for AR/VR and gaming. It uses three mechanisms—Visual Motion Binding, Spatial Track Encoder, and Residual Track RoPE—to coordinate motion and sound location. Experiments using the new StereoWorld-29K dataset demonstrate improved spatial audio alignment over previous models.
StereoBind is a new framework for joint video-audio generation that addresses the need for spatially immersive stereo sound, especially in AR/VR and interactive gaming. Unlike prior models, which often ignored dynamic stereo effects, StereoBind explicitly binds the motion of visual sources to their corresponding stereo audio locations.

The framework uses motion tracks to synchronize visual and audio spatial positioning through three complementary mechanisms:
- Visual Motion Binding: Establishes source-aware audiovisual correspondence.
- Spatial Track Encoder: Captures the absolute positions of moving sources.
- Residual Track RoPE: Models the relative motion of sources over time.
To supervise and evaluate the system, the researchers introduced StereoWorld-29K, a large-scale dataset with paired video, stereo audio, and motion tracks, as well as StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind achieves improved spatial alignment of stereo audio with visual motion, without compromising overall audio-visual quality.
