Skip to content
Book Open access

Orchestrating Multimodal Control for Unified Audio-Visual Synthesis

Jul 2026 · Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Courses · 0 citations · 6 references

Abstract

Audio-video foundation models trained at scale implicitly encode a vast repertoire of perceptual and physical knowledge: motion, identity, environmental sound, lip dynamics, light, and material. The practical bottleneck is no longer what such a model can synthesize, but what a user can ask of it. This course presents multimodal control as the means to unlock that latent capability on demand. Once a creative idea is framed as a task, a pairing of audio and visual conditioning signals with the desired output, a short training run on a single machine teaches the model to expose the corresponding behavior. Because the pretraining covers so much ground, the space of attainable controls is in practice open-ended: almost any capability the model already “knows” can be made addressable. We build the course on LTX-2 [HaCohen et al. 2026], an asymmetric dual-stream audio-video foundation model. Heterogeneous conditioning signals are injected into the appropriate stream and learned with the AVControl framework [Ben-Yosef et al. 2026]. We illustrate the framework with several recent applications, all obtained from the same training recipe rather than from bespoke, task-specific architectures: cross-lingual video dubbing, in which the model jointly generates translated speech and the corresponding lip motion while preserving speaker identity, non-speech audio, and visual context [Chen et al. 2026]; SDR-to-HDR video generation via latent alignment with a logarithmic perceptual encoding [Ken Korem et al. 2026]; and a range of finer-grained controls, active-speaker selection (who is speaking when several people share a frame), background replacement, depth- and pose-conditioned generation, and audio transformations, drawn from the AVControl paper itself [Ben-Yosef et al. 2026]. The result is a blueprint for moving from prompting stochastic generators to teaching pretrained models the specific capabilities that a production workflow requires.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.