Audio-video foundation models trained at scale implicitly encode a vast repertoire of perceptual and physical knowledge: motion, identity, environmental sound, lip dynamics, light, and material. The practical bottleneck is no longer what such a model can synthesize, but what a user can ask of it. This course presents m...
Naomi ken korem, Matan Ben Yosef· Proceedings of the Special I...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.