Building 3D Representations and Generating Motions From a Single Image via Video-Generation
This work proposes a framework known as Video-Generation Environment Representation (VGER), which leverages the advances of large-scale video generation models to generate a moving camera video conditioned on the input image, and demonstrates its ability to produce smooth motions that account for the captured geometry of a scene, all from a single RGB input image.