This work shows that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass, enabling interactive 4D-controllable streaming generation for the first time.
Abstract
Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing approach captures only part of this: camera-parameter methods steer the viewpoint but cannot move objects, 2D-trajectory methods act in the image plane and ignore depth and occlusion, and recent 3D methods add geometry but run only offline at a fixed length. In particular, none combines 3D-consistent control of both camera and objects with real-time, streaming generation. Here we show that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass. To learn this interface at scale, we mine in-the-wild video for 3D motion supervision, yielding OpenVidHD-Motion3D, and encode it with a lightweight Geometric Motion Head that plugs into a pretrained video diffusion model. Because this encoder is temporally separable, we distill the model into a causal streaming student that generates arbitrarily long video in four denoising steps at memory independent of length. This unified design surpasses prior camera-only, 2D, and offline-3D methods in motion-control precision while covering modalities they address only in isolation. 4DStreamCtrl runs at 20 FPS on a single high-end GPU for 480p video and stays temporally coherent over hundreds of frames, enabling, to our knowledge, interactive 4D-controllable streaming generation for the first time. More broadly, grounding generation in explicit 3D geometry with efficient causal inference points toward interactive world models with closed-loop spatiotemporal control, from controllable simulators to real-time visual imagination for embodied agents.
This work proposes TARS, a 3D-free video re-shooting paradigm that provides more accurate and temporally consistent camera control than prior methods, and introduces self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction.
Jiwen Liu, Shujuan Li, Xiaohan Li et al.· arXiv.org· 0 citations
CameraAnything is introduced, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters, and a scalable synthetic pipeline is developed that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations.
Yixuan Li, Yanhong Zeng, K. Cheng et al.· arXiv.org· 1 citation
A new benchmark for evaluating camera pose estimation is introduced, called ORBIT, to leverage online panoramic 360{\deg} video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery.
S. Sabour, Linyi Jin, Richard Tucker et al.· 0 citations
StateFlow is presented, a state-centric framework for generative previsualization that uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired.
Yuyang Yin, Zixiang Li, Longxuan Deng et al.· 1 citation
D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-position correspondence as a practical 4D rendering condition.
Junhao Chen, Mingjin Chen, He Zhang et al.· 0 citations
RynnWorld-4D is introduced, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process, and achieves state-of-the-art performance on real-world dexterous bimanual manipulation tasks, particularly excelling in tasks demanding spatial precision and temporal coordination.
Haoyu Zhao, Xingyue Zhao, Siteng Huang et al.· 1 citation
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.
Microsoft Research Blog· microsoft.comJul 30, 2026
Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 29, 2026
The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.