G2TAM: Geometry Grounded Track Anything Model
The Geometry Grounded Tracking Anything Model is proposed, a unified framework for promptable instance tracking in 3D using only unordered RGB images or videos and delivers strong cross-view consistency, promptable instance spatial tracking, video object segmentation and spatial reconstruction, establishing a foundation for interactive, geometry-grounded spatial reasoning.