ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
ConsiSpace is proposed, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal, and utilizes unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability.