Skip to content
Preprint

StereoPatch: Patch-Aligned RGB-Depth Fusion for Spatial Perception in Robot Manipulation

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

StereoPatch is introduced, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction and suggests that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction.

Abstract

Recent advances in robot imitation learning have produced visuomotor policies that predict actions directly from visual observations. Yet visually similar scenes can require different actions as target position, object height, or contact geometry changes. Pretrained RGB features may map these geometrically distinct states to similar policy inputs, while simply adding depth requires the policy to learn RGB-depth correspondence from the same limited demonstrations used to learn control. We introduce StereoPatch, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction. On a shared 2-D patch grid, asymmetric cross-attention incorporates depth information into the corresponding RGB features before action decoding. The resulting StereoPatch Tokens provide a geometry-aware visual representation that can condition general visuomotor policies without changing their underlying learning objectives. Across six real-robot tasks, StereoPatch achieves higher closed-loop success than appearance-only, geometry-only, raw RGB-D, and late-fusion baselines. Additional experiments across three simulation suites evaluate compatibility across visuomotor policy architectures, spatial generalization, and operating limits. Results suggest that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction, rather than supplying it as an independent modality. Project page: https://aus.bot/research/stereopatch/.

View source

Similar papers

Preprint Sep 2026

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

This work introduces \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow.

Wen-Bo Li, Yi-Teng Chen, Wen-Hao Li et al. · 0 citations
Preprint Aug 2026

VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation

Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provid...

Jiaying Chen, Wen-Long Dong, Yan Huang et al. · 0 citations
Preprint Sep 2026

ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning

Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each nove...

Chiyoung Kim, Min-Sung Choi, Jin-Ho Ju et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild

Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-...

Hong-Yu Ma, Hai-Rong Qu, Shi-Qi Zhao et al. · 0 citations
Preprint Aug 2026

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.

Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al. · 1 citation
Preprint Sep 2026

AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations

Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sens...

Soham Patil, O. Gunjal, Sourabh P. Bhosale et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.