Skip to content

FRAM: Trajectory-Guided Visual Feature Selection for Compact Language-Conditioned Robot Manipulation

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

This work proposes the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input and selects visual information based on future motion to obtain both high performance and robustness in a small robot policy.

Abstract

Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of $\pi_0$ with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.

View source

Similar papers

Preprint Aug 2026

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

TrAct is proposed, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction, enabling more accurate world modeling and stronger robot generalization.

Zhi-Hang Cao, Howard Ji, Kevin Zhang et al. · 3 citations
#computer vision Preprint Sep 2026

MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model

MaskVLA, a masking-based fine-tuning strategy that randomly masking a small portion of the main camera's visual information leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance.

Yuxuan Jiang, Jia-Ying Huang, Ge Wang et al. · 0 citations
Preprint Sep 2026

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...

Jun-Le Li, Weixian Waylon Li, Fu-Xiang Wu et al. · 1 citation
Preprint Sep 2026

FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation

Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future obs...

Khang Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen et al. · 0 citations
Preprint Sep 2026

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction, establishes learning functionally structured intermediate...

Yu-Peng Zheng, Xiang Li, Song-En Gu et al. · 2 citations
Preprint Sep 2026

H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs)...

Xiong-Feng Peng, Lu Xu, Yan-Dong Wang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.