Planetary scene classification plays a fundamental role in geomorphological analysis and autonomous exploration missions. However, planetary terrains exhibit high intraclass structural variability, and their analysis relies on an extremely limited set of annotated samples, making exhaustive premission labeling impractical. Therefore, recognition systems should be able to identify unseen classes with few (and sometimes no) training examples. This naturally motivates the adoption of the zero-shot learning (ZSL) paradigm for planetary scene classification. Existing solutions either require large-scale pretraining or employ heterogeneous and attention-intensive designs, limiting their practicality in data-scarce and resource-constrained planetary environments. To address these issues, we propose HiL-SSM, a hierarchical interactive linear state-space modeling framework for zero-shot planetary scene classification. It employs a unified, attention-free architecture based on structured state-space models (SSMs), enabling joint optimization of visual representation learning and semantic alignment within a single backbone framework. Importantly, it does not rely on large-scale pretraining. A hierarchical stage-wise interaction mechanism is introduced to progressively refine visual–semantic correspondence across multiple representation levels, enabling stronger alignment between geomorphological structures and semantic descriptors. Experiments on the ZSMars dataset demonstrate that the proposed framework achieves favorable classification performance under multiple seen/unseen splits while balancing computational complexity and accuracy.
Xiaomeng Tan, Changbin Xue, Bobo Xi et al.· IEEE Transactions on Geoscie...· 0 citations
Vision-Language-Action (VLA) model plays a crucial role in embodied decision making. While practical deployment requires fast inference under limited onboard computation, a full forward pass through the vision-language model makes such deployment challenging. To address this issue, existing methods typically employ lightweight techniques to compress the backbone. However, these information-lossy methods degrade spatial representations for action generation. In contrast, rate-distortion principles aim to reduce computation while retaining control-sufficient information. Inspired by this insight, we introduce Effective-Edge Flow, an action-aligned attribution measure that quantifies the marginal contribution of token interactions across network depth. This analysis reveals a consistent depth asymmetry, with visual evidence dominating early layers and linguistic reasoning sustaining task-relevant influence into deeper layers. Building on this structure, we propose MoDeVLA, the first rate-distortion driven efficient VLA model that performs token-wise depth allocation via Mixture-of-Depth Conditioning and integrates shallow visual-spatial with deep textual-logical features for action conditioning. Extensive real-robot evaluations across 20 tasks and multiple embodiments demonstrate that MoDeVLA preserves task performance while reducing latency by about 38% and FLOPs by 86% on edge device NVIDIA Jetson Orin, highlighting its strong ability for embodied systems deployment.
Weiying Xie, Qingcheng Zeng, Zihan Meng et al.· Proceedings of the 32nd ACM...· 0 citations
Vision-Language-Action (VLA) model plays a crucial role in embodied decision making. While practical deployment requires fast inference under limited onboard computation, a full forward pass through the vision-language model makes such deployment challenging. To address this issue, existing methods typically employ lightweight techniques to compress the backbone. However, these information-lossy methods degrade spatial representations for action generation. In contrast, rate-distortion principles aim to reduce computation while retaining control-sufficient information. Inspired by this insight, we introduce Effective-Edge Flow, an action-aligned attribution measure that quantifies the marginal contribution of token interactions across network depth. This analysis reveals a consistent depth asymmetry, with visual evidence dominating early layers and linguistic reasoning sustaining task-relevant influence into deeper layers. Building on this structure, we propose MoDeVLA, the first rate-distortion driven efficient VLA model that performs token-wise depth allocation via Mixture-of-Depth Conditioning and integrates shallow visual-spatial with deep textual-logical features for action conditioning. Extensive real-robot evaluations across 20 tasks and multiple embodiments demonstrate that MoDeVLA preserves task performance while reducing latency by about 38% and FLOPs by 86% on edge device NVIDIA Jetson Orin, highlighting its strong ability for embodied systems deployment.
Weiying Xie, Qingcheng Zeng, Zihan Meng et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.