Skip to content
Review Open access

Vision‐Language‐Action Models for Embodied Artificial Intelligence: A Comprehensive Survey

Aug 2026 · Journal of Field Robotics · 0 citations · 31 references

TL;DR

A comprehensive review of VLA models for Embodied AI from an action‐generation perspective and proposes an action‐generation‐centered taxonomy that categorizes VLA models into three paradigms: direct policy learning, generative action modeling, and reasoning‐guided modeling.

Abstract

Embodied artificial intelligence (Embodied AI) aims to develop agents that can perceive, reason, and act through continuous interaction with physical world. To achieve this goal, agents must not only understand visual scenes and language instructions, but also transform multimodal information into executable actions. This requirement calls for models that can integrate perception, semantic reasoning, and action generation within a unified decision‐making framework. Recent advances in large language models (LLMs) have provided important technical foundations for this integration, leading to emergence of vision‐language‐action (VLA) models as a central paradigm for embodied systems. Although existing surveys have reviewed Embodied AI from perspectives such as robotic systems, simulation environments and task settings, few have systematically analyzed action‐generation mechanisms of VLA models. This limitation makes it difficult to distinguish VLA architectures. To address this gap, we provide a comprehensive review of VLA models for Embodied AI from an action‐generation perspective. We first formulate VLA models as embodied decision‐making systems that map visual observations and language instructions to action sequences through interaction with dynamic environments. Building on this formulation, we review historical evolution of VLA models. Furthermore, we propose an action‐generation‐centered taxonomy that categorizes VLA models into three paradigms: direct policy learning, generative action modeling, and reasoning‐guided modeling. For each paradigm, we analyze representative methods, action‐generation mechanisms, advantages and limitations, thereby providing a coherent framework for comparing diverse VLA architectures. Beyond model architectures, we summarize key resources and benchmarking settings and discuss recent applications. Finally, we identify major challenges and future directions. By organizing VLA models around the central problem of action generation, we aim to provide a structured technical map for future research toward more general, reliable, and physically grounded Embodied AI.

Read PDF

Similar papers

Review Open access Aug 2026

Intent-Driven Embodied Artificial Intelligence

Embodied Artificial Intelligence (Embodied AI) has emerged as a promising paradigm for developing more general and adaptive intelligent systems, emphasizing that intelligence emerges from continuous interaction among perception, cognition, and action in real-world environments. Recent advances increasingly integrate large language models and multimodal learning into embodied agents; however, most existing approaches remain correlation-driven, relying on implicit objectives, task-specific rewards, or prompt-level instructions. As a consequence, intent is rarely represented explicitly, limiting causal coherence, long-horizon consistency, and robust value alignment in open-world settings. In this Review, we synthesize recent progress in Embodied AI and articulate Intent-Driven Embodied Artificial Intelligence (IDEAI) as a system-level organizing framework in which intent functions as an explicit, revisable, and verifiable mediating construct between human goals, environmental constraints, and agent behavior. Building on this synthesis, we propose a four-layer conceptual organization-semantic grounding, concept generation and learning, intent modeling, and value alignment-that clarifies how explicit intent mediates perception, cognition, and action in embodied systems. We analyze how existing techniques address recurring failure modes along the intent-to-execution pipeline and highlight the limitations that arise when intent remains implicit. By making intent explicit, revisable, and value-constrained where such structure is needed, IDEAI supports interpretable decision-making, adaptive task decomposition, and value-consistent behavior in open-ended, human-interactive, and safety-critical embodied domains, providing a unifying perspective for advancing Embodied AI toward robust, socially deployable intelligent systems.

Nanning Zheng · 0 citations
Review Open access Jul 2026

Large Language Models for Task Planning in Embodied AI: A Survey

A structured taxonomy is presented that organizes existing work into three complementary paradigms that represent dominant architectural tendencies in current LLM-based embodied task planning research, and compares these paradigms along dimension of accuracy, robustness, scalability, efficiency, and sim-to-real transfer.

Zhen Zhang · 0 citations
Preprint Aug 2026

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Capek 0.5 is presented, an embodied vision-language model built around an execution-centric capability taxonomy that improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.

Ying Chen, Weizhen Li, Zhe Hu et al. · 0 citations
Open access Aug 2026

PFEA: a VLM-based high-level natural language planning and feedback embodied agent for human-centered AI

A closed-loop framework for planning and evaluation of a vision-language model-based robotic manipulation agent operating in tabletop object rearrangement and manipulation tasks and demonstrates the potential of closed-loop vision-language planning for human-centered robotic manipulation.

Wenbin Ding, Jun Chen, Mingjia Chen et al. · 0 citations
Review Sep 2026

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enabling robots to work in increasingly complex environments. However, these paradigms are typically developed in isolation, resulting in fragmented systems that struggle with generalization, long-horizon temporal reasoning and planning, and deployment in unstructured environments. In this survey, we present a unified perspective on robot learning by organizing the existing methods along three complementary axes: understanding through representation learning, acting through VLA models, and reasoning through world models. We introduce a structured taxonomy that captures key design choices in environment representation, policy learning, and predictive modeling, and summarize the recent progress in these domains. Beyond classifying the existing works, we analyze how these components interact, discuss common limitations, and highlight emerging trends towards more integrated systems. Through this lens, we identify the challenges in the domain of robot learning, including uncertainty quantification, out-of-distribution generalization, cross-embodiment transfer, long-context understanding, and long-horizon planning. We argue that these challenges arise not only from limitations within individual components but also from the lack of integration across perception, action, and reasoning. Building on this analysis, we outline future directions towards unified, physically grounded, and probabilistic robot learning to develop robust robotic systems that maintain consistent internal representations and support decision making over extended interactions in real-world environments.

Shaunak A. Mehta, Ananya Hazarika, Hao-Chen Zhang et al. · 0 citations
Review Open access Aug 2026

Vision-and-Language Navigation: A Component-Centric Survey of Interactions, Coupling, and Deployment

Vision-and-Language Navigation (VLN) is a representative task in embodied artificial intelligence, requiring agents to perceive, understand, and make navigation decisions in partially observable environments according to natural language instructions. As research has expanded from early discrete simulation benchmarks to continuous control, interactive clarification, open-vocabulary perception, and real-world robotic deployment, VLN has evolved from a path-following multimodal task into an important research area connecting language understanding, environment modeling, spatial reasoning, and embodied execution. Existing surveys mainly organize the literature by timeline, model paradigm, or benchmark, while paying less attention to the internal components of VLN systems and their functional coupling. In this survey, we revisit VLN from a component-internal perspective, viewing it as a navigation system composed of internal components such as instructions, environment representations, and embodied agents, and organizing existing work around the functions and interactions of these components. Specifically, we summarize task definitions, datasets, and evaluation settings, and review representative methods and technical progress in instruction understanding and action generation, instruction–environment alignment, and robot–environment interaction understanding. We further discuss key trends as VLN moves from closed benchmarks toward open-world and real-world deployment, including reasoning-enhanced planning, open-vocabulary and online semantic mapping, long-horizon memory and structured spatial representation, and sim-to-real transfer across platforms. We hope this survey provides a clearer component-level analytical framework for understanding the evolution of internal VLN capabilities and for informing future method design and embodied-system deployment.

Xiang-Xun Wu, Yin-Sheng Wu, Xiaojiang Peng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.