面向具身智能的视觉—语言—动作模型研究进展
Survey on embodied intelligence: advances in vision-language-action models
- 2026年31卷第8期 页码:2731-2773
收稿:2025-11-17,
修回:2026-04-03,
录用:2026-04-07,
网络首发:2026-04-07,
纸质出版:2026-08-16
DOI: 10.11834/jig.250544
移动端阅览

浏览全部资源
扫码关注微信
收稿:2025-11-17,
修回:2026-04-03,
录用:2026-04-07,
网络首发:2026-04-07,
纸质出版:2026-08-16
移动端阅览
具身智能作为人工智能与机器人学交叉的前沿领域,旨在使智能体通过与物理世界的动态交互来感知、推理并执行任务。然而,传统基于深度学习的级联式感知—控制模型在开放、动态环境下泛化能力不足,且高度依赖大规模标注数据。近年来,视觉—语言—动作模型(vision-language-action models, VLA)通过融合视觉感知、语言理解与动作生成,为具身智能的研究与应用提供了新的推动力。本文系统梳理了基于VLA的具身智能研究进展,从发展历程、模型架构、系统分类、训练与评估等方面展开综述。首先,本文追溯了视觉与语言基础模型的演进脉络,并阐述VLA概念的提出背景;随后,本文深入剖析VLA的关键技术模块,包括视觉编码、语言表征及动作词元化与解码机制;在此基础上,引入系统架构分类法,将现有工作归纳为单系统、双系统与层次化3类,并分析其设计权衡与适用场景;此外,本文总结了模型的预训练与后训练策略,并梳理了仿真及真实环境下的主流评测基准;最后,分析了VLA在实时推理效率、数据质量、环境泛化性与安全伦理等维度面临的挑战,并展望从被动感知到主动推理、持续学习、场景泛化与可靠部署等未来方向。本文旨在为相关研究者提供系统的技术参考,推动VLA在开放世界具身任务中的理论发展和应用落地。本文提及的算法、数据集和评估指标已汇总至
https://github.com/DefaultRui/vision-language-action-models-for-embodied-AI
https://github.com/DefaultRui/vision-language-action-models-for-embodied-AI
。
Embodied artificial intelligence aims to endow agents with the ability to perceive, reason, and act through continuous interaction with the physical world. In contrast with disembodied intelligence, which mostly operates on static data and symbolic tasks, embodied intelligence must handle sensorimotor coupling, environmental uncertainty, long-horizon decision-making, and real-time control in open-world settings. These requirements expose the limitations of traditional perception-planning-control pipelines. Conventional robotic systems are often built from cascaded modules for visual perception, language understanding, state estimation, task planning, and low-level control. Although such pipelines benefit from modular engineering and interpretability, they freque
ntly suffer from error accumulation across stages, brittle generalization to novel tasks and scenes, and heavy dependence on manually designed interfaces, reward functions, and domain-specific annotations. Recent advances in large-scale foundation models have opened a new path for embodied intelligence by enabling the integration of visual grounding, language reasoning, and action generation within a shared model space. This trend has led to the rapid emergence of vision-language-action (VLA) models, which aim to translate multimodal observations and task instructions directly into executable actions or action-conditioned policies, narrowing the gap between semantic understanding and physical execution.This survey presents a systematic review of VLA models in embodied intelligence from the perspectives of historical development, technical architecture, system design, training strategies, evaluation protocols, and future trends. We first revisit the technological trajectory that made VLA possible. The evolution of large vision models and large language models has substantially improved semantic representation, multimodal alignment, and knowledge transfer, while progress in robot learning has highlighted the importance of scalable policy learning from heterogeneous data sources. Within this context, VLA models can be considered the convergence of three lines of research: multimodal foundation models for perception and reasoning, imitation and reinforcement learning for embodied decision-making, and robot action modeling for continuous control. Their significance lies not merely in adding language prompts to robot policies, but also in establishing a unified interface through which high-level task intent, environmental perception, and low-level action generation can be jointly optimized.A central objective of this survey is to clarify the internal structure of VLA systems. We analyze the key technical modules that constitute current VLA pipelines. On the perception side, VLA models rely on visual encoders to transfo
rm raw observations into compact and task-relevant representations. These observations may include RGB images, depth maps, multi-view input, temporal image sequences, or structured spatial representations, such as 3D or bird’s-eye-view features. The choice of visual representation directly influences the model’s capability to capture object semantics, spatial relations, affordances, and scene dynamics. On the language side, textual encoders provide instruction understanding, goal specification, contextual reasoning, and cross-modal alignment. Language not only describes tasks but also offers a flexible interface for compositional generalization, abstract planning, and human-robot interaction. The crucial problem of action representation lies between perception and control. Compared with visual and textual tokens, action signals are continuous, temporally dense, and strongly constrained by embodiment. Therefore, we pay special attention to action tokenization and decoding, reviewing how actions can be represented as discrete tokens, continuous vectors, chunked trajectories, latent action codes, or generative distributions. We further compare the major decoding paradigms, including direct regression, classification after discretization, autoregressive sequence modeling, diffusion-based policy generation, and flow-based generation, and then discuss their respective trade-offs in precision, expressivity, stability, and inference speed.Beyond module-level analysis, this survey proposes a system-level taxonomy that categorizes existing VLA models into three broad classes: single systems, dual-system architectures, and hierarchical systems. Single systems integrate perception, reasoning, and action generation within a unified model, and they are attractive for end-to-end deployment, parameter sharing, and simplified optimization. However, they may experience difficulties in balancing high-level deliberation with fast control. Dual-system architectures address this issue by explicitly separating slow reasoning from fast
acting, frequently coupling a deliberative module for semantic planning with a lightweight controller for reactive execution. This design is particularly relevant to robotics because reasoning and motor control operate under different temporal constraints. Hierarchical systems introduce structured intermediate abstractions, such as subgoals, skills, waypoints, object-centric plans, or symbolic task graphs, to improve modularity, interpretability, and reuse across tasks. We argue that this taxonomy provides a useful perspective for understanding the design space of VLA models and analyzing how different systems trade off generalization, efficiency, and robustness.The survey also reviews the training and evaluation ecosystem that supports VLA research. Current VLA models are typically trained through multistage pipelines that involve internet-scale visual-language corpora, simulated interactive data, teleoperated demonstrations, and real robot trajectories. Pretraining equips the model with broad perceptual and linguistic priors, while post-training and task-specific adaptation align those priors with embodiment, action semantics, and control requirements. We summarize commonly used strategies, such as supervised imitation learning, behavior cloning, instruction tuning, parameter-efficient fine-tuning, model distillation, reinforcement learning augmentation, and hybrid training paradigms that combine offline and online data. We further discuss the role of simulator-based benchmarks and real-world evaluations. Simulation enables scalable and reproducible testing under controlled settings, while real robot experiments reveal embodiment gaps, latency constraints, and deployment failures that are not fully captured by simulation alone. In reviewing the literature, we emphasize that cross-paper comparisons must be interpreted with caution because reported results are often influenced by differences in robot morphology, sensor configurations, action horizons, training data quality, evaluation metrics, hardware platforms,
and implementation details.Despite their remarkable progress, VLA models still face several fundamental challenges. First, real-time inference remains a major bottleneck. Many powerful multimodal models rely on autoregressive decoding or iterative generative sampling, which may be too slow for high-frequency control. Second, action learning suffers from limited robotic data, inconsistent annotation quality, and poor cross-platform transfer, especially when embodiments differ substantially. Third, environmental generalization is far from solved. Models trained on curated datasets may degrade sharply in cluttered scenes, long-horizon tasks, dynamic environments, or safety-critical settings where perception is ambiguous and recovery from failure is required. Fourth, current VLA systems still exhibit limited causal reasoning and world modeling capabilities. Many methods align observations and instructions effectively but lack deep predictive understanding of object dynamics, interaction consequences, and long-term task structure. Fifth, evaluation remains fragmented. Existing studies often emphasize task success rate while underreporting execution frequency, safety violations, intervention cost, recovery capability, and robustness under distribution shift. Finally, broader issues of trustworthiness, including interpretability, failure diagnosis, ethical risk, and human-centered deployment, have become increasingly important as VLA systems move from laboratory demonstrations toward real applications.In the future, we argue that the next stage of VLA research will be defined by several converging directions. One is the shift from passive perception-conditioned action generation to active reasoning-driven embodied intelligence, in which agents can query, explore, verify, and adapt during execution. Another is the integration of richer world models that support prediction, imagination, and planning across time and modality. Lifelong and continual learning will also be essential for agents that are operating in evolving
environments. In addition, future VLA systems are likely to expand beyond fixed single-robot manipulation toward integrated mobile manipulation, long-horizon household assistance, multi-agent collaboration, and broader open-world autonomy. Achieving these objectives will require not only stronger models, but also better data curation, more principled action representations, more realistic evaluation standards, and more reliable safety mechanisms. By synthesizing the development of VLA models from architecture to deployment, this survey aims to provide a structured reference for researchers and practitioners, and to support the advancement of embodied intelligence toward more general, robust, and trustworthy interactions in the physical world. Furthermore, a comprehensively curated list of open-source VLA algorithms, datasets, and simulation benchmarks discussed in this survey is updated at
https://github.com/DefaultRui/vision-language-action-models-for-embodied-AI
https://github.com/DefaultRui/vision-language-action-models-for-embodied-AI
.
Abdin M , Aneja J , Behl H , Bubeck S , Eldan R , Gunasekar S , et al . 2024 . Phi-4 Technical Report [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.08905.pdf https://arxiv.org/pdf/2412.08905.pdf
Achiam J , Adler S , Agarwal S , Ahmad L , Akkaya I , Aleman F L , et al . 2023 . GPT-4 technical report [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2303.08774.pdf https://arxiv.org/pdf/2303.08774.pdf
Agarwal N , Ali A , Bala M , Balaji Y , Barker E , Cai T , et al . 2025 . Cosmos world foundation model platform for physical AI [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.03575.pdf https://arxiv.org/pdf/2501.03575.pdf
Alayrac J B , Donahue J , Luc P , Miech A , Barr I , Hasson Y , et al . 2022 . Flamingo: a visual language model for few-shot learning // Proceedings of the 36th International Conference on Neural Information Processing Systems . New Orleans, USA : Curran Associates Inc.: #1723
Asif U , Tang J B and Harrer S . 2018 . GraspNet: an efficient convolutional neural network for real-time grasp detection for low-powered devices // Proceedings of the 27th International Joint Conference on Artificial Intelligence . Stockholm, Sweden : ijcai.org: 4875 - 4882 [ DOI: 10.24963/ijcai.2018/677 http://dx.doi.org/10.24963/ijcai.2018/677 ]
Bai J Z , Bai S , Chu Y F , Cui Z Y , Dang K , Deng X D , et al . 2023 . Qwen technical report [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2309.16609.pdf https://arxiv.org/pdf/2309.16609.pdf
Belkhale S , Ding T L , Xiao T , Sermanet P , Vuong Q , Tompson J , et al . 2024 . RT-H: action hierarchies using language // Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.049 http://dx.doi.org/10.15607/RSS.2024.XX.049 ]
Beyer L , Steiner A , Pinto A S , Kolesnikov A , Wang X , Salz D , et al . 2024 . PaliGemma: a versatile 3B VLM for transfer [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2407.07726.pdf https://arxiv.org/pdf/2407.07726.pdf
Bi J X , Ma K Y , Hao C , Shou M Z and Soh H . 2025 . VLA-touch: enhancing vision-language-action models with dual-level tactile feedback [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.17294.pdf https://arxiv.org/pdf/2507.17294.pdf
Bjorck J , Castañeda F , Cherniadev N , Da X Y , Ding R Y , Fan L X , et al . 2025 . GR00T N1: an open foundation model for generalist humanoid robots [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.14734.pdf https://arxiv.org/pdf/2503.14734.pdf
Black K , Brown N , Darpinian J , Dhabalia K , Driess D , Esmail A , et al . 2025 . π 0 . 5 : a vision-language-action model with open-world generalization [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2504.16054.pdf https://arxiv.org/pdf/2504.16054.pdf
Black K , Brown N , Driess D , Esmail A , Equi M , Finn C , et al . 2024 . π 0 : a vision-language-action flow model for general robot control [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2410.24164.pdf https://arxiv.org/pdf/2410.24164.pdf
Brohan A , Brown N , Carbajal J , Chebotar Y , Dabis J , Finn C , et al . 2023 . RT-1: robotics transformer for real-world control at scale // Proceedings of Robotics: Science and Systems XIX . Daegu, Korea(South) : MIT Press [ DOI: 10.15607/RSS.2023.XIX.025 http://dx.doi.org/10.15607/RSS.2023.XIX.025 ]
Bruce J , Dennis M D , Edwards A , Parker-Holder J , Shi Y G , Hughes E , et al . 2024 . Genie: generative interactive environments // Proceedings of the 41st International Conference on Machine Learning . Vienna, Austria : PMLR: 4603 - 4623
Bu Q W , Cai J S , Chen L , Cui X Q , Ding Y , Feng S Y , et al . 2025a . AgiBot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.06669.pdf https://arxiv.org/pdf/2503.06669.pdf
Bu Q W , Li H Y , Chen L , Cai J S , Zeng J , Cui H M , et al . 2024 . Towards synergistic, generalized, and efficient dual-system for robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2410.08001.pdf https://arxiv.org/pdf/2410.08001.pdf
Bu Q W , Yang Y T , Cai J S , Gao S Y , Ren G H , Yao M Q , et al . 2025b . UniVLA: learning to act anywhere with task-centric latent actions [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.06111.podf https://arxiv.org/pdf/2505.06111.podf
Cen J , Yu C H , Yuan H J , Jiang Y M , Huang S T , Guo J Y , et al . 2025 . WorldVLA: towards autoregressive action world model [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.21539.pdf https://arxiv.org/pdf/2506.21539.pdf
Chameleon Team . 2025 . Chameleon: mixed-modal early-fusion foundation models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2405.09818.pdf https://arxiv.org/pdf/2405.09818.pdf
Cheang C , Chen S J , Cui Z R , Hu Y D , Huang L Q , Kong T , et al . 2025 . GR-3 technical report [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.15493.pdf https://arxiv.org/pdf/2507.15493.pdf
Cheang C L , Chen G Z , Jing Y , Kong T , Li H , Li Y F , et al . 2024 . GR-2: a generative video-language-action model with web-scale knowledge for robot manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2410.06158.pdf https://arxiv.org/pdf/2410.06158.pdf
Chen H , Liu J M , Gu C Y , Liu Z Y , Zhang R R , Li X Q , et al . 2025a . Fast-in-slow: a dual-system foundation model unifying fast manipulation within slow reasoning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.01953.pdf https://arxiv.org/pdf/2506.01953.pdf
Chen H P , Li S F , Fan J M , Duan A Q , Yang C G , Navarro-Alarcon D , et al . 2025b . Human-in-the-loop robot learning for smart manufacturing: a human-centric perspective . IEEE Transactions on Automation Science and Engineering , 22 : 11062 - 11086 [ DOI: 10.1109/TASE.2025.3528051 http://dx.doi.org/10.1109/TASE.2025.3528051 ]
Chen J Y , Song W X , Ding P X , Zhou Z Y , Zhao H , Tang F L , et al . 2026 . Unified diffusion VLA: vision-language-action model via joint discrete denoising diffusion process // Proceedings of the 14th International Conference on Learning Representations . Rio de Janeiro, Brazil : OpenReview.net
Chen J Y , Zhu H Y , He X L , Wang Y F , Zhou J J , Chang W Z , et al . 2025c . DeepVerse: 4D autoregressive video generation as a world model [EB/OL]. [ 2025-06-11 ]. https://doi.org/10.48550/arXiv.2506.01103 https://doi.org/10.48550/arXiv.2506.01103
Chen X Y , Wei H X , Zhang P S , Zhang C H , Wang K X , Guo Y J , et al . 2025d . villa-X: enhancing latent action modeling in vision-language-action models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.23682.pdf https://arxiv.org/pdf/2507.23682.pdf
Chen Y , Ge Y Y , Li Y Z , Ge Y X , Ding M Y , Shan Y , et al . 2024 . Moto: latent motion token as the bridging language for robot manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.04445.pdf https://arxiv.org/pdf/2412.04445.pdf
Chen Y H , Tian S , Liu S G , Zhou Y T , Li H R and Zhao D B . 2025e . ConRFT: a reinforced fine-tuning method for VLA models via consistency policy //Proceedings of Robotics: Science and Systems. Los Angeles, USA: RSS Foundation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/XXI.019.pdf https://arxiv.org/pdf/XXI.019.pdf
Chen Z J , Niu R L , Kong H and Wang Q . 2025f . TGRPO: fine-tuning vision-language-action model via trajectory-wise group relative policy optimization [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.08440.pdf https://arxiv.org/pdf/2506.08440.pdf
Cheng A C , Ji Y D , Yang Z J , Gongye Z T , Zou X Y , Kautz J , et al . 2024 . NaVILA: legged robot vision-language-action model for navigation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.04453.pdf https://arxiv.org/pdf/2412.04453.pdf
Cherepanov E , Kachaev N , Kovalev A K and Panov A I . 2025 . Memory, benchmark and robots: a benchmark for solving complex tasks with reinforcement learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2502.10550.pdf https://arxiv.org/pdf/2502.10550.pdf
Cherti M , Beaumont R , Wightman R , Wortsman M , Ilharco G , Gordon C , et al . 2023 . Reproducible scaling laws for contrastive language-image learning // Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Vancouver, Canada : IEEE: 2818 - 2829 [ DOI: 10.1109/CVPR52729.2023.00276 http://dx.doi.org/10.1109/CVPR52729.2023.00276 ]
Chi C , Xu Z J , Feng S Y , Cousineau E , Du Y L , Burchfiel B , et al . 2025 . Diffusion policy: visuomotor policy learning via action diffusion . The International Journal of Robotics Research , 44 ( 10/11 ): 1684 - 1704 [ DOI: 10.1177/02783649241273668 http://dx.doi.org/10.1177/02783649241273668 ]
Choi J W , Yoon Y , Ong H , Kim J and Jang M . 2024 . LoTa-Bench: benchmarking language-oriented task planners for embodied agents // Proceedings of the 12th International Conference on Learning Representations . Vienna, Austria : OpenReview.net
Cui C , Ding P X , Song W X , Bai S H , Tong X Y , Ge Z R , et al . 2025 . OpenHelix: a short survey, empirical analysis, and open-source dual-system VLA model for robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.03912.pdf https://arxiv.org/pdf/2505.03912.pdf
Damen D , Doughty H , Farinella G M , Furnari A , Kazakos E , Ma J , et al . 2022 . Rescaling egocentric vision: collection, pipeline and challenges for EPIC-KITCHENS-100 . International Journal of Computer Vision , 130 ( 1 ): 33 - 55 [ DOI: 10.1007/s11263-021-01531-2 http://dx.doi.org/10.1007/s11263-021-01531-2 ]
Dang R H , Yuan Y Q , Zhang W Q , Xin Y F , Zhang B Q , Li L , et al . 2025 . ECBench: can multi-modal foundation models understand the egocentric world? A holistic embodied cognition benchmark // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Nashville, USA : IEEE: 24593 - 24602 [ DOI: 10.1109/CVPR52734.2025.02290 http://dx.doi.org/10.1109/CVPR52734.2025.02290 ]
Dasari S , Ebert F , Tian S , Nair S , Bucher B , Schmeckpeper K , et al . 2019 . RoboNet: large-scale multi-robot learning // Proceedings of the 3rd Annual Conference on Robot Learning . Osaka, Japan : PMLR: 885 - 897
Davisson L D . 1972 . Rate-distortion theory and application . Proceedings of the IEEE , 60 ( 7 ): 800 - 808 [ DOI: 10.1109/proc.1972.8779 http://dx.doi.org/10.1109/proc.1972.8779 ]
Deng S L , Yan M , Wei S L , Ma H X , Yang Y X , Chen J Y , et al . 2025 . GraspVLA: a grasping foundation model pre-trained on billion-scale synthetic action data [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.03233.pdf https://arxiv.org/pdf/2505.03233.pdf
Devlin J , Chang M W , Lee K and Toutanova K . 2019 . BERT: pre-training of deep bidirectional transformers for language understanding // Proceedings of 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, USA : ACL: 4171 - 4186 [ DOI: 10.18653/v1/N19-1423 http://dx.doi.org/10.18653/v1/N19-1423 ]
Dey S , Zaech J N , Nikolov N , Van Gool L and Paudel D P . 2025 . ReVLA: reverting visual domain limitation of robotic foundation models // Proceedings of 2025 IEEE International Conference on Robotics and Automation (ICRA) . Atlanta, USA : IEEE: 8679 - 8686 [ DOI: 10.1109/ICRA55743.2025.11128823 http://dx.doi.org/10.1109/ICRA55743.2025.11128823 ]
Ding P X , Ma J F , Tong X Y , Zou B H , Luo X X , Fan Y G , et al . 2025a . Humanoid-VLA: towards universal humanoid control with visual integration [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2502.14795.pdf https://arxiv.org/pdf/2502.14795.pdf
Ding P X , Zhao H , Zhang W J , Song W X , Zhang M , Huang S T , et al . 2025b . QUAR-VLA: vision-language-action model for quadruped robots // Proceedings of the 18th European Conference on Computer Vision . Milan, Italy : Springer: 352 - 367 [ DOI: 10.1007/978-3-031-72652-1_21 http://dx.doi.org/10.1007/978-3-031-72652-1_21 ]
Doshi R , Walke H R , Mees O , Dasari S and Levine S . 2024 . Scaling cross-embodied learning : one policy for manipulation, navigation, locomotion and aviation// Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 496 - 512
Dosovitskiy A , Beyer L , Kolesnikov A , Weissenborn D , Zhai X H , Unterthiner T , et al . 2021 . An image is worth 16 × 16 words: transformers for image recognition at scale //Proceedings of the 9th International Conference on Learning Representations. [s.l.]: OpenReview.net.
Driess D , Xia F , Sajjadi M S M , Lynch C , Chowdhery A , Ichter B , et al . 2023 . PaLM-E: an embodied multimodal language model // Proceedings of the 40th International Conference on Machine Learning . Honolulu, USA : PMLR: 8469 - 8488
Ebert F , Yang Y L , Schmeckpeper K , Bucher B , Georgakis G , Daniilidis K , et al . 2022 . Bridge data: boosting generalization of robotic skills with cross-domain datasets // Proceedings of Robotics: Science and Systems XVIII . New York, USA : MIT Press [ DOI: 10.15607/RSS.2022.XVIII.063 http://dx.doi.org/10.15607/RSS.2022.XVIII.063 ]
Ehsani K , Gupta T , Hendrix R , Salvador J , Weihs L , Zeng K H , et al . 2024 . SPOC: imitating shortest paths in simulation enables effective navigation and manipulation in the real world // Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Seattle, USA : IEEE: 16238 - 16250 [ DOI: 10.1109/CVPR52733.2024.01537 http://dx.doi.org/10.1109/CVPR52733.2024.01537 ]
Esser P , Rombach R and Ommer B . 2021 . Taming transformers for high-resolution image synthesis // Proceedings of 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, USA : IEEE: 12868 - 12878 [ DOI: 10.1109/CVPR46437.2021.01268 http://dx.doi.org/10.1109/CVPR46437.2021.01268 ]
Fang H S , Fang H J , Tang Z Y , Liu J R , Wang C X , Wang J B , et al . 2024 . RH20T: a comprehensive robotic dataset for learning diverse skills in one-shot // Proceedings of 2024 IEEE International Conference on Robotics and Automation . Yokohama, Japan : IEEE: 653 - 660 [ DOI: 10.1109/ICRA57147.2024.10611615 http://dx.doi.org/10.1109/ICRA57147.2024.10611615 ]
Feng M T , Shen J H , Wu Z J , Peng W X , Zhong H , Guo Y L , et al . 2025 . Advancements in 3D vision understanding using multimodal large language models . Journal of Image and Graphics , 30 ( 6 ): 1744 - 1791
冯明涛 , 沈军豪 , 武子杰 , 彭伟星 , 钟杭 , 郭裕兰 , 等 . 2025 . 多模态大模型驱动的三维视觉理解技术前沿进展. 中国图象图形学报 , 30 ( 6 ): 1744 - 1791 [ DOI: 10.11834/jig.240588 http://dx.doi.org/10.11834/jig.240588 ]
Feng T , Wang W G and Yang Y . 2025 . A survey of world models for autonomous driving [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.11260.pdf https://arxiv.org/pdf/2501.11260.pdf
Firoozi R , Tucker J , Tian S , Majumdar A , Sun J K , Liu W Y , et al . 2025 . Foundation models in robotics: applications, challenges, and the future . The International Journal of Robotics Research , 44 ( 5 ): 701 - 739 [ DOI: 10.1177/02783649241281508 http://dx.doi.org/10.1177/02783649241281508 ]
Gao C K , Liu Z X , Chi Z H , Huang J S , Fei X , Hou Y W , et al . 2025 . VLA-OS: structuring and dissecting planning representations and paradigms in vision-language-action models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.17561.pdf https://arxiv.org/pdf/2506.17561.pdf
Ghosh D , Walke H , Pertsch K , Black K , Mees O , Dasari S , et al . 2024 . Octo: an open-source generalist robot policy // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.090 http://dx.doi.org/10.15607/RSS.2024.XX.090 ]
Gong J Y , Lou Y J , Liu F Q , Zhang Z W , Chen H M , Zhang Z Z , et al . 2023 . Scene point cloud understanding and reconstruction technologies in 3D space . Journal of Image and Graphics , 28 ( 6 ): 1741 - 1766
龚靖渝 , 楼雨京 , 柳奉奇 , 张志伟 , 陈豪明 , 张志忠 , 等 . 2023 . 三维场景点云理解与重建技术. 中国图象图形学报 , 28 ( 6 ): 1741 - 1766 [ DOI: 10.11834/jig.230004 http://dx.doi.org/10.11834/jig.230004 ]
Goyal R , Kahou S E , Michalski V , Materzynska J , Westphal S , Kim H , et al . 2017 . The “something something” video database for learning and evaluating visual common sense // Proceedings of 2017 IEEE International Conference on Computer Vision . Venice, Italy : IEEE: 5843 - 5851 [ DOI: 10.1109/ICCV.2017.622 http://dx.doi.org/10.1109/ICCV.2017.622 ]
Grauman K , Westbury A , Byrne E , Cartillier V , Chavis Z , Furnari A , et al . 2025 . Ego 4 D: around the world in 3 600 hours of egocentric video. IEEE Transactions on Pattern Analysis and Machine Intelligence , 47 ( 11 ): 9468 - 9509 [ DOI: 10.1109/TPAMI.2024.3381075 http://dx.doi.org/10.1109/TPAMI.2024.3381075 ]
Grauman K , Westbury A , Torresani L , Kitani K , Malik J , Afouras T , et al . 2024 . Ego-Exo4 D: understanding skilled human activity from first- and third-person perspectives// Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Seattle, USA : IEEE: 19383 - 19400 [ DOI: 10.1109/CVPR52733.2024.01834 http://dx.doi.org/10.1109/CVPR52733.2024.01834 ]
Guo Y F , Yu Z T , Liu A S , Zhou W B , Qiao T , Li B , et al . 2025 . Recent progress of the security research for multimodal large models . Journal of Image and Graphics , 30 ( 6 ): 2051 - 2081
郭园方 , 余梓彤 , 刘艾杉 , 周文柏 , 乔通 , 李斌 , 等 . 2025 . 多模态大模型安全研究进展. 中国图象图形学报 , 30 ( 6 ): 2051 - 2081 [ DOI: 10.11834/jig.250067 http://dx.doi.org/10.11834/jig.250067 ]
Guo Y J , Zhang J K , Chen X Y , Ji X , Wang Y J , Hu Y C , et al . 2025 . Improving vision-language-action model with online reinforcement learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.16664.pdf https://arxiv.org/pdf/2501.16664.pdf
Han B O , Kim J and Jang J . 2024 . A dual process VLA: efficient robotic manipulation leveraging VLM [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2410.15549.pdf https://arxiv.org/pdf/2410.15549.pdf
Hao P , Zhang C F , Li D Z , Cao X G , Hao X S , Cui S W , et al . 2025 . TLA: tactile-language-action model for contact-rich manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.08548.pdf https://arxiv.org/pdf/2503.08548.pdf
He K M , Chen X L , Xie S N , Li Y H , Dollr P and Girshick R . 2022 . Masked autoencoders are scalable vision learners // Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition . New Orleans, USA : IEEE: 15979 - 15988 [ DOI: 10.1109/CVPR52688.2022.01553 http://dx.doi.org/10.1109/CVPR52688.2022.01553 ]
He K M , Gkioxari G , Dollr P and Girshick R . 2017 . Mask R-CNN // Proceedings of 2017 IEEE International Conference on Computer Vision . Venice, Italy : IEEE: 2980 - 2988 [ DOI: 10.1109/ICCV.2017.322 http://dx.doi.org/10.1109/ICCV.2017.322 ]
He K M , Zhang X Y , Ren S Q and Sun J . 2016 . Deep residual learning for image recognition // Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition . Las Vegas, USA : IEEE: 770 - 778 [ DOI: 10.1109/CVPR.2016.90 http://dx.doi.org/10.1109/CVPR.2016.90 ]
Hirose N , Shah D , Sridhar A and Levine S . 2024 . SACSoN: scalable autonomous control for social navigation . IEEE Robotics and Automation Letters , 9 ( 1 ): 49 - 56 [ DOI: 10.1109/LRA.2023.3329626 http://dx.doi.org/10.1109/LRA.2023.3329626 ]
Hong Y N , Zhen H Y , Chen P H , Zheng S H , Du Y L , Chen Z F , et al . 2023 . 3D-LLM: injecting the 3D world into large language models // Proceedings of the 37th International Conference on Neural Information Processing Systems . New Orleans, USA : Curran Associates Inc.: #900
Hu E J , Shen Y L , Wallis P , Allen-Zhu Z , Li Y Z , Wang S A , et al . 2022 . LoRA: low-rank adaptation of large language models //Proceedings of the 10th International Conference on Learning Representations. [s.l.]: OpenReview .net
Hu Y D , Lin F Q , Zhang T , Yi L and Gao Y . 2023 . Look before you leap: unveiling the power of GPT-4V in robotic vision-language planning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2311.17842.pdf https://arxiv.org/pdf/2311.17842.pdf
Huang C P , Wu Y H , Chen M H , Wang Y C F and Yang F E . 2025a . ThinkAct: vision-language-action reasoning via reinforced visual latent planning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.16815.pdf https://arxiv.org/pdf/2507.16815.pdf
Huang J L , Wang S , Lin F Q , Hu Y H , Wen C and Gao Y . 2025b . Tactile-VLA: unlocking vision-language-action model's physical knowledge for tactile generalization [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.09160.pdf https://arxiv.org/pdf/2507.09160.pdf
Huang J Y , Yong S L , Ma X J , Linghu X K , Li P H , Wang Y , et al . 2024a . An embodied generalist agent in 3D world // Proceedings of the 41st International Conference on Machine Learning . Vienna, Austria : PMLR: 20413 - 20451
Huang S Y , Jiang Z K , Dong H , Qiao Y , Gao P and Li H S . 2023b . Instruct2Act: mapping multi-modality instructions to robotic actions with large language model [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2305.11176.pdf https://arxiv.org/pdf/2305.11176.pdf
Huang W L , Wang C , Li Y Z , Zhang R H and Li F F . 2024b . ReKep: spatio-temporal reasoning of relational keypoint constraints for robotic manipulation // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 4573 - 4602
Huang W L , Wang C , Zhang R H , Li Y Z , Wu J J and Li F F . 2023c . VoxPoser: composable 3D value maps for robotic manipulation with language models // Proceedings of the 7th Conference on Robot Learning . Atlanta, USA : PMLR: 540 - 562
Huang W L , Xia F , Xiao T , Chan H , Liang J , Florence P , et al . 2022 . Inner monologue: embodied reasoning through planning with language models // Proceedings of the 6th Conference on Robot Learning . Auckland, New Zealand : PMLR: 1769 - 1782
Hung C Y , Sun Q , Hong P F , Zadeh A , Li C , Tan U X , et al . 2025 . NORA: a small open-sourced generalist vision language action model for embodied tasks [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2504.19854.pdf https://arxiv.org/pdf/2504.19854.pdf
Hurst A , Lerer A , Goucher A P , Perelman A , Ramesh A , Clark A , et al . 2024 . GPT-4o system card [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2410.21276.pdf https://arxiv.org/pdf/2410.21276.pdf
Ichter B , Brohan A , Chebotar Y , Finn C , Hausman K , Herzog A , et al . 2022 . Do as I can, not as I say: grounding language in robotic affordances // Proceedings of the 6th Conference on Robot Learning . Auckland, New Zealand : PMLR: 287 - 318
James S , Ma Z C , Arrojo D R and Davison A J . 2020 . RLBench: the robot learning benchmark and learning environment . IEEE Robotics and Automation Letters , 5 ( 2 ): 3019 - 3026 [ DOI: 10.1109/LRA.2020.2974707 http://dx.doi.org/10.1109/LRA.2020.2974707 ]
Jang E , Irpan A , Khansari M , Kappler D , Ebert F , Lynch C , et al . 2021 . BC-Z: zero-shot task generalization with robotic imitation learning // Proceedings of the 5th Conference on Robot Learning . London, UK : PMLR: 991 - 1002
Jatavallabhula K M , Iyer G and Paull L . 2019 . gradSLAM: dense SLAM meets automatic differentiation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/1910.10672.pdf https://arxiv.org/pdf/1910.10672.pdf
Ji Y H , Tan H J , Shi J Y , Hao X S , Zhang Y , Zhang H Y , et al . 2025 . RoboBrain: a unified brain model for robotic manipulation from abstract to concrete // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Nashville, USA : IEEE: 1724 - 1734 [ DOI: 10.1109/CVPR52734.2025.00168 http://dx.doi.org/10.1109/CVPR52734.2025.00168 ]
Jia X G , Wang Q , Donat A , Xing B W , Li G , Zhou H Y , et al . 2024 . MaIL: improving imitation learning with selective state space models // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 3888 - 3907
Jian M W , Ling Y K , Zhang H R , Zhang L S and Ma J J . 2026 . Embodied intelligence-driven distracted driving detection: a framework and research prospects [J/OL]. Journal of Image and Graphics , 1 - 16 . https://doi.org/10.11834/jig.250514 ( https://doi.org/10.11834/jig.250514(
蹇木伟 , 凌钰坤 , 张昊然 , 张琳松 , 马嘉骏 . 2026 . 具身智能驱动的分心驾驶检测: 框架研究与前沿展望 [J/OL]. 中国图象图形学报 , 1 - 16 . [ 2026-5-29 ] https://doi.org/10.11834/jig.250514 https://doi.org/10.11834/jig.250514
Jiang J P , Xiao W Y , Lin Z Y , Zhang H Z , Ren T X , Gao Y , et al . 2025a . SOLAMI: social vision-language-action modeling for immersive interaction with 3D autonomous characters // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, USA : IEEE: 26887 - 26898 [ DOI: 10.1109/CVPR52734.2025.02504 http://dx.doi.org/10.1109/CVPR52734.2025.02504 ]
Jiang S , Fang X L , Roy N , Lozano-Pérez T , Kaelbling L P a nd Ancha S . 2025b . Streaming flow policy: simplifying diffusion/flow-matching policies by treating action trajectories as flow trajectories // Proceedings of the 9th Conference on Robot Learning . Seoul, Korea(South) : PMLR: 238 - 257
Jiang Y F , Gupta A , Zhang Z C , Wang G Z , Dou Y Q , Chen Y J , et al . 2022 . VIMA: general robot manipulation with multimodal prompts [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2210.03094.pdf https://arxiv.org/pdf/2210.03094.pdf
Jiang Z Y , Xie Y Q , Lin K , Xu Z J , Wan W K , Mandlekar A , et al . 2025c . DexMimicGen: automated data generation for bimanual dexterous manipulation via imitation learning // Proceedings of 2025 IEEE International Conference on Robotics and Automation (ICRA) . Atlanta, USA : IEEE: 16923 - 16930 [ DOI: 10.1109/ICRA55743.2025.11127809 http://dx.doi.org/10.1109/ICRA55743.2025.11127809 ]
Jones J , Mees O , Sferrazza C , Stachowicz K , Abbeel P and Levine S . 2025 . Beyond sight: finetuning generalist robot policies with heterogeneous sensors via language grounding // Proceedings of 2025 IEEE International Conference on Robotics and Automation . Atlanta, USA : IEEE: 5961 - 5968 [ DOI: 10.1109/ICRA55743.2025.11127987 http://dx.doi.org/10.1109/ICRA55743.2025.11127987 ]
Jülg T , Burgard W and Walter F . 2025 . Refined policy distillation: from VLA generalists to RL experts [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.05833.pdf https://arxiv.org/pdf/2503.05833.pdf
Kawaharazuka K , Oh J , Yamada J , Posner I and Zhu Y K . 2025 . Vision-language-action models for robotics: a review towards real-world applications . IEEE Access , 13 : 162467 - 162504 [ DOI: 10.1109/ACCESS.2025.3609980 http://dx.doi.org/10.1109/ACCESS.2025.3609980 ]
Kerbl B , Kopanas G , Leimkuehler T and Drettakis G . 2023 . 3D Gaussian splatting for real-time radiance field rendering . ACM Transactions on Graphics , 42 ( 4 ): # 139 [ DOI: 10.1145/3592433 http://dx.doi.org/10.1145/3592433 ]
Khazatsky A , Pertsch K , Nair S , Balakrishna A , Dasari S , Karamcheti S , et al . 2024 . DROID: a large-scale in-the-wild robot manipulation dataset // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.120 http://dx.doi.org/10.15607/RSS.2024.XX.120 ]
Kim H , Kang J , Kang H , Cho M , Kim S J and Lee Y . 2025a . UniSkill: imitating human videos via cross-embodiment skill representations [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.08787.pdf https://arxiv.org/pdf/2505.08787.pdf
Kim M J , Finn C and Liang P . 2025b . Fine-tuning vision-language-action models: optimizing speed and success [EB/OL]. [ 2025-02-27 ]. https://arxiv.org/pdf/2502.19645.pdf https://arxiv.org/pdf/2502.19645.pdf
Kim M J , Pertsch K , Karamcheti S , Xiao T , Balakrishna A , Nair S , et al . 2024 . OpenVLA: an open-source vision-language-action model // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 2679 - 2713
Kirillov A , Mintun E , Ravi N , Mao H Z , Rolland C , Gustafson L , et al . 2023 . Segment anything // Proceedings of 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . Paris, France : IEEE: 3992 - 4003 [ DOI: 10.1109/ICCV51070.2023.00371 http://dx.doi.org/10.1109/ICCV51070.2023.00371 ]
Lee J , Duan J F , Fang H Q , Deng Y Q , Liu S , Li B Y , et al . 2025 . MolmoAct: action reasoning models that can reason in space [EB/OL]. [ 2025-10-21 ]. https://arxiv.org/pdf/2508.07917.pdf https://arxiv.org/pdf/2508.07917.pdf
Li C M , Wen J J , Peng Y X , Peng Y and Zhu Y C . 2026 . PointVLA: injecting the 3D world into vision-language-action models . IEEE Robotics and Automation Letters , 11 ( 3 ): 2506 - 2513 [ DOI: 10.1109/LRA.2026.3653303 http://dx.doi.org/10.1109/LRA.2026.3653303 ]
Li C S , Zhang R H , Wong J , Gokmen C , Srivastava S , Martín-Martín R , et al . 2022a . BEHAVIOR-1K: a human-centered, embodied AI benchmark with 1 000 everyday activities and realistic simulation // Proceedings of the 6th Conference on Robot Learning . Auckland, New Zealand : PMLR: 80 - 93
Li H R , Chen Y H , Cui W B , Liu W H , Liu K , Zhou M C , et al . 2026 . Survey of vision-language-action models for embodied manipulation . Acta Automatica Sinica , 52 ( 1 ): 18 - 51
李浩然 , 陈宇辉 , 崔文博 , 刘卫恒 , 刘锴 , 周明才 , 等 . 2026 . 面向具身操作的视觉-语言-动作模型综述. 自动化学报 , 52 ( 1 ): 18 - 51 [ DOI: 10.16383/j.aas.c250394 http://dx.doi.org/10.16383/j.aas.c250394 ]
Li J N , Li D X , Savarese S and Hoi S C H . 2023 . BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models // Proceedings of the 40th International Conference on Machine Learning . Honolulu, USA : PMLR: 19730 - 19742
Li Q X , Liang Y B , Wang Z Y , Luo L , Chen X , Liao M Z , et al . 2024a . CogACT: a foundational vision-language-action model for synergizing cognition and action in robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2411.19650.pdf https://arxiv.org/pdf/2411.19650.pdf
Li S , Puig X , Paxton C , Du Y L , Wang C , Fan L X , et al . 2022b . Pre-trained language models for interactive decision-making // Advances in Neural Information Processing Systems 35 . New Orleans, USA : Neural Information Processing Systems Foundation: 31199 - 31212
Li S L , Gao L S , Cao J W and Hu Y B . 2025a . Graph-fused vision-language-action for policy reasoning in multi-arm robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2509.07957.pdf https://arxiv.org/pdf/2509.07957.pdf
Li S L , Wang J , Dai R , Ma W Y , Ng W Y , Hu Y B , et al . 2025c . RoboNurse-VLA: robotic scrub nurse system based on vision-language-action model // Proceedings of 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . Hangzhou, China : IEEE: 3986 - 3993 [ DOI: 10.1109/IROS60139.2025.11246030 http://dx.doi.org/10.1109/IROS60139.2025.11246030 ]
Li X H , Liu M H , Zhang H B , Yu C J , Xu J , Wu H T , et al . 2024b . Vision-language foundation models as effective robot imitators // Proceedings of the 12th International Conference on Learning Representations . Vienna, Austria : OpenReview.net
Li X L , Hsu K , Gu J Y , Mees O , Pertsch K , Walke H R , et al . 2024c . Evaluating real-world robot manipulation policies in simulation // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 3705 - 3728
Li Y , Deng Y Q , Zhang J , Jang J , Memmel M , Garrett C R , et al . 2025d . HAMSTER: hierarchical action models for open-world robot manipulation // Proceedings of the 13th International Conference on Learning Representations . Singapore, Singapore : OpenReview.net
Li Y B , Gong Z Y , Li H Y , Huang X Q , Kang H L , Bai G P , et al . 2025e . Robotic visual instruction // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, USA : IEEE: 12155 - 12165 [ DOI: 10.1109/CVPR52734.2025.01135 http://dx.doi.org/10.1109/CVPR52734.2025.01135 ]
Li Y L , Yan G , Macaluso A , Ji M Z Y , Zou X Y and Wang X L . 2025f . Integrating LMM planners and 3D skill policies for generalizable manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.18733.pdf https://arxiv.org/pdf/2501.18733.pdf
Li W , Zhang R S , Shao R , He J and Nie L Q . 2025g . CogVLA: cognition-aligned vision-language-action model via instruction-driven routing & sparsification //Advances in Neural Information Processing Systems 38. DiegoSan, USA: Neural Information Processing Systems Foundation [EB/OL]. [ 2025-11-17 ]. https://openreview.net/forum?id=Fg9HufTI0K https://openreview.net/forum?id=Fg9HufTI0K
Liang J , Huang W L , Xia F , Xu P , Hausman K , Ichter B , et al . 2023 . Code as policies: language model programs for embodied control // Proceedings of 2023 IEEE International Conference on Robotics and Automation . London, UK : IEEE: 9493 - 9500 [ DOI: 10.1109/ICRA48891.2023.10160591 http://dx.doi.org/10.1109/ICRA48891.2023.10160591 ]
Liang Z X , Mu Y , Ma H B , Tomizuka M , Ding M Y and Luo P . 2024 . SkillDiffuser: interpretable hierarchical planning via skill abstractions in diffusion-based task execution // Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Seattle, USA : IEEE: 16467 - 16476 [ DOI: 10.1109/CVPR52733.2024.01558 http://dx.doi.org/10.1109/CVPR52733.2024.01558 ]
Lin F Q , Nai R Q , Hu Y D , You J C , Zhao J M and Gao Y . 2025a . OneTwoVLA: a unified vision-language-action model with adaptive reasoning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.11917.pdf https://arxiv.org/pdf/2505.11917.pdf
Lin J Y , Taherin A , Akbari A , Akbari A , Lu L , Chen G Y , et al . 2025b . VOTE: vision-language-action optimization with trajectory ensemble voting [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.05116.pdf https://arxiv.org/pdf/2507.05116.pdf
Lin T , Li G , Zhong Y L , Zou Y W , Du Y X , Liu J T , et al . 2025c . Evo-0: vision-language-action model with implicit spatial understanding [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.00416.pdf https://arxiv.org/pdf/2507.00416.pdf
Liu B , Zhu Y F , Gao C K , Feng Y H , Liu Q , Zhu Y K , et al . 2023a . LIBERO: benchmarking knowledge transfer for lifelong robot learning // Proceedings of the 37th International Conference on Neural Information Processing Systems . New Orleans, USA : Curran Associates Inc.: #1939
Liu H T , Li C Y , Wu Q Y and Lee Y J . 2023b . Visual instruction tuning // Proceedings of the 37th International Conference on Neural Information Processing Systems . New Orleans, USA : Curran Associates Inc.: #1516
Liu J M , Chen H , An P J , Liu Z Y , Zhang R R , Gu C Y , et al . 2025a . HybridVLA: collaborative diffusion and autoregression in a unified vision-language-action model [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.10631.pdf https://arxiv.org/pdf/2503.10631.pdf
Liu J M , Liu M Z , Wang Z Y , An P J , Li X Q , Zhou K C , et al . 2024a . RoboMamba: efficient vision-language-action model for robotic reasoning and manipulation // Proceedings of the 38th International Conference on Neural Information Processing Systems . Vancouver, Canada : Curran Associates Inc.: #1266
Liu S L , Zeng Z Y , Ren T H , Li F , Zhang H , Yang J , et al . 2025b . Grounding DINO: marrying DINO with grounded pre-training for open-set object detection // Proceedings of the 18th European Conference on Computer Vision . Milan, Italy : Springer: 38 - 55 [ DOI: 10.1007/978-3-031-72970-6_3 http://dx.doi.org/10.1007/978-3-031-72970-6_3 ]
Liu S M , Wu L X , Li B G , Tan H K , Chen H Y , Wang Z Y , et al . 2025c . RDT-1B: a diffusion foundation model for bimanual manipulation // Proceedings of the 13th International Conference on Learning Representations . Singapore, Singapore : OpenReview.net
Liu S Y , Du J W , Xiang S C , Wang Z B and Luo D S . 2024b . ReLEP: a novel framework for real-world long-horizon embodied planning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2409.15658.pdf https://arxiv.org/pdf/2409.15658.pdf
Liu Z , Mao H Z , Wu C Y , Feichtenhofer C , Darrell T and Xie S N . 2022 . A ConvNet for the 2020s // Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition . New Orleans, USA : IEEE: 11966 - 11976 [ DOI: 10.1109/CVPR52688.2022.01167 http://dx.doi.org/10.1109/CVPR52688.2022.01167 ]
Liu Z Y , Gu Y C , Zheng S X , Fu Y W , Xue X Y and Jiang Y G . 2025d . TriVLA: a triple-system-based unified vision-language-action model with episodic world modeling for general robot control [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.01424.pdf https://arxiv.org/pdf/2507.01424.pdf
Lu G X , Guo W K , Zhang C B , Zhou Y H , Jiang H N , Gao Z F , et al . 2025 . VLA-RL: towards masterful and general robotic manipulation with scalable reinforcement learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.18719.pdf https://arxiv.org/pdf/2505.18719.pdf
Lu J S , Clark C , Lee S , Zhang Z C , Khosla S , Marten R , et al . 2024 . Unified-IO 2: scaling autoregressive multimodal models with vision, language, audio, and action // Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Seattle, USA : IEEE: 26429 - 26445 [ DOI: 10.1109/CVPR52733.2024.02497 http://dx.doi.org/10.1109/CVPR52733.2024.02497 ]
Luo H , Feng Y C , Zhang W P , Zheng S P , Wang Y , Yuan H Q , et al . 2025a . Being-H0: vision-language-action pretraining from large-scale human videos [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.15597.pdf https://arxiv.org/pdf/2507.15597.pdf
Lynch C , Wahid A , Tompson J , Ding T L , Betker J , Baruch R , et al . 2022 . Interactive language: talking to robots in real time [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2210.06407.pdf https://arxiv.org/pdf/2210.06407.pdf
Ma Y E , Song Z X , Zhuang Y Z , Hao J Y and King I . 2024 . A survey on vision-language-action models for embodied AI [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2405.14093.pdf https://arxiv.org/pdf/2405.14093.pdf
Ma Y J , Sodhani S , Jayaraman D , Bastani O , Kumar V and Zhang A . 2023 . VIP: towards universal visual reward and representation via value-implicit pre-training // Proceedings of the 11th International Conference on Learning Representations . Kigali, Rwanda : OpenReview.net
Makoviychuk V , Wawrzyniak L , Guo Y R , Lu M , Storey K , Macklin M , et al . 2021 . Isaac Gym: high performance GPU based physics simulation for robot learning //Proceedings of the 35th Conference on Neural Information Processing Systems. [s.l.]: Curran Associates, Inc.: 1 - 10
Mandlekar A , Nasiriany S , Wen B W , Akinola I , Narang Y , Fan L X , et al . 2023 . MimicGen: a data generation system for scalable robot learning using human demonstrations // Proceedings of the 7th Conference on Robot Learning . Atlanta, USA : PMLR: 1820 - 1864
Mandlekar A , Xu D F , Wong J , Nasiriany S , Wang C , Kulkarni R , et al . 2021 . What matters in learning from offline human demonstrations for robot manipulation // Proceedings of the 5th Conference on Robot Learning . London, UK : PMLR: 1678 - 1690
Mandlekar A , Zhu Y K , Garg A , Booher J , Spero M , Tung A , et al . 2018 . ROBOTURK: a crowdsourcing platform for robotic skill learning through imitation // Proceedings of the 2nd Conference on Robot Learning . Zürich, Switzerland : PMLR: 879 - 893
Mao W X , Zhong W H , Jiang Z , Fang D , Zhang Z Y , Lan Z H , et al . 2024 . RoboMatrix: a skill-centric hierarchical framework for scalable robot task planning and execution in open-world [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.00171.pdf https://arxiv.org/pdf/2412.00171.pdf
Mees O , Hermann L , Rosete-Beas E and Burgard W . 2022 . CALVIN: a benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks . IEEE Robotics and Automation Letters , 7 ( 3 ): 7327 - 7334 [ DOI: 10.1109/LRA.2022.3180108 http://dx.doi.org/10.1109/LRA.2022.3180108 ]
Mildenhall B , Srinivasan P P , Tancik M , Barron J T , Ramamoorthi R and Ng R . 2022 . NeRF: representing scenes as neural radiance fields for view synthesis . Communications of the ACM , 65 ( 1 ): 99 - 106 [ DOI: 10.1145/3503250 http://dx.doi.org/10.1145/3503250 ]
Monache S D , Moleman R and Özcan E . 2024 . Soundstorm, a collaborative ideation game for sound-driven design // Proceedings of the 19th International Audio Mostly Conference: Explorations in Sonic Cultures . Milan, Italy : ACM: 479 - 486 [ DOI: 10.1145/3678299.3678348 http://dx.doi.org/10.1145/3678299.3678348 ]
Mu Y , Zhang Q L , Hu M K , Wang W H , Ding M Y , Jin J , et al . 2023 . EmbodiedGPT: vision-language pre-training via embodied chain of thought // Proceedings of the 37th International Conference on Neural Information Processing Systems . New Orleans, USA : Curran Associates Inc.: #1090
Nair S , Rajeswaran A , Kumar V , Finn C and Gupta A . 2022 . R 3 M: a universal visual representation for robot manipulation// Proceedings of the 6th Conference on Robot Learning . Auckland, New Zealand: PMLR: 892 - 909
Nasiriany S , Maddukuri A , Zhang L C , Parikh A , Lo A , Joshi A , et al . 2024 . RoboCasa: large-scale simulation of everyday tasks for generalist robots [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2406.02523.pdf https://arxiv.org/pdf/2406.02523.pdf
Nematollahi I , DeMoss B , Chandra A L , Hawes N , Burgard W and Posner I . 2025 . LUMOS: language-conditioned imitation learning with world models // Proceedings of 2025 IEEE International Conference on Robotics and Automation . Atlanta, USA : IEEE: 8219 - 8225 [ DOI: 10.1109/ICRA55743.2025.11127988 http://dx.doi.org/10.1109/ICRA55743.2025.11127988 ]
Niu D T , Sharma Y , Biamby G , Quenum J , Bai Y T , Shi B F , et al . 2024 . LLARVA: vision-action instruction tuning enhances robot learning // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 3333 - 3355
O’Neill A , Rehman A , Maddukuri A , Gupta A , Padalkar A , Lee A , et al . 2024 . Open X-embodiment: robotic learning datasets and RT-X models // Proceedings of 2024 IEEE International Conference on Robotics and Automation . Yokohama, Japan : IEEE: 6892 - 6903 [ DOI: 10.1109/ICRA57147.2024.10611477 http://dx.doi.org/10.1109/ICRA57147.2024.10611477 ]
Oquab M , Darcet T , Moutakanni T , Vo H V , Szafraniec M , Khalidov V , et al . 2024 . DINOv2: learning robust visual features without supervision . Transactions on Machine Learning Research , 2024 : 1 - 17
Ouyang L , Wu J , Jiang X , Almeida D , Wainwright C L , Mishkin P , et al . 2022 . Training language models to follow instructions with human feedback // Proceedings of the 36th International Conference on Neural Information Processing Systems . New Orleans, USA : Curran Associates Inc.: #2011
Patratskiy M A , Kovalev A K and Panov A I . 2025 . Spatial traces: enhancing VLA models with spatial-temporal understanding [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2508.09032.pdf https://arxiv.org/pdf/2508.09032.pdf
Pertsch K , Stachowicz K , Ichter B , Driess D , Nair S , Vuong Q , et al . 2025 . FAST: efficient action tokenization for vision-language-action models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.09747.pdf https://arxiv.org/pdf/2501.09747.pdf
Pumacay W , Singh I , Duan J , Krishna R , Thomason J and Fox D . 2024 . The COLOSSEUM: a benchmark for evaluating generalization for robotic manipulation // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.133 http://dx.doi.org/10.15607/RSS.2024.XX.133 ]
Qi C R , Su H , Mo K C and Guibas L J . 2017a . PointNet: deep learning on point sets for 3D classification and segmentation // Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition . Honolulu, USA : IEEE: 77 - 85 [ DOI: 10.1109/CVPR.2017.16 http://dx.doi.org/10.1109/CVPR.2017.16 ]
Qi C R , Yi L , Su H and Guibas L J . 2017b . PointNet++: deep hierarchical feature learning on point sets in a metric space // Proceedings of the 31st International Conference on Neural Information Processing Systems . Long Beach, USA : Curran Associates Inc.: 5105 - 5114
Qi Z K , Dong R P , Zhang S C , Geng H R , Han C R , Ge Z , et al . 2025a . ShapeLLM: universal 3D object understanding for embodied interaction // Proceedings of the 18th European Conference on Computer Vision . Milan, Italy : Springer: 214 - 238 [ DOI: 10.1007/978-3-031-72775-7_13 http://dx.doi.org/10.1007/978-3-031-72775-7_13 ]
Qi Z K , Zhang W Y , Ding Y F , Dong R P , Yu X Q , Li J W , et al . 2025b . SOFAR: language-grounded orientation bridges spatial reasoning and object manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2502.13143.pdf https://arxiv.org/pdf/2502.13143.pdf
Qu D L , Song H M , Chen Q Z , Yao Y Q , Ye X Y , Ding Y , et al . 2025 . SpatialVLA: exploring spatial representations for visual-language-action model [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.15830v1.pdf https://arxiv.org/pdf/2501.15830v1.pdf
Radford A , Kim J W , Hallacy C , Ramesh A , Goh G , Agarwal S , et al . 2021 . Learning transferable visual models from natural language supervision //Proceedings of the 38th International Conference on Machine Learning. [s.l.]: PMLR: 8748 - 8763
Raffel C , Shazeer N , Roberts A , Lee K , Narang S , Matena M , et al . 2020 . Exploring the limits of transfer learning with a unified text-to-text transformer . The Journal of Machine Learning Research , 21 ( 1 ): #140
Ranasinghe K , Li X , Mata C , Park J and Ryoo M S . 2025 . Pixel motion as universal representation for robot control [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.07817.pdf https://arxiv.org/pdf/2505.07817.pdf
Rashid A , Sharma S , Kim C M , Kerr J , Chen L Y , Kanazawa A , et al . 2023 . Language embedded radiance fields for zero-shot task-oriented grasping // Proceedings of the 7th Conference on Robot Learning . Atlanta, USA : PMLR: 178 - 200
Reed S E , Zolna K , Parisotto E , Gómez Colmenarejo S , Novikov A , Barth-Maron G , et al . 2022 . A generalist agent . Transactions on Machine Learning Research , 2022
Reimers N and Gurevych I . 2019 . Sentence-BERT: sentence embeddings using siamese BERT-networks // Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing . Hong Kong, China : ACL: 3982 - 3992 [ DOI: 10.18653/v1/D19-1410 http://dx.doi.org/10.18653/v1/D19-1410 ]
Reuss M , Yagmurlu Ö E , Wenzel F and Lioutikov R . 2024 . Multimodal diffusion transformer: learning versatile behavior from multimodal goals // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.121 http://dx.doi.org/10.15607/RSS.2024.XX.121 ]
Röder F , Benad J , Eppe M and Banerjee P K . 2025 . Dynamics-aligned latent imagination in contextual world models for zero-shot generalization // Proceedings of the 39th Annual Conference on Neural Information Processing Systems . San Diego, USA : Curran Associates, Inc.
Ryoo M S , Piergiovanni A J , Arnab A , Dehghani M and Angelova A . 2021 . TokenLearner: adaptive space-time tokenization for videos //Proceedings of the 35th International Conference on Neural Information Processing Systems. [s.l.]: Curran Associates Inc .: #979
Sameni S , Kafle K , Tan H and Jenni S . 2024 . Building vision-language models on solid foundations with masked distillation // Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Seattle, USA : IEEE: 14216 - 14226 [ DOI: 10.1109/CVPR52733.2024.01348 http://dx.doi.org/10.1109/CVPR52733.2024.01348 ]
Sanh V , Debut L , Chaumond J and Wolf T . 2019 . DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/1910.01108v1.pdf https://arxiv.org/pdf/1910.01108v1.pdf
Sapkota R , Cao Y , Roumeliotis K I and Karkee M . 2025 . Vision-language-action models: concepts, progress, applications and challenges [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.04769v1.pdf https://arxiv.org/pdf/2505.04769v1.pdf
Sautenkov O , Yaqoot Y , Lykov A , Mustafa M A , Tadevosyan G , Akhmetkazy A , et al . 2025 . UAV-VLA: vision-language-action system for large scale aerial mission generation // Proceedings of the 20th ACM/IEEE International Conference on Human-Robot Interaction . Melbourne, Australia : IEEE: 1588 - 1592 [ DOI: 10.1109/HRI61500.2025.10974117 http://dx.doi.org/10.1109/HRI61500.2025.10974117 ]
Savva M , Kadian A , Maksymets O , Zhao Y L , Wijmans E , Jain B , et al . 2019 . Habitat: a platform for embodied AI research // Proceedings of 2019 IEEE/CVF International Conference on Computer Vision . Seoul, Korea(South) : IEEE: 9338 - 9346 [ DOI: 10.1109/ICCV.2019.00943 http://dx.doi.org/10.1109/ICCV.2019.00943 ]
Schulman J , Wolski F , Dhariwal P , Radford A and Klimov O . 2017 . Proximal policy optimization algorithms [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/1707.06347.pdf https://arxiv.org/pdf/1707.06347.pdf
Shafiullah N M M , Rai A , Etukuru H , Liu Y Q , Misra I , Chintala S , et al . 2023 . On bringing robots home [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2311.16098.pdf https://arxiv.org/pdf/2311.16098.pdf
Shao R , Li W , Zhang L S , Zhang R S , Liu Z Y , Chen R , et al . 2025 . Large VLM-based vision-language-action models for robotic manipulation: a survey [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2508.13073.pdf https://arxiv.org/pdf/2508.13073.pdf
Sharma P , Mohan L , Pinto L and Gupta A . 2018 . Multiple interactions made easy (MIME): large scale demonstrations data for imitation // Proceedings of the 2nd Annual Conference on Robot Learning . Zürich, Switzerland : PMLR: 906 - 915
Sharshar A , Khan L U , Ullah W and Guizani M . 2025 . Vision-language models for edge networks: a comprehensive survey . IEEE Internet of Things Journal , 12 ( 16 ): 32701 - 32724 [ DOI: 10.1109/JIOT.2025.3579032 http://dx.doi.org/10.1109/JIOT.2025.3579032 ]
Shentu Y D , Wu P , Rajeswaran A and Abbeel P . 2024 . From LLMs to actions: latent codes as bridges in hierarchical robot control // Proceedings of 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems . Abu Dhabi, United Arab Emirates : IEEE: 8539 - 8546 [ DOI: 10.1109/IROS58592.2024.10801683 http://dx.doi.org/10.1109/IROS58592.2024.10801683 ]
Shi L X , Ichter B , Equi M R , Ke L Y M , Pertsch K , Vuong Q , et al . 2025 . Hi robot: open-ended instruction following with hierarchical vision-language-action models // Proceedings of the 42nd International Conference on Machine Learning . Vancouver, Canada : PMLR: 54919 - 54933
Shorinwa O , Tucker J , Smith A , Swann A , Chen T , Firoozi R , et al . 2024 . Splat-MOVER : multi-stage, open-vocabulary robotic manipulation via editable gaussian splatting// Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 4748 - 4770
Shridhar M , Manuelli L and Fox D . 2021 . CLIPort: what and where pathways for robotic manipulation // Proceedings of the 5th Conference on Robot Learning . London, UK : PMLR: 894 - 906
Shridhar M , Thomason J , Gordon D , Bisk Y , Han W , Mottaghi R , et al . 2020 . ALFRED: a benchmark for interpreting grounded instructions for everyday tasks // Proceedings of 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Seattle, USA : IEEE: 10737 - 10746 [ DOI: 10.1109/CVPR42600.2020.01075 http://dx.doi.org/10.1109/CVPR42600.2020.01075 ]
Shukor M , Aubakirova D , Capuano F , Kooijmans P , Palma S , Zouitine A , et al . 2025 . SmolVLA: a vision-language-action model for affordable and efficient robotics [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.01844.pdf https://arxiv.org/pdf/2506.01844.pdf
Singh H , Das R J , Han M , Nakov P and Laptev I . 2024 . MALMM: multi-agent large language models for zero-shot robotics manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2411.17636.pdf https://arxiv.org/pdf/2411.17636.pdf
Singh I , Blukis V , Mousavian A , Goyal A , Xu D F , Tremblay J , et al . 2023 . ProgPrompt: generating situated robot task plans using large language models // Proceedings of 2023 IEEE International Conference on Robotics and Automation . London, UK : IEEE: 11523 - 11530 [ DOI: 10.1109/ICRA48891.2023.10161317 http://dx.doi.org/10.1109/ICRA48891.2023.10161317 ]
Song C H , Sadler B M , Wu J M , Chao W L , Washington C and Su Y . 2023 . LLM-planner: few-shot grounded planning for embodied agents with large language models // Proceedings of 2023 IEEE/CVF International Conference on Computer Vision . Paris, France : IEEE: 2986 - 2997 [ DOI: 10.1109/ICCV51070.2023.00280 http://dx.doi.org/10.1109/ICCV51070.2023.00280 ]
Song J M , Meng C L and Ermon S . 2021 . Denoising diffusion implicit models //Proceedings of the 9th International Conference on Learning Representations. [s.l.]: OpenReview .net
Song W X , Chen J Y , Ding P X , Zhao H , Zhao W , Zhong Z D , et al . 2025a . Accelerating vision-language-action model integrated with action chunking via parallel decoding [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.02310v1.pdf https://arxiv.org/pdf/2503.02310v1.pdf
Song W X , Chen J Y , Li W X , He X , Zhao H , Cui C , et al . 2025b . RationalVLA: a rational vision-language-action model with dual system [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.10826.pdf https://arxiv.org/pdf/2506.10826.pdf
Song Z R , Ouyang G X , Li M Z , Ji Y H , Wang C X , Xu Z X , et al . 2026 . ManipLVM-R1: reinforcement learning for reasoning in embodied manipulation with large vision-language models // Proceedings of the 40th AAAI Conference on Artificial Intelligence . Singapore, Singapore : AAAI Press: 18558 - 18566 [ DOI: 10.1609/aaai.v40i22.38922 http://dx.doi.org/10.1609/aaai.v40i22.38922 ]
Sun Q , Fang Y X , Wu L , Wang X L and Cao Y . 2023 . EVA-CLIP: improved training techniques for CLIP at scale [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2303.15389.pdf https://arxiv.org/pdf/2303.15389.pdf
Tan M X and Le Q V . 2019 . EfficientNet: rethinking model scaling for convolutional neural networks // Proceedings of the 36th International Conference on Machine Learning . Long Beach, USA : PMLR: 6105 - 6114
Tan S H , Dou K R , Zhao Y and Krähenbühl P . 2025a . Interactive post-training for vision-language-action models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.17016.pdf https://arxiv.org/pdf/2505.17016.pdf
Tan X D , Yang Y X , Ye P , Zheng J L , Bai B Z , Wang X Y , et al . 2025b . Think twice, act once: token-aware compression and action reuse for efficient inference in vision-language-action models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.21200.pdf https://arxiv.org/pdf/2505.21200.pdf
Touvron H , Lavril T , Izacard G , Martinet X , Lachaux M A , Lacroix T , et al . 2023 . LLaMA: open and efficient foundation language models [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf
van den Oord A , Vinyals O and Kavukcuoglu K . 2017 . Neural discrete representation learning // Proceedings of the 31st International Conference on Neural Information Processing Systems . Long Beach, USA : Curran Associates Inc.: 6309 - 6318
Van Vo T , Nguyen T Q , Nguyen K M , Nguyen D H M and Vu M N . 2025 . ReFineVLA: reasoning-aware teacher-guided transfer fine-tuning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.19080.pdf https://arxiv.org/pdf/2505.19080.pdf
Wang C , Xia F , Yu W H , Zhang T N , Zhang R H , Liu C K , et al . 2025a . Chain-of-modality: learning manipulation programs from multimodal human videos with vision-language-models // Proceedings of 2025 IEEE International Conference on Robotics and Automation (ICRA) . Atlanta, USA : IEEE: 6527 - 6535 [ DOI: 10.1109/ICRA55743.2025.11128270 http://dx.doi.org/10.1109/ICRA55743.2025.11128270 ]
Wang H Y , Xiong C Y , Wang R P and Chen X L . 2025b . BitVLA: 1-bit vision-language-action models for robotics manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.07530.pdf https://arxiv.org/pdf/2506.07530.pdf
Wang S S , Yu R C , Yuan Z H , Yu C , Gao F , Wang Y , et al . 2025c . Spec-VLA: speculative decoding for vision-language-action models with relaxed acceptance // Proceedings of 2025 Conference on Empirical Methods in Natural Language Processing . Suzhou, China : ACL: 26928 - 26940 [ DOI: 10.18653/v1/2025.emnlp-main.1367 http://dx.doi.org/10.18653/v1/2025.emnlp-main.1367 ]
Wang T W , Han C , Liang J C , Yang W H , Liu D F , Zhang L X , et al . 2025d . Exploring the adversarial vulnerabilities of vision-language-action models in robotics [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2411.13587.pdf https://arxiv.org/pdf/2411.13587.pdf
Wang W G , Yang Y and Pan Y H . 2025e . Visual knowledge in the big model era: retrospect and prospect . Frontiers of Information Technology and Electronic Engineering , 26 ( 1 ): 1 - 19 [ DOI: 10.1631/FITEE.2400250 http://dx.doi.org/10.1631/FITEE.2400250 ]
Wang Y H , Ding P X , Li L X , Cui C , Ge Z R , Tong X Y , et al . 2026 . VLA-adapter: an effective paradigm for tiny-scale vision-language-action model // Proceedings of the 40th AAAI Conference on Artificial Intelligence . Singapore, Singapore : AAAI Press: 18638 - 18646 [ DOI: 10.1609/aaai.v40i22.38931 http://dx.doi.org/10.1609/aaai.v40i22.38931 ]
Wang Y Q , Li X H , Wang W X , Zhang J B , Li Y Y , Chen Y T , et al . 2025f . Unified vision-language-action model [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.19850.pdf https://arxiv.org/pdf/2506.19850.pdf
Wang Y T , Zhu H Y , Liu M Y , Yang J G , Fang H S and He T . 2025g . VQ-VLA: improving vision-language-action models via scaling vector-quantized action tokenizers // Proceedings of 2025 IEEE/CVF International Conference on Computer Vision . Honolulu, USA : IEEE: 11089 - 11099
Wei J L , Yuan S S , Li P F , Hu Q D , Gan Z X and Ding W C . 2024 . OccLLaMA: an occupancy-language-action generative world model for autonomous driving [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2409.03272.pdf https://arxiv.org/pdf/2409.03272.pdf
Wen C , Lin X Y , So J , Chen K , Dou Q , Gao Y , et al . 2024a . Any-point trajectory modeling for policy learning // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.092 http://dx.doi.org/10.15607/RSS.2024.XX.092 ]
Wen J J , Zhu M J , Zhu Y C , Tang Z B , Li J M , Zhou Z Y , et al . 2025a . Diffusion-VLA: generalizable and interpretable robot foundation model via self-generated reasoning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.03293.pdf https://arxiv.org/pdf/2412.03293.pdf
Wen J J , Zhu Y C , Li J M , Tang Z B , Shen C M and Feng F F . 2025b . DexVLA: vision-language model with plug-in diffusion expert for general robot control [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2502.05855.pdf https://arxiv.org/pdf/2502.05855.pdf
Wen J J , Zhu Y C , Li J M , Zhu M J , Wu K , Xu Z Y , et al . 2024b . TinyVLA: towards fast, data-efficient vision-language-action models for robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2409.12514.pdf https://arxiv.org/pdf/2409.12514.pdf
Wu D , Fan J X , Zang J Z , Wang G B , Yin W , Li W H , et al . 2025a . Reinforced reasoning for embodied planning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.22050.pdf https://arxiv.org/pdf/2505.22050.pdf
Wu H J , Zhang H , Zhang P S , Wang J and Wang C . 2025b . HiBerNAC: hierarchical brain-emulated robotic neural agent collective for disentangling complex manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.08296.pdf https://arxiv.org/pdf/2506.08296.pdf
Wu H T , Jing Y , Cheang C , Chen G Z , Xu J F , Li X H , et al . 2024a . Unleashing large-scale video generative pre-training for visual robot manipulation // Proceedings of the 12th International Conference on Learning Representations . Vienna, Austria : OpenReview.net
Wu K , Hou C K , Liu J M , Che Z P , Ju X Z , Yang Z Q , et al . 2024b . RoboMIND: benchmark on multi-embodiment intelligence normative data for robot manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.13877.pdf https://arxiv.org/pdf/2412.13877.pdf
Wu Z Y , Zhou Y H , Xu X W , Wang Z W and Yan H B . 2025c . MoManipVLA: transferring vision-language-action models for general mobile manipulation // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, USA : IEEE: 1714 - 1723 [ DOI: 10.1109/CVPR52734.2025.00167 http://dx.doi.org/10.1109/CVPR52734.2025.00167 ]
Xia F , Shen W B , Li C S , Kasimbeg P , Tchapmi M E , Toshev A , et al . 2020 . Interactive gibson benchmark: a benchmark for interactive navigation in cluttered environments . IEEE Robotics and Automation Letters , 5 ( 2 ): 713 - 720 [ DOI: 10.1109/LRA.2020.2965078 http://dx.doi.org/10.1109/LRA.2020.2965078 ]
Xiang F B , Qin Y Z , Mo K C , Xia Y K , Zhu H , Liu F C , et al . 2020 . SAPIEN: a SimulAted part-based interactive ENvironment // Proceedings of 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Seattle, USA : IEEE: 11094 - 11104 [ DOI: 10.1109/CVPR42600.2020.01111 http://dx.doi.org/10.1109/CVPR42600.2020.01111 ]
Xu C , Li Q Y , Luo J L and Levine S . 2024a . RLDG: robotic generalist policy distillation via reinforcement learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.09858.pdf https://arxiv.org/pdf/2412.09858.pdf
Xu R T , Zhang J , Guo M H , Wen Y P , Yang H T , Lin M , et al . 2025 . A 0 : an affordance-aware hierarchical model for general robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2504.12636.pdf https://arxiv.org/pdf/2504.12636.pdf
Xu Z Y , Wu K , Wen J J , Li J M , Liu N , Che Z P , et al . 2024b . A survey on robotics with foundation models: toward embodied AI [EB/OL]. [ 2024-02-05 ]. https://arxiv.org/pdf/2402.02385.pdf https://arxiv.org/pdf/2402.02385.pdf
Yan F , Liu F F , Zheng L M , Zhong Y F , Huang Y Y , Guan Z C , et al . 2024 . RoboMM: all-in-one multimodal large model for robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.07215v1.pdf https://arxiv.org/pdf/2412.07215v1.pdf
Yang J , Glossop C , Bhorkar A , Shah D , Vuong Q , Finn C , et al . 2024a . Pushing the limits of cross-embodiment learning for manipulation and navigation // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/rss.2024.xx.093 http://dx.doi.org/10.15607/rss.2024.xx.093 ]
Yang L H , Kang B Y , Huang Z L , Xu X G , Feng J S and Zhao H S . 2024b . Depth anything: unleashing the power of large-scale unlabeled data // Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Seattle, USA : IEEE: 10371 - 10381 [ DOI: 10.1109/CVPR52733.2024.00987 http://dx.doi.org/10.1109/CVPR52733.2024.00987 ]
Yang R H , Yu Q X , Wu Y C , Yan R , Li B R , Cheng A C , et al . 2025a . EgoVLA: learning vision-language-action models from egocentric human videos [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.12440.pdf https://arxiv.org/pdf/2507.12440.pdf
Yang Y , Sun J X , Kou S Q , Wang Y H and Deng Z J . 2025b . LoHoVLA: a unified vision-language-action model for long-horizon embodied tasks [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.00411.pdf https://arxiv.org/pdf/2506.00411.pdf
Yang Y , Zhuang Y T and Pan Y H . 2022 . The review of visual knowledge: a new pivot for cross-media intelligence evolution . Journal of Image and Graphics , 27 ( 9 ): 2574 - 2588
杨易 , 庄越挺 , 潘云鹤 . 2022 . 视觉知识: 跨媒体智能进化的新支点 . 中国图象图形学报 , 27 ( 9 ): 2574 - 2588 [ DOI: 10.11834/jig.211264 http://dx.doi.org/10.11834/jig.211264 ]
Yang Z , Chen Y , Wang J K , Manivasagam S , Ma W C , Yang A J , et al . 2023 . UniSim: a neural closed-loop sensor simulator // Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Vancouver, Canada : IEEE: 1389 - 1399 [ DOI: 10.1109/CVPR52729.2023.00140 http://dx.doi.org/10.1109/CVPR52729.2023.00140 ]
Ye S , Jang J , Jeon B , Joo S J , Yang J W , Peng B L , et al . 2025 . Latent action pretraining from videos // Proceedings of the 13th International Conference on Learning Representations . Singapore, Singapore : OpenReview.net
Yu F , Chen H F , Wang X , Xian W Q , Chen Y Y , Liu F C , et al . 2020 . BDD100K: a diverse driving dataset for heterogeneous multitask learning // Proceedings of 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Seattle, USA : IEEE: 2633 - 2642 [ DOI: 10.1109/CVPR42600.2020.00271 http://dx.doi.org/10.1109/CVPR42600.2020.00271 ]
Yu J W , Liu H R , Yu Q J , Ren J J , Hao C , Ding H T , et al . 2025 . ForceVLA: enhancing VLA models with a force-aware MoE for contact-rich manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.22159.pdf https://arxiv.org/pdf/2505.22159.pdf
Yu S , Sohn K , Kim S and Shin J . 2023 . Video probabilistic diffusion models in projected latent space // Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Vancouver, Canada : IEEE: 18456 - 18466 [ DOI: 10.1109/CVPR52729.2023.01770 http://dx.doi.org/10.1109/CVPR52729.2023.01770 ]
Yu T H , Quillen D , He Z P , Julian R , Hausman K , Finn C , et al . 2019 . Meta-world: a benchmark and evaluation for multi-task and meta reinforcement learning // Proceedings of the 3rd Annual Conference on Robot Learning . Osaka, Japan : PMLR: 1094 - 1100
Yuan C B , Wen C , Zhang T and Gao Y . 2024a . General flow as foundation affordance for scalable robot learning // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 1541 - 1566
Yuan W T , Duan J F , Blukis V , Pumacay W , Krishna R , Murali A , et al . 2024b . RoboPoint: a vision-language model for spatial affordance prediction in robotics // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 4005 - 4020
Yue Y , Wang Y L , Kang B Y , Han Y Z , Wang S Z , Song S J , et al . 2024 . DeeR-VLA: dynamic inference of multimodal large language models for efficient robot execution // Proceedings of the 38th International Conference on Neural Information Processing Systems . Vancouver, Canada : Curran Associates Inc.: #1803
Zawalski M , Chen W , Pertsch K , Mees O , Finn C and Levine S . 2024 . Robotic control via embodied chain-of-thought reasoning // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 3157 - 3181
Ze Y J , Zhang G , Zhang K N , Hu C Y , Wang M H and Xu H Z . 2024 . 3D diffusion policy: generalizable visuomotor policy learning via simple 3D representations // Proceedings of Robotics: Science and Systems XX . Delft, the Netherlands : MIT Press [ DOI: 10.15607/RSS.2024.XX.067 http://dx.doi.org/10.15607/RSS.2024.XX.067 ]
Zhai X H , Mustafa B , Kolesnikov A and Beyer L . 2023 . Sigmoid loss for language image pre-training // Proceedings of 2023 IEEE/CVF International Conference on Computer Vision . Paris, France : IEEE: 11941 - 11952 [ DOI: 10.1109/ICCV51070.2023.01100 http://dx.doi.org/10.1109/ICCV51070.2023.01100 ]
Zhang B R , Zhang Y H , Ji J M , Lei Y S , Dai J , Chen Y P , et al . 2025a . SafeVLA: towards safety alignment of vision-language-action model via constrained learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.03480v2.pdf https://arxiv.org/pdf/2503.03480v2.pdf
Zhang C F , Hao P , Cao X G , Hao X S , Cui S W and Wang S . 2025b . VTLA: vision-tactile-language-action model with preference learning for insertion manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.09577.pdf https://arxiv.org/pdf/2505.09577.pdf
Zhang D P , Sun J , Hu C H , Wu X Y , Yuan Z L , Zhou R , et al . 2025c . Pure vision language action (VLA) models: a comprehensive survey [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2509.19012.pdf https://arxiv.org/pdf/2509.19012.pdf
Zhang F and Gienger M . 2024 . Affordance-based Robot Manipulation with Flow Matching [EB/OL]. [ 2025-10-29 ]. https://arxiv.org/pdf/2505.21851.pdf https://arxiv.org/pdf/2505.21851.pdf
Zhang H C , Yu H N , Zhao L , Choi A , Bai Q X , Yang B , et al . 2025d . SLIM: sim-to-real legged instructive manipulation via long-horizon visuomotor learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2501.09905.pdf https://arxiv.org/pdf/2501.09905.pdf
Zhang H Y , Zhuang Z F , Zhao H , Ding P X , Lu H C and Wang D L . 2025e . ReinboT: amplifying robot visual-language manipulation with reinforcement learning // Proceedings of the 42nd International Conference on Machine Learning . Vancouver, Canada : PMLR: 77254 - 77271
Zhang J H , Chen Y R , Xu Y M , Huang Z , Zhou Y P , Yuan Y J , et al . 2025f . 4D-VLA: spatiotemporal vision-language-action pretraining with cross-scene calibration [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2506.22242.pdf https://arxiv.org/pdf/2506.22242.pdf
Zhang J H , Luo Y S , Anwar A , Sontakke S A , Lim J J , Thomason J , et al . 2025g . ReWiND: language-guided rewards teach robot policies without new demonstrations [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.10911.pdf https://arxiv.org/pdf/2505.10911.pdf
Zhang J K , Guo Y J , Chen X Y , Wang Y J , Hu Y C , Shi C M , et al . 2024a . HiRT: enhancing robotic control with hierarchical robot transformers // Proceedings of the 8th Conference on Robot Learning . Munich, Germany : PMLR: 933 - 946
Zhang J K , Guo Y J , Hu Y C , Chen X Y , Zhu X and Chen J Y . 2025h . UP-VLA: a unified understanding and prediction model for embodied agent // Proceedings of the 42nd International Conference on Machine Learning . Vancouver, Canada : PMLR: 74911 - 74922
Zhang J Y , Xu W Q , Yu Z J , Xie P F , Tang T T and Lu C W . 2025i . DexTOG: learning task-oriented dexterous grasp with language condition . IEEE Robotics and Automation Letters , 10 ( 2 ): 995 - 1002 [ DOI: 10.1109/LRA.2024.3518116 http://dx.doi.org/10.1109/LRA.2024.3518116 ]
Zhang J Z , Wang K Y , Wang S A , Li M H , Liu H R , Wei S L , et al . 2024b . Uni-NaVid: a video-based vision-language-action model for unifying embodied navigation tasks [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.06224v1.pdf https://arxiv.org/pdf/2412.06224v1.pdf
Zhang K D , Xu R T , Ren P Z , Lin J F , Wu H F , Lin L , et al . 2025j . RoBridge: a hierarchical architecture bridging cognition and execution for general robotic manipulation [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.01709.pdf https://arxiv.org/pdf/2505.01709.pdf
Zhang R Y , Dong M H , Zhang Y , Heng L , Chi X W , Dai G L , et al . 2026 . MoLe-VLA: dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation // Proceedings of the 40th AAAI Conference on Artificial Intelligence . Singapore, Singapore : AAAI Press: 18764 - 18772 [ DOI: 10.1609/aaai.v40i22.38945 http://dx.doi.org/10.1609/aaai.v40i22.38945 ]
Zhang S D , Xu Z , Liu P J , Yu X P , Li Y , Gao Q H , et al . 2024c . VLABench: a large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2412.18194.pdf https://arxiv.org/pdf/2412.18194.pdf
Zhang W Q , Wang M N , Liu G G , Xu H X , Jiang Y W , Shen Y L , et al . 2025k . Embodied-reasoner: synergizing visual search, reasoning, and action for embodied interactive tasks [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2503.21696.pdf https://arxiv.org/pdf/2503.21696.pdf
Zhang X , Zhang D , Li S M , Zhou Y Q and Qiu X P . 2024d . SpeechTokenizer: unified speech tokenizer for speech language models // Proceedings of the 12th International Conference on Learning Representations . Vienna, Austria : OpenReview.net
Zhang Y J , Zhang R Q , Zhou H J , Qi J , Yu Z F and Huang T J . 2025 . Research status and development trends of vision foundation models . Journal of Image and Graphics , 30 ( 1 ): 1 - 24
张燚钧 , 张润清 , 周华健 , 齐骥 , 余肇飞 , 黄铁军 . 2025 . 视觉基础模型研究现状与发展趋势 . 中国图象图形学报 , 30 ( 1 ): 1 - 24 [ DOI: 10.11834/jig.230911 http://dx.doi.org/10.11834/jig.230911 ]
Zhao Q Q , Lu Y , Kim M J , Fu Z P , Zhang Z Y , Wu Y C , et al . 2025a . CoT-VLA: visual chain-of-thought reasoning for vision-language-action models // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, USA : IEEE: 1702 - 1713 [ DOI: 10.1109/CVPR52734.2025.00166 http://dx.doi.org/10.1109/CVPR52734.2025.00166 ]
Zhao T Z , Kumar V , Levine S and Finn C . 2023 . Learning fine-grained bimanual manipulation with low-cost hardware // Proceedings of Robotics: Science and Systems XIX . Daegu, Korea(South) : MIT Press [ DOI: 10.15607/RSS.2023.XIX.016 http://dx.doi.org/10.15607/RSS.2023.XIX.016 ]
Zhao W , Ding P X , Zhang M , Gong Z F , Bai S H , Zhao H , et al . 2025b . VLAS: vision-language-action model with speech instructions for customized robot manipulation // Proceedings of the 13th International Conference on Learning Representations . Singapore, Singapore : OpenReview.net
Zhao W , Li G S , Gong Z F , Ding P X , Zhao H and Wang D L . 2025c . Unveiling the potential of vision-language-action models with open-ended multimodal instructions [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.11214.pdf https://arxiv.org/pdf/2505.11214.pdf
Zhen H Y , Qiu X W , Chen P H , Yang J C , Yan X , Du Y L , et al . 2024 . 3D-VLA: a 3D vision-language-action generative world model // Proceedings of the 41st International Conference on Machine Learning . Vienna, Austria : PMLR: 61229 - 61245
Zheng J L , Li J X , Liu D X , Zheng Y N , Wang Z H , Ou Z H , et al . 2025a . Universal actions for enhanced embodied foundation models // Proceedings of 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, USA : IEEE: 22508 - 22519 [ DOI: 10.1109/CVPR52734.2025.02096 http://dx.doi.org/10.1109/CVPR52734.2025.02096 ]
Zheng R J , Liang Y Y , Huang S Y , Gao J F , Daumé III H , Kolobov A , et al . 2025b . TraceVLA: visual trace prompting enhances spatial-temporal awareness for generalist robotic policies // Proceedings of the 13th International Conference on Learning Representations . Singapore, Singapore : OpenReview.net
Zhi P Y , Zhang Z Y , Zhao Y , Han M Z , Zhang Z Y , Li Z T , et al . 2025 . Closed-loop open-vocabulary mobile manipulation with GPT-4V // Proceedings of 2025 IEEE International Conference on Robotics and Automation . Atlanta, USA : IEEE: 4761 - 4767 [ DOI: 10.1109/ICRA55743.2025.11127975 http://dx.doi.org/10.1109/ICRA55743.2025.11127975 ]
Zhong Y F , Bai F S , Cai S F , Huang X C , Chen Z , Zhang X W , et al . 2025a . A survey on vision-language-action models: an action tokenization perspective [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2507.01925.pdf https://arxiv.org/pdf/2507.01925.pdf
Zhong Z D , Yan H D , Li J F , Liu X C , Gong X , Song W X , et al . 2025b . FlowVLA: thinking in motion with a visual chain of thought [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2508.18269.pdf https://arxiv.org/pdf/2508.18269.pdf
Zhou J S , Wang J S , Ma B R , Liu Y S , Huang T J and Wang X L . 2024 . Uni 3 D : exploring unified 3D representation at scale //Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria : OpenReview.net
Zhou X C , Han X Y , Yang F , Ma Y P , Tresp V and Knoll A . 2026 . OpenDriveVLA: towards end-to-end autonomous driving with large vision language action model // Proceedings of 2026 AAAI Conference on Artificial Intelligence . Singapore, Singapore : AAAI Press: 13782 - 13790 [ DOI: 10.1609/aaai.v40i16.38386 http://dx.doi.org/10.1609/aaai.v40i16.38386 ]
Zhou X Y , Tie G Y , Zhang G W , Wang H C , Zhou P and Sun L C . 2025a . BadVLA: towards backdoor attacks on vision-language-action models via objective-decoupled optimization // Proceedings of the 39th Annual Conference on Neural Information Processing Systems . San Diego, USA : Curran Associates, Inc.
Zhou Z Y , Zhu Y C , Wen J J , Shen C M and Xu Y . 2025b . Vision-language-action model with open-world embodied reasoning from pretrained knowledge [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2505.21906v1.pdf https://arxiv.org/pdf/2505.21906v1.pdf
Zhou Z Y , Zhu Y C , Zhu M J , Wen J J , Liu N , Xu Z Y , et al . 2025c . ChatVLA: unified multimodal understanding and robot control with vision-language-action model // Proceedings of 2025 Conference on Empirical Methods in Natural Language Processing . Suzhou, China : ACL: 5377 - 5395 [ DOI: 10.18653/v1/2025.emnlp-main.273 http://dx.doi.org/10.18653/v1/2025.emnlp-main.273 ]
Zhu Y K , Wong J , Mandlekar A and Martín-Martín R . 2020 . Robosuite: a modular simulation framework and benchmark for robot learning [EB/OL]. [ 2025-11-17 ]. https://arxiv.org/pdf/2009.12293v1.pdf https://arxiv.org/pdf/2009.12293v1.pdf
Zitkovich B , Yu T H , Xu S C , Xu P , Xiao T , Xia F , et al . 2023 . RT-2: vision-language-action models transfer web knowledge to robotic control // Proceedings of the 7th Conference on Robot Learning . Atlanta, USA : PMLR: 2165 - 2183
相关作者
相关机构
京公网安备11010802024621