Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Open access Sep 2026

AgenticRL: Agentic Reinforcement Learning with Self-Refinement for Complex UAV Navigation

Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but still relies heavily on time consuming manual reward design and fine tuning. Existing automated reward generation and refinement methods reduce this effort, yet often lack task-level behavioral diagnosis for directing subsequen...

Roohan Ahmed Khan, Yasheerah Yaqoot, Amir Atef Habel et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation

Contact-rich manipulation demands both high-level semantic reasoning and the safe regulation of high-frequency contact dynamics. While Vision-Language-Action (VLA) models provide unprecedented semantic generalization, their low-rate outputs lack the reliability required for direct plant authority in force-sensitive tas...

Haofan Cao, Zhaoyang Li, Zhichao You · 0 citations
#artificial intelligence Preprint Open access Sep 2026

REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception

Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and robustness to extreme lighting. However, existing learning-based approaches for event processing are typically confined to narrow, task-specific silos and lack the ability to generaliz...

Vincenzo Polizzi, David B. Lindell, Jonathan Kelly · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Evolving Skill Modules under a Fixed Planner: Versioning, Rollback, and Runtime Governance for Long-Lived Robot Systems

Robots deployed for long periods keep improving their skills, and each update changes a released system. We treat this as a software-lifecycle problem: a fixed decision layer dispatches versioned skill modules and a runtime layer was built to screen each action. On six robosuite tasks we report three negative results a...

Xue Qin, Simin Luan, Cong Yang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation

While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. This leaves a significant gap in real-time, sound-centric manipulation where fleeting environmental acoustics provide critical state verifica...

Chang Nie, Tianchen Deng, Guangming Wang et al. · 0 citations

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

HerMES is proposed, a holistic risk-aware end-to-end multimodal driving framework that explicitly incorporates long-tail semantic knowledge into trajectory planning and demonstrates consistent improvements over representative recent baselines in overall planning performance and across diverse safety-critical scenarios.

Wei-Zhe Tang, Jun-Wei You, Jia-Xi Liu et al. · 5 citations
#artificial intelligence Preprint Sep 2025

Benchmarking Autonomous Driving Planners Across Leaderboards: A Unified CARLA-Based Evaluation

This study presents a comparative case study of representative motion planning methods drawn from major benchmark ecosystems, including CARLA, nuPlan, and the Waymo Open Dataset to highlight the strengths and weaknesses of current approaches.

Merve Atasever, Alfredo Reina Corona, Zhuo-Chen Liu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. Choosing well requires knowing how much the robot's sensors reveal about the cause and how reliable the robot's own diagnosis is. We build a simulated benchmark...

Eshika Pathak, L. Krishna · 0 citations
#artificial intelligence Preprint Open access Sep 2026

ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction

Manipulating objects requires understanding not only their motion, but also the physical properties that determine it. For articulated objects, these include inertia, friction, and mechanisms such as springs or door closers, whose effects can vary with configuration and velocity. Such properties are not directly observ...

Tim Engelbracht, Ren\'e Zurbr\"ugg, Mayank Mittal et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies

Vision-language-action (VLA) policies solve the same manipulation task through different action interfaces, but task success alone does not establish whether their physical executions agree. We study cross-policy end-effector geometry in 15,000 closed-loop LIBERO rollouts from four policies. The primary clean-condition...

Xing-Yu Lin, Zhuang Li, Zhong-Run Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated t...

Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation

PA-RL, a reinforcement-learning framework that uses artificial potential fields as the action representation, is proposed and is the only method to reach a 100% evaluation success rate within the allotted training time, while the best baseline reaches 92.6%.

Xin-Yu Liu, Gokhan Solak, Arash Ajoudani · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.