Skip to content

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

MILER, an end-to-end policy framework with zero-shot sim-to-real transfer of both perception and control, is presented and extensively evaluated on a diverse test track comprising numerous challenges.

Abstract

Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird's-eye-view representation consistent with that of the MLR simulator. The actions generated by the policy network are not applied directly to the real vehicle. Instead, we employ a trajectory-alignment strategy that enables zero-shot sim-to-real transfer of both perception and control. We extensively evaluate the proposed framework on a diverse test track comprising numerous challenges, including various obstacles, hairpin curves, velocities of up to 33.6 km/h, and off-road sections. In total, we drove 17.3 km with two different vehicles on a 3.0 km test track without human intervention, thereby demonstrating the effectiveness of our approach. Furthermore, the entire software stack runs on a Jetson AGX Orin.

View source

Similar papers

Preprint Sep 2026

Sim-to-Real Aware End-to-End Learning Environment for Micromobility

While end-to-end autonomous driving systems show promise, their application to micromobility vehicles is hindered by simulators failing to capture specific kinematics, such as differential drives and omni-wheels. This paper pro- poses a sim-to-real-aware, vehicle-specific end-to-end learning environment for the WHILL M...

Shouma Amano, Takuya Azumi · 0 citations
Open access Sep 2026

FROA-Drive: Failure-Routed Offline Adaptation for Lightweight Vision–Language–Action Autonomous Driving

Vision–Language–Action (VLA) models have shown strong potential for end-to-end autonomous driving, yet their post-training commonly relies on expensive simulator interaction or global policy updates. For an already competent pretrained policy, targeted correction is substantially more economical than repeated simulator...

Yun-Han Xu, Ao Xu · 0 citations
Preprint Sep 2026

V2X-WAM: A Cooperative World Action Model for End-to-End Autonomous Driving

Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of th...

Jun-Wei You, Wei-Zhe Tang, Can Wang et al. · 0 citations
Open access Aug 2026

PPO-Based Sim-to-Real Maples’ Navigation for TurtleBot3 Mobile Robots in Unknown and Dynamic Indoor Environments

The results of this study show that a PPO-based policy trained only in simulation and fine-tuned only on a small number of real-world tasks can compete with the classical and alternative DRL baselines in unknown and dynamic environments.

Nabeel Muhamed, Khaleel Ali Khudhur · 0 citations
Conference Open access Sep 2026

Self-Improving Autonomous Vehicles via Real-World Reinforcement Learning

End-to-end autonomous driving systems have demonstrated advantages over traditional modular systems. Despite this progress, these end-to-end systems still struggle to be deployed in real-world driving environments, as they inevitably encounter undertrained scenarios in which autonomous vehicles may take unsafe actions....

Daehyeok Kwon, Seung-Woo Seo, Sang-Hyun Lee · 0 citations
#artificial intelligence Preprint Sep 2026

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

This work presents Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations.

Wen-Kang Qin, Yu-Kun Zhou, Noah Shen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.