Skip to content
Preprint

Autonomously Acquiring Robot Manipulation Skills with Language-Driven Quality-Diversity

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

This paper proposes an approach designed to output diverse motion primitive archives by autonomously leveraging quality-diversity algorithms, only requiring a free-form description of the task in common language, and proposes an autonomous exploration mechanism able to reliably output sets of functionals covering the fitness and behavior descriptor (BD) space.

Abstract

Quality-diversity (QD) algorithms have been gaining traction in robot learning, where diverse motion primitive libraries allow robots to adapt zero-shot to constraints at deployment time. However, such methods typically require expert designers to write the success condition, fitness and diversity metrics, and this strongly limits the robot's autonomy. On the other hand, existing LLM-based reward-shaping techniques allow robots to learn autonomously but only output single high-performing solutions, limiting the robot's adaptability. In this paper, we propose an approach designed to output diverse motion primitive archives by autonomously leveraging quality-diversity algorithms, only requiring a free-form description of the task in common language. To address the difficulty of designing relevant fitness and diversity metrics, we propose an autonomous exploration mechanism able to reliably output sets of functionals covering the fitness and behavior descriptor (BD) space. First, we pose policy exploration as a functional design problem, where the functional spaces are lower-dimensional than the full BD and fitness spaces, and propose an LLM-based exploration scheme to sample from these low-dimensional spaces without any task-specific prompts, fine-tuning or expert intervention. We adapt a multi-BD variant of the MAP-Elites success (MES) algorithm, designed to leverage the heterogeneous BD samples. Finally, through experiments based on the genesis simulator, we show that our method effectively generates archives of diverse motion primitives, outperforming classical QD algorithms with inferred and hand-written parametrizations on a set of $4$ robotic manipulation tasks.

View source

Similar papers

Preprint Aug 2026

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

This work uses Sample-based Model Predictive Control entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets and validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies.

Martin Schuck, Maks Sorokin, S. Manni et al. · 1 citation
Preprint Aug 2026

Adaptation of Generalist Robot Policies with Minimal Data

A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful be...

Shreyas Kowshik, Sreyas Venkataraman, L. Wang et al. · 1 citation
Preprint Sep 2026

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents'broader engineerin...

Hai-Tong Ma, Chen-Xiao Gao, Ru-Shi Qiang et al. · 0 citations
#large language models Open access Aug 2026

Physics filtering favors the generalization of robot learning

It is shown that robots can generalize effectively under dynamics uncertainties even with limited training data by leveraging a feedback mechanism, namely PhyFilter, that corrects learning outputs with physics-filtered learning residuals, and shows that physics-filtered feedback can serve as a powerful alternative to m...

Jin-Dou Jia, Shixu Han, Meng Wang et al. · 2 citations
#small language model Preprint Sep 2026

DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation

Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM...

Vincenzo Pomponi, Rocco Felici, Paolo Franceschi et al. · 0 citations
Preprint Sep 2026

SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation

Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement...

Saksham Singh, Zhe-Yuan Hu, Max Sobol Mark et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.