This thesis investigates Active Learning from Demonstrations (Active LfD) as a principled approach to reduce distribution shift while accounting for human factors in realistic human-in-the-loop settings and develops an Active LfD algorithm that operates with offline demonstrations and guides the resulting demonstration distribution toward a more balanced one, or more generally toward any specified target distribution, so as to improve robot learning.
Abstract
Learning from Demonstrations (LfD) enables robots to acquire new skills by imitating a domain expert, most commonly a human. In practice, however, human demonstrations are typically available only under a limited budget and rarely cover the full range of situations a robot may encounter. This limitation often induces a distribution shift between the states represented in offline demonstrations and those visited by the robot during online exploration, which can substantially degrade performance. Interactive Imitation Learning (IIL) mitigates this shift by keeping the human teachers in the learning loop, allowing them to provide online input such as feedback and corrections. Yet these benefits come at the cost of sustained human supervision and can be undermined by noise, inconsistency, and errors in human teaching.
This thesis investigates Active Learning from Demonstrations (Active LfD) as a principled approach to reduce distribution shift while accounting for human factors in realistic human-in-the-loop settings. In Active LfD, the robot learner does not passively consume demonstrations. Instead, it optimizes its query decisions to selectively request demonstrations from a human teacher. By determining when and what to query, the learner aims to reduce reliance on continuous monitoring, focus human effort where it is most beneficial, and limit the impact of suboptimal human teaching decisions.
The thesis is organized into three parts. In Part I, I lay the foundation for how a robot learner can actively shape human teaching by revisiting conventional offline LfD through the lens of demonstration distributions. After identifying biased human teaching strategies under unguided conditions, I develop an Active LfD algorithm that operates with offline demonstrations and guides the resulting demonstration distribution toward a more balanced one, or more generally toward any specified target distribution, so as to improve robot learning. Importantly, however, a ``balanced'' distribution (e.g., a uniform distribution used as a default) is not necessarily optimal for learning. The most beneficial demonstration distribution depends on the robot’s evolving policy and its online exploration, and is therefore difficult to specify a priori. To address this coupling, I present a second Active LfD algorithm that leverages online demonstrations, iteratively deciding both the timing and content of queries as learning progresses in order to optimize the demonstration distribution for the learner.
In Part II, I explore how Active LfD extends to more realistic and complex scenarios of learning from human teachers. I first extend the paradigm to transfer learning for complex manipulation tasks, spanning a range of policy transfer scenarios. I then consider the practical case of non-expert human teachers, who may provide imperfect demonstrations relative to the task objective. To accommodate imperfect teaching, I develop an Active LfD framework that optimizes the query sequence by jointly accounting for both the expected improvement in the robot’s policy induced by a query and the teacher’s capability to provide informative guidance in the queried region.
Part III closes the loop of human-robot interactive learning by examining how robot query design influences human teaching beyond user experience alone. I leverage Curriculum Learning to design an Active LfD algorithm that benefits both robot learning and human teaching, encouraging a reciprocal loop between the robot learner and the human teacher.
A framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation is presented, extending the widely used Gaussian Mixture Model and Gaussian Mixture Regression approach for learning from demonstration.
Alperen Kenan, Paul A. Bremner, Manuel Giuliani· 1 citation
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models'performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models. Webpage: https://roboreel.github.io
This study investigates the combination of a state-of-the-art reinforcement learning (RL) algorithm with human demonstrations to learn how to open a door with minimal task-specific engineering on an articulated soft robot arm and shows that combining LfD with RL results in both better performance and more robust behaviors.
Laurenz Elstner, Erik Kyrkjebø, M. Stoelen· Frontiers in Robotics and AI· 0 citations
Human-robot collaboration describes the process of humans and autonomous agents working together to accomplish common goals. This process is facilitated best when robot policies, or behaviors in different situations, are made transparent to humans. Demonstration-based explanations have been a focus of human-robot collaboration research, and the field has frequently drawn upon literature from education to improve how humans are taught robot policies. However, no single teaching method has been proven effective across domains, difficulties, learners, and other variables; the question of how humans can most effectively be taught robot policies remains open. In traditional classrooms, learners are shown erroneous examples, in which they reflect on and correct incorrect responses to understand common pitfalls when learning a concept. We propose using erroneous examples to teach robot policies, extending an existing policy teaching framework. We conduct a user study in which participants view incorrect demonstrations of robot behavior and correct the actions to align with the actual policy. Our findings suggest that viewing these incorrect demonstrations and verbalizing one's reasoning in predicting a robot's actions improves retention of the policy over time, in agreement with the effect of erroneous examples in classrooms. We also categorize participants into distinct learning styles and establish that participants using inverse reinforcement learning-like reasoning perform best on policy prediction tasks. With this work, we aim to advance the methods by which robots educate humans on their policies.
Rithika Narayan, Suresh Kumaar Jayaraman, H. Admoni· 0 citations
HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills.
Guangyan Chen, Meiling Wang, Te Cui et al.· arXiv.org· 2 citations
A training method for HIL online reinforcement learning for real robots that automatically switches between learning from interventions and on-policy self-improvement, reducing the policy--target-sample gap that otherwise induces execution-time distribution shift.
Zihang Wang, Yishan Wang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.