Jul 2026· International Journal of Precision Engineering and Manufacturing· 2 citations· 18 references
TL;DR
An LLM-based Voice-to-Action (VTA) system that converts spoken user commands into grounded robot behaviors for an indoor mobile manipulator, LeeAhn 2, and suggests that LLM-grounded spoken interfaces can reduce operator burden and improve accessibility for indoor service robots.
Abstract
Voice interaction has been actively studied in human–robot interaction (HRI) for decades, yet deploying spoken interfaces on physical mobile manipulators remains challenging because language is ambiguous, tasks are long-horizon, and robot actions must be grounded to perception and motion under real-world uncertainties. Recent large language models (LLMs) offer a practical way to interpret open-ended spoken requests, but their non-deterministic outputs and limited transparency can hinder safe and reproducible robot execution. This paper presents an LLM-based Voice-to-Action (VTA) system that converts spoken user commands into grounded robot behaviors for an indoor mobile manipulator, LeeAhn 2. The system combines speech transcription with an LLM that produces structured, skill-level action plans aligned with a predefined library of robot capabilities, including vision-based seeking, wheeled navigation, and manipulation. To improve reliability, we incorporate interface constraints that restrict generated actions to executable skills and enable recovery from common failures during execution. We evaluate the proposed system through simulation experiments and real-world trials on a representative search task, reporting component-level performance across seeking, navigation, and manipulation. The simulation experiments provide repeatable analysis under controlled conditions, while the real-world trials demonstrate practical applicability on the LeeAhn 2 mobile manipulator and reveal limitations such as latency and occasional plan/execution failures. Our results suggest that LLM-grounded spoken interfaces can reduce operator burden and improve accessibility for indoor service robots.
Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.
The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.
Natural language interfaces can lower the expertise barrier for operating robotic manipulators by allowing users to express goals in everyday speech. This paper presents RANLP-Arm, a modular bilingual voice-to-manipulation system that converts spoken Thai or English instructions into real-time robotic arm actions for tabletop pick-and-place tasks. The system integrates Gemini-based live voice interaction with function calling, stereo 3D object perception and segmentation, persistent object tracking with EMA smoothing and coordinate locking, camera-to-robot calibration with optional IDW residual correction, and TCP/IP robot control in a unified Python architecture for real-time operation. We evaluate the system on a physical robot using 100 pick-and-place trials and 40 bilingual voice-commanded trials. The affine calibration model achieves a mean correspondence error of 3.72 mm on 13 calibration points, while end-to-end experiments achieve 82.0% task completion, 100.0% intent recognition, and 95.0% voice-to-task success. All observed failures were caused by grasp instability rather than language understanding, perception, or calibration errors. These results show that bilingual voice-driven manipulation is practical for tabletop pick-and-place, while the main remaining limitation lies in end-effector robustness.
Supakorn Thavornvong, A. Kitsommart, Mahannop Thabua et al.· 2026 23rd International Conf...· 0 citations
This paper proposes an HRI agentic artificial intelligence system—integrating a construction domain-specific, noise-robust automatic speech recognition (ASR) agent and a vision-language model (VLM)-based robotic control agent—to reliably transcribe speech in noisy environments and parse instructions for robot navigation in acoustically challenging construction jobsites.
Oscar Poudel, Rayan H. Assaad, Mohamad Awada· Journal of computing in civi...· 0 citations
AnthroDial is presented, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment and shows that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.
Wentao Liu, Si-Yu Song, Xi Chen et al.· arXiv.org· 0 citations
In conversational Human-Robot Interaction, robots typically remain silent during user speech and reply only after a pause, making interaction feel unnatural. In contrast, humans signal that they listen through active behavior. To overcome this, we present a system in which a social robot conveys active listening through non-verbal backchannels grounded in interactional intents. The system combines two ideas: (i) a dual-stage framework separating the user’s communicative intent (Speaker Intent) from the robot’s interactional stance (Listener Intent), mapping the latter to non-verbal reactions; and (ii) a parallel pipeline whose chunk-level branch generates non-verbal feedback during speech while a turn-level branch produces the verbal reply at turn end. We deployed this system on a robot and conducted a usability study (N = 6) in a hotel-negotiation task. We found that participants considered the system usable and could interpret gestures. We contribute a ready-to-deploy intent-aware system to enable active listening for robots.
Yang Sun, Jan Leusmann, Michael A. Hedderich· Message Understanding Confer...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.