Skip to content

Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation

Jun 2026 · arXiv.org · Vol abs/2606.11891 · 0 citations · 23 references
Computer Science

TL;DR

It is argued that critic architecture deserves explicit treatment as a design variable in multi-objective humanoid RL, and specify the single-variable ablation needed to establish its causal contribution.

Abstract

Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within one policy. A natural design choice is between a single (unified) critic that estimates the combined value of all objectives and separate (dual) critics with disjoint reward signals. We compare the two on the Unitree G1 humanoid in NVIDIA Isaac Lab. In the standing mode of a standardized evaluation, the dual-critic run reaches targets 3.5x faster (6.5 vs. 22.6 simulation steps), achieves 2x the throughput (14.3 vs. 7.0 validated reaches per 1,000 steps) and a higher validated reach rate (65.2% vs. 53.8%) than the unified-critic run. That evaluation pins the fingers open for every policy, whereas the unified run had trained driving its own. When the unified run drives its own fingers, with nothing else changed, the standing-mode gap falls from 3.5x to 1.3x in speed and from 2x to 1.1x in throughput. This is a single re-evaluation of a single checkpoint, and we do not generalize from it. Adding five anti-gaming reward mechanisms to the dual critic did not raise validated reach rate (60.9% vs. 65.2%). The two runs differ not only in the critic but also in the PPO update rule (one summed advantage under one likelihood ratio, versus a per-stream advantage and a ratio per actor), and further in curriculum, arm action dimensionality, finger control and reward weights; each is a single run. The measurement therefore cannot separate the critic from the update rule. We argue that critic architecture deserves explicit treatment as a design variable in multi-objective humanoid RL, and specify the single-variable ablation needed to establish its causal contribution. Code, checkpoints and a project page: https://mturan33.github.io/critic-architecture-matters/

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL...

J. Sahoo, Saurabh Kumar, Surya Prakash S.K. et al. · 0 citations
Preprint Oct 2026

Beyond Task Reward: A Controller-Restriction Protocol for Evaluating Embodiment-Dependent Competence

Co-design methods optimize a robot's body and controller jointly and judge the result by one number, the task reward of the fully optimized pair. That number cannot separate morphologies whose competence depends on the controller to very different degrees. We evaluate a morphology by restricting its controller instead,...

Si-Yuan Zhang · 0 citations
Conference Aug 2026

Reinforcement Learning for Bipedal Locomotion Using Minimal Instrumentation with a Single Inertial Measurement Unit

This work presents a simulation-based validation framework for locomotion control on a custom-built 13 DoF bipedal robot using only signals derivable from a 6-axis inertial measurement unit (3D angular velocity and 3D gravity vector projection) as actor observations. The system employs the Genesis World simulator and t...

Juan Esteban Gomez Lopez, Yesid Eugenio Santafe Ramon · 0 citations
Preprint Aug 2026

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

A unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA) is introduced that enables multiple actors to share a centralized multi-head critic and substantially improves both sample efficiency and policy performance.

Changhao Li, Yifang Zhang, Heng Zhang et al. · 0 citations
Preprint Sep 2026

Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies

GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and fee...

Galbot Team Xuchuan Chen, Xiao-Qi Cheng, Yu Deng et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.