Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-langua...
This work proposes a calibration-free procedure for the estimation of contact surface normals using Universal Photometric Stereo neural networks, and demonstrates that with sufficient illumination settings surface normals could be estimated using a model trained solely on synthetic data.
This work introduces manifold-stable flow matching (MSFM), which can start from an arbitrary ambient prior, not necessarily supported on the manifold, and guarantees manifold invariance and transverse convergence to the manifold within a desired time window.
Amirhossein Nazerian, A. Pezeshki, Jian-Guo Zhao· 0 citations
Adjoint Guidance Flow is proposed, which is a deterministic optimal control problem, whose optimal guidance is a costate that carries the terminal critic gradient back through the remaining flow, and regress the guidance network onto this costate while keeping both the VLA and critic frozen.
Jeongsol Kim, Youngjun Jun, Kyumin Choi et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work proposes and evaluates State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications, and showcases the features’ performance in learning physically-based billiard shots demonstrating high action precision and discovering multiple sho...
N. H. Kim, Markus Kirjonen, Perttu Hämäläinen· Motion in Games· 0 citations
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read, then measure how far reinforcement learning falls s...
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safet...
Bo-Yang Li, Matthew Kim, Sylvia L. Herbert· 0 citations
When the human reference scores zero on a metric, the released GTRS-Dense label generator for NAVSIM marks every candidate trajectory in the scene as passing it. NAVSIM's authors introduced this human-reference forgiveness to avoid penalizing contextually justified maneuvers when scoring one trajectory, and warned that...
Jiaxuan Guo, Jingxin Yang, Jiaqi Ye et al.· 0 citations
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matters. Holding the forwa...
Xing-Jian Li, Jian-Hua Z. Huang, Jun-Li Duan· 0 citations
World models simulate the consequences of action candidates, but good planning need not preserve every physical distinction required for accurate prediction. We formalize this gap through a hierarchy of mechanism, response, and decision sufficiency. Given a candidate set, the planning query determines which physical va...
Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu et al.· 0 citations
This work builds on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and proposes Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL.
Bo-Yan Li, Matthew Kim, Sylvia L. Herbert· 0 citations
Vision-Language-Action (VLA) models leverage large-scale pretraining to ultimately achieve generalist manipulation. Deployed VLA policies must support continual learning to acquire new tasks over time. Teaching a VLA a new task generally requires finetuning it on demonstrations of that task. However, naively finetuning...
Aayushi Shrivastava, Xunlan Zhou, Hong-Ru Zhao et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.