A causal-transformer warm-start for the computationally dominant attitude-manipulator stage of a two-stage sequential convex programming (SCP) framework is developed and demonstrated that learned warm-starts can reduce onboard optimization cost while steering SCP toward basins containing feasible, low-cost solutions.
Yuji Takubo, M. Adang, Mac Schwager et al.· arXiv.org· 2 citations
Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions and visual observations to actions in a single forward pass. While conceptually simple, this formulation entangles instruction comprehension,...
Weilong Guo, Yuchen Wang, Renping Zhou et al.· 0 citations
Hybrid aerial--ground robots can use thrust to cross obstacles that impede wheel-driven motion, but deciding how much thrust to apply during contact remains challenging. We present an energy-aware reinforcement learning framework that jointly commands wheels, tilt servos and propellers through a single continuous polic...
Jiaxing Li, Ishaan Bhimwal, Wen Tian et al.· 0 citations
SAIL is a framework that reframes robot imitation as an iterative refinement problem capable of scaling with test-time compute, and utilizes Monte Carlo Tree Search, where each node is a complete trajectory and edges correspond to trajectory refinements.
Full-stack human-robot collaboration (HRC) can become brittle when replacing a planner, partner model, coordination policy, or controller changes the physical meaning of cross-layer signals. We introduce C2C, an object-centric cognition-to-control architecture that preserves these meanings through physical contracts. R...
Hao Zhang, Yisen Li, Ruize Geng et al.· 0 citations
Across five real-robot settings with all policies running on the same Jetson AGX Thor, HybridFlow improves normalized task performance over 16-step Diffusion Policy by 13-68 points with approximately eightfold lower action-generation latency.
Zhen Dong, Fu-Lin Chen, Jin-Na Fu et al.· 0 citations
While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large language models (LLMs) and vision language models (VLMs), opens the possibility of teams that infer mission-relevant semantics and subtasks given high-level natural language spec...
Zachary Ravichandran, Fernando Cladera, Ankit Prabhu et al.· 0 citations
Addressing the challenge of ensuring safety in ever-changing and unpredictable environments, particularly in the swiftly advancing realm of autonomous driving in today's 5G wireless communication world, we present Navigation Secure (NavSecure). This vision-based navigation framework merges the strengths of world models...
Hong Ding, Ziming Wang, Yi Ding et al.· 0 citations
RankQ, an offline-to-online Q-learning objective that augments temporal-difference learning with a self-supervised multi-term ranking loss to enforce structured action ordering is proposed, which shapes the Q-function such that action gradients are directed toward higher-quality behaviors.
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a...
Hao-Ran Yuan, Ze-Kai Wang, Boning Shao et al.· 0 citations
Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries - where trunks and primary branches are trained into approximately planar walls. We introduce an end...
Abhinav Jain, Cindy Grimm, Stefan Lee· IEEE Transactions on Field R...· 0 citations
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accuracy and compute. Potential fields learned directly from images remove that dependency but inherit the classical weakness...
Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.