UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models
UniMem is presented, a framework that unifies high-level, multimodal memory and low-level control under one backbone that outperforms fixed-interval image sampling baselines in simulation and hierarchical baselines in hardware, while offering faster inference and a simple training pipeline for easy adoption.