Vision Language Models (VLMs) extend large language models with visual perception, enabling complex vision tasks at the edge. Low-rank adaptation (LoRA) adapters offer a lightweight method to inject domain-specific knowledge into a shared base VLM, making them attractive for edge serving where concurrent mobile clients...
Wei-Jun Wang, Liang Mi, Jing-Han Chen et al.· IEEE Transactions on Mobile...· 0 citations
Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training...
Liang Mi, Wei-Jun Wang, Bo-Wen Gao et al.· 0 citations
Zetta is presented, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen, and shows that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.