Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
Light-Omni is introduced, a multimodal agent framework for reflexive and lightweight video understanding that achieves semantically aligned retrieval and reflexive responses while avoiding iterative reasoning and serves as a memory system to enhance both the performance and efficiency of existing MLLMs.