2026
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
This work proposes SkyAnchor, an MLLM with two designs to the above challenges: a Semantics-Aware Token Router that preserves small-target under a reduced visual-token budget, and a Hierarchical Memory Bank that keeps the target consistently understood on streams.
Penglei Sun, Yehua Huang, Zhuoli Tao et al.
· arXiv.org · 0 citations