Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
This work proposes SkyAnchor, an MLLM with two designs to the above challenges: a Semantics-Aware Token Router that preserves small-target under a reduced visual-token budget, and a Hierarchical Memory Bank that keeps the target consistently understood on streams.