Skip to content

Author

Yaoxian Song

We have 1 of 14 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

2026

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

This work proposes SkyAnchor, an MLLM with two designs to the above challenges: a Semantics-Aware Token Router that preserves small-target under a reduced visual-token budget, and a Hierarchical Memory Bank that keeps the target consistently understood on streams.

Penglei Sun, Yehua Huang, Zhuoli Tao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.