Skip to content

Author

Ya-Min Mao

We have 2 of 9 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

R4DSG introduces a relative 4D scene graph memory for long egocentric video, built on recent RGB-only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, and produces a retrieval-ready memory directly usable for long-horizon question answering.

Ke Ma, Ya-Min Mao, Weiming Li et al. · 0 citations
Preprint Sep 2026

H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs)...

Xiong-Feng Peng, Lu Xu, Yan-Dong Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.