Skip to content

Author

Xueying Sun

We have 2 of 20 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Frequency-aware elastic prototype boundary learning for long-tailed scene graph generation

Scene graph generation (SGG) addresses the task of detecting objects in an image and predicting the relationships among them. Although prototype-based methods have recently achieved clear progress on long-tailed SGG, fine-grained low-frequency predicates remain difficult to recognize because their relation features often exhibit larger intra-class variation and more dispersed distributions, making them easily confused with semantically similar high-frequency coarse-grained predicates under a unified prototype-matching rule. To alleviate this issue, we propose a frequency-aware elastic prototype boundary learning framework, termed SGE-Net. Under fixed relation prototypes, the framework learns relation-category-specific boundary scales through explicit frequency compensation and frequency-adaptive virtual sampling, so that relation prediction can exploit not only prototype-center matching but also category-dependent decision-boundary information. During inference, we further introduce elastic boundary-aware distance calibration, enabling the boundary information learned during training to better distinguish relation categories that are easily confused under prototype matching. In addition, we combine visual and semantic features with dynamic gating to provide more reliable relation features for the above boundary learning. Experiments and analyses on Visual Genome and Open Images V6 demonstrate that the proposed method achieves consistent gains in both long-tailed relation prediction and overall evaluation metrics.

Binghao Wang, Xueying Sun, Hanzhu Dai et al. · 0 citations
Preprint Aug 2026

NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

NebulaVLA is presented, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity and introduces GESTURE-7, a unified language-grounded action representation.

Congyu Zhao, Shuai Tian, Xu Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.