Skip to content

Author

C. Yuan

We have 5 of 111 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward 3D foundation models. However, their inherent reliance on static-scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critical gap, we propose SPAR, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation. Furthermore, we introduce a dynamic-region-aware end-to-end training paradigm that structurally couples motion estimation with multi-view visual and semantic learning. This unified approach enables the network to inherently resolve motion conflicts and distill multi-view consistent, temporally stable scene representations from dynamic inputs. Extensive experiments on the challenging D-RE10K benchmark demonstrate that SPAR achieves state-of-the-art performance. Our end-to-end approach achieves exceptional novel view synthesis quality, yielding a PSNR of 22.15 dB and 23.33 dB given only 3 and 4 input views respectively. Despite being trained in a self-supervised manner, our model achieves an mIoU of 88.5% for motion mask prediction. Furthermore, our analysis reveals a strong inter-task synergy between photometric scene reconstruction and semantic understanding, where semantic synthesis learning consistently enhances photometric fidelity in novel view rendering. Code will be available at https://github.com/dmucby/SPAR.

Boyu Cai, Li Yang, Yan Xu et al. · 0 citations
Jul 2026

MMAgent-R2: Learning to Rerank and Reject for Agentic mRAG

MMAgent-R$^2$, an agentic mRAG framework that integrates visual reranking and active rejection as its internal verification mechanism, is proposed and achieves joint optimization of external retrieval, internal verification, and answer generation via GRPO training.

Tao Zhang, Ziqi Zhang, Zongyang Ma et al. · 0 citations
Jul 2026

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

Gaussian Mixture Modeling for Event-Aware Visual Allocation is proposed, which leverages Gaussian Mixture Models to model event-level structure from discrete frame-wise observations and achieves comparable performance to baseline selection methods while utilizing only approximately half of the visual token budget.

Yifan Lu, Ziqi Zhang, C. Yuan et al. · 0 citations

SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names

A Semantic-Retrieval-Augmented Detector (SRA-Det) is proposed that uses an attention-based module to retrieve multiple semantic facets from token-level text features, and a soft-min matching rule that behaves like a differentiable logical AND over these facets, ensuring that all key attributes are satisfied.

Li Yang, Boyu Cai, Wei Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.