Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

Unsupervised 2D Image-Based 3D Model Retrieval via Decision Boundary Alignment and Graph Semantic Propagation

Unsupervised 2D image-based 3D model retrieval (IBMR) aims to retrieve semantically relevant 3D shapes for a given 2D image query when 3D annotations are unavailable. This setting is challenging due to severe modality gaps, category-imbalanced mini-batches, inconsistent cross-domain decision boundaries, and mismatched semantic neighborhood structures. In this paper, we propose a unified framework that integrates Category-Aligned Sampling (CAS), Decision Boundary Alignment (DBA), and Graph Semantic Propagation (GSP) into a single optimization paradigm. CAS constructs category-consistent mini-batches to stabilize crossmodal learning. Built upon CAS, DBA leverages a masked Margin Disparity Discrepancy to regularize cross-domain class decision boundaries via an adversarial min-max objective, encouraging discriminative separation beyond marginal feature matching. To complement boundary-level regularization, GSP builds a crossdomain affinity graph over 2D and 3D samples and propagates supervision-induced relational structure through semantic message passing, explicitly preserving instance-level neighborhood consistency that is critical for retrieval. Extensive experiments on MI3DOR and MI3DOR-2 demonstrate consistent improvements over representative unsupervised IBMR baselines.

Nian Hu, Yibo Zhao, Xinhui Li et al. · 0 citations
Jul 2026

A Temporal Action Detection Framework Based on Multi-Scale Temporal-Channel Collaboration

Temporal action detection is a crucial task in the field of video understanding, aiming to localize and recognize the category of actions in untrimmed videos along with their precise start and end times. Despite the success achieved by existing methods, two significant challenges remain: (1) difficulty in modeling long-range dependencies between video segments, leading to inaccurate localization of complex action boundaries; and (2) over-reliance on temporal dimension modeling, with insufficient exploration of the representational capabilities within the channel dimension, which limits model performance. To address these challenges, we propose a Multi-scale Temporal-Channel Collaborative (MTCC) detection framework. First, we construct a Long-Range Dependency Enhancement (LRDE) module based on the Mamba architecture with state space models (SSMs), which efficiently capture long-range temporal dependencies in video sequences. Second, leveraging the advantages of multi-scale CNNs, a Multi-Scale Boundary Awareness (MSBA) module is designed to extract local features and enhance the model's sensitivity to action boundaries. Finally, a Cross-Channel Information Fusion (CCIF) module is designed to extract features from the channel dimension of video. To better integrate features at different scales, a Multi-scale Detection Head (MDH) is employed to dynamically fuse the feature pyramid. Our systematic evaluations on the benchmark datasets THUMOS14 and ActivityNet-1.3 yield impressive results, demonstrating the effectiveness of our proposed method.

Yibo Zhao, Wen Zhang, Chunjie Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.