Skip to content

Author

Nicu Sebe

We have 10 of 28 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

Kinematics-Centric Continuous Sign Language Retrieval with Gloss-Guided Boundary-Aware Alignment

Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sig...

Chang Liu, Ke Han, Davide Talon et al. · 0 citations
Review Oct 2026

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scal...

Jian Liu, Wei Sun, Zhen-Qi Dai et al. · 0 citations
Review Sep 2026

The Past Frames the Future: Memory for Autoregressive Video Generation

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands,...

Harold Haodong Chen, Rong-Jin Guo, Di-Sen Lan et al. · 0 citations
Preprint Sep 2026

Reconstructing the Dynamic World: A Representation-Centric View of 4D Scene Reconstruction

4D scene reconstruction aims to recover the evolving geometry, appearance, and motion of dynamic environments from visual observations. Despite substantial progress in neural scene representations, reconstructing dynamic scenes remains challenging due to non-rigid motion, occlusions, temporal inconsistencies, and the t...

Zi-Ren Gong, Guo Chen, Yong-Jian Li et al. · 0 citations
Review Aug 2026

Human-Centric Intelligence in the Era of Foundation Models: A Survey

A full-spectrum human context taxonomy is introduced that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agen...

Yang Chen, Tianqi Wang, Xiao-Wen Jiang et al. · 0 citations
Preprint Aug 2026

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence, contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alp...

Fei Ma, Ze-Bang Cheng, Ming-Hui Li et al. · 0 citations
Conference Open access Sep 2026

Gradient Enhancement Task Aware Post-training Quantization

This paper introduces Gradient Enhancement Task Aware Post-training Quantization, i.e., GTAQ, to address the generalization issue of Large Language Models, and extensively evaluates the LLaMA family of language models on WikiText, C4, and MMLU.

Yi-Hua Shao, Yang-Yang Gu, Min-Xi Yan et al. · 1 citation

Cross Domain Test Time Scaling: Scale Knowledge and Reasoning on Cross Domains

Cross-Domain TTS is proposed, a novel framework that enables task-tailored scaling in broader domains and achieves an improvement of up to 17% in pass@1 accuracy while reducing inference latency and saving up to 30% in token consumption.

Minxi Yan, Yihua Shao, Yanling Pan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.