Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sig...
Chang Liu, Ke Han, Davide Talon et al.· 0 citations
Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scal...
Jian Liu, Wei Sun, Zhen-Qi Dai et al.· 0 citations
Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands,...
Harold Haodong Chen, Rong-Jin Guo, Di-Sen Lan et al.· 0 citations
4D scene reconstruction aims to recover the evolving geometry, appearance, and motion of dynamic environments from visual observations. Despite substantial progress in neural scene representations, reconstructing dynamic scenes remains challenging due to non-rigid motion, occlusions, temporal inconsistencies, and the t...
Zi-Ren Gong, Guo Chen, Yong-Jian Li et al.· 0 citations
Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).
Shaghayegh Kolli, S. Emami, Moreno D'Incà et al.· 0 citations
GrainGS is a dynamic Gaussian framework that combines a hierarchical anchor scaffold with per-Gaussian deformation that achieves high reconstruction quality, real-time novel view synthesis, and compact storage.
A full-spectrum human context taxonomy is introduced that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agen...
Yang Chen, Tianqi Wang, Xiao-Wen Jiang et al.· 0 citations
HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence, contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alp...
Fei Ma, Ze-Bang Cheng, Ming-Hui Li et al.· 0 citations
This paper introduces Gradient Enhancement Task Aware Post-training Quantization, i.e., GTAQ, to address the generalization issue of Large Language Models, and extensively evaluates the LLaMA family of language models on WikiText, C4, and MMLU.
Yi-Hua Shao, Yang-Yang Gu, Min-Xi Yan et al.· Proceedings of the Thirty-Fi...· 1 citation
Cross-Domain TTS is proposed, a novel framework that enables task-tailored scaling in broader domains and achieves an improvement of up to 17% in pass@1 accuracy while reducing inference latency and saving up to 30% in token consumption.
Minxi Yan, Yihua Shao, Yanling Pan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.