The application of data mining technologies in the early identification of children with Autism Spectrum Disorder (ASD) has gained prominence. Facial image analysis has emerged as a popular method for its efficiency and scalability. However, current approaches often classify facial images into ASD categories without elucidating the specific contributions of different facial areas to the outcomes. To address this gap, we propose a novel method for ASD detection using facial images while concurrently identifying significant facial areas. Our approach integrates a Pre-trained Image Encoder to extract semantic information from the original image, a Gated Fusion Module to dynamically regulate the contribution of each pixel, and a scoring layer to predict ASD scores based on the fused feature map. Experimental validation on a publicly available dataset showcases the efficacy of our method, demonstrating commendable performance in terms of precision and recall metrics.
Mangna Fang, Yangyang Fang, Ran Wei et al.· International Conference on...· 0 citations
Transformers have achieved remarkable performance in video-based 3D human pose estimation, yet their high computational cost hinders deployment on resource-constrained devices. To balance accuracy and efficiency, this paper proposes an efficient plug-and-play joint enhancement framework, FG-Net, which integrates the Frequency Enhancement Module (FEM) and Gaussian Enhancement Module (GEM) to boost the performance of video pose Transformers. On this basis, FEM calibrates the semantic consistency of multi-scale features through frequency-domain detail enhancement and deformable spatial alignment, compensating for information loss caused by sampling. GEM constructs graph attention based on human skeletal topology, combined with temporal Gaussian smoothing and residual fusion, to adaptively strengthen joint structures, suppress noise and temporal jitter. The two modules work in synergy, enabling the model to achieve efficient inference while being more robust to low-quality video inputs such as occlusion and motion blur. Experimental results on the public dataset Human3.6M demonstrate that the proposed method achieves 39.87 mm MPJPE, obtaining state-of-the-art accuracy with lower computational complexity. The framework is generic and can be seamlessly integrated into mainstream video pose Transformers, providing an effective solution for real-time 3D human pose estimation in resource-constrained scenarios.
Shaojie Cheng, Dazheng Zhou, Ran Wei et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.