Generalized zero-shot learning (GZSL) aims to recognize novel categories by leveraging semantic knowledge transferred from previously observed categories, where learning transferable features for effective visual-semantic alignment plays a critical role. Conventional methods typically utilize all shared attributes to learn semantically related features. However, not all attributes are related to a specific instance, and the inconsistency may result in less effective image features. Additionally, the visual features extracted by the pre-trained backbone may be unsuitable for a particular sample, reducing the discrimination. To address these problems, we propose a novel instance-level visual and semantic adaptation framework that effectively adapts pre-trained image features for the GZSL tasks. From the semantic adaptation aspect, we propose to choose instancespecific attributes for dynamic semantic prompt tuning. From the visual adaptation aspect, we construct instance visual proto-types to produce channel attention, which adaptively strengthens crucial visual features for each sample. Extensive experiments on three GZSL benchmark datasets demonstrate that our approach achieves the new state-of-the-art performance.
Hua-Jie Jiang, Zheng-Xian Li, Xi-Chao Yu et al.· IEEE Transactions on Image P...· 0 citations
Dementia affects over 57 million people worldwide and places an immense burden on informal caregivers, yet current AI tools remain largely fragmented across isolated tasks and modalities. Recent large language models (LLMs) and vision-language models (VLMs) offer promising capabilities for dementia support, but adapting them to this safety-critical, multimodal, and deeply individualized care domain raises challenges that general-purpose AI surveys do not address. In this paper, we present a challenge-driven survey that organizes the rapidly growing literature on LLMs and VLMs for dementia care around three core adaptation challenges: (1) Knowledge Grounding, which anchors model outputs to verified clinical knowledge through retrieval augmented generation, knowledge graphs, and constrained training to mitigate hallucination risk; (2) Multimodal Understanding, which fuses visual, audio, and sensor data with language to reason about the multimodal inherent of daily care; and (3) Personalization, which adapts model behavior to individual patient histories, caregiver needs, and unique disease progression over time via persistent memory and biography-driven interaction. We review over 20 recent methods, identify cross-cutting architectural patterns, and survey available datasets and benchmarks. Finally, we highlight critical open challenges including the need for standardized evaluation protocols, longitudinal deployment studies, and tighter integration between clinical workflows and foundation model capabilities.
Afrouz Sheikholeslami, Dexuan Ding, Amin Beheshti et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.