Condensed Test-Time Adaptation of VLMs for Action Recognition
A novel training-free Condensed Dynamic Adapter C ON DA is proposed, which leverages vision-text alignment to guide vision-vision alignment and is compatible with arbitrary VLM and generalizes well across complex scenarios, such as long-term and egocentric scenarios.
Wenxuan Ge, Hongyu Qu, Rui Yan et al.
· 0 citations