Cross-Modal Feature Adapter for Few-Shot Human Activity Recognition.
A new cross-modal feature adapter is designed, which fine-tunes a pre-trained CLIP image encoder to effectively align image-sensor pairs and adaptively blend the old knowledge inherited from the original zero-shot CLIP with the new knowledge adapted from few-shot training samples, which makes training converges faster whilst forming a streamlined time series sensor encoder.