Edge-Efficient Compositional Recognition via Disentangled Prompt Tuning of Frozen Vision-Language Models
This work proposes Progressively Disentangled and Recurrent Prompt Tuning (PDRPT), an edge-efficient framework that decouples object and state updates before joint refinement, suppresses traction force from highly-entangled prompts, and preserves alignment with the natural language space of CLIP.