C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes
This work introduces C-GAP Caption-Guided Augmentation and Prompting), a detector-agnostic, annotation-free framework that operates in two phases, and establishes a composite caption baseline combining per-image scene descriptions with class-quantity context, which is shown to outperforms scene-description only or class-quantity-only prompts across multiple open-vocabulary architectures and benchmarks.