Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection
The TaSe (Talk in Pieces, See in Whole) framework is introduced with three main contributions: a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embeddings into subspace compositions; and aggregating disentangled components into hierarchically structured embeddings guided by the proposed hierarchical objectives.