DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning
DINO-A is presented, an adaptation of self-distillation from vision to general audio representation learning, and it is traced to two mechanisms: the interaction between DINO's high-dimensional projection space and FSD50K's limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2 does not.