Where Do Multilingual Vision-Language Encoders Fail on Low-Resource Languages?
A front-layer trunk that pulls each language's projection toward the parallel-content centroid corroborates the diagnosis at training time, with consistent gains across three further benchmarks while preserving HRL performance.
Donghoon Han, Sunghyun Moon, Aidyn Zhakatayev et al.
· 0 citations