Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding
This work demonstrates that cross-lingual word-to-speech mappings can be learned directly from visual grounding without transcriptions or explicit model training.