A Large Vision-Language Embedding Model trained on English multimodal tasks to support multilingual inputs is adapted and a new benchmark is introduced to evaluate the multilingual and multimodal capabilities of embedding models.
Abstract
Large Language Models are increasingly used for embedding extraction. In fact, there are many approaches that try to optimize the embedding representations that these models can learn, exploiting the knowledge gained from extensive pre-training and large parameter counts. However, most works currently focus on the English language and textual input only, reflecting the trend of current Large Language Model training corpora. Recently, several Large Vision-Language Models, which are Large Language Models capable of processing multimodal signals in input, have been released. Yet, their training procedure still remains predominantly based on English data. This limitation also affects the evaluation step, with embedding benchmarks that provide limited coverage for low-resource languages. To address these challenges, we adapt a Large Vision-Language Embedding Model trained on English multimodal tasks to support multilingual inputs. Furthermore, we introduce a new benchmark to evaluate the multilingual and multimodal capabilities of embedding models.
The results show that multilingual transfer is the dominant factor in extremely low-resource Bantu translation while eliminating the need for heuristic proxy selection, and all systems fail to preserve tonal diacritics, highlighting an open challenge.
Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu et al.· 0 citations
MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
Uri Katz, Omer Goldman, Tomasz Limisiewicz et al.· 0 citations
This work investigates how MLLMs learn novel concepts by introducing concepts in unimodal pre-training or multimodal fine-tuning, and evaluates the model’s ability to generalize between the two settings, and evaluates how deeply a model maps concepts across modalities.
S. Boppana, Tian Yun, C. Curto et al.· 0 citations
Results demonstrate that MoLGE consistently outperforms dense multilingual baselines with a minimal increase in trainable parameters, and suggest that structured language specialization provides an effective pathway for massively scaling language coverage of multilingual ASR.
An inference-time multilingual steering method that uses pretrained sparse autoencoders to identify and strengthen target-language-related features that are decoded into steering signals and injected into the model's hidden states without additional training is proposed.
It is observed that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities and proposes SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informa...
Yucheng Wang, Qihui Zhu, Yang Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.