Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.
Andrea Caciolai, P. L. Cabot, Chierh Cheng et al.· 0 citations
This paper presents a method for aggressively pruning experts from modern mixture-of-experts LLMs while incurring negligible degradation in translation quality, and shows that translation requires only a fraction of the LLM, enabling substantial compression of the MoE blocks that contain over 90% of parameters.
Liu O. Martin, Lucas Bandarkar, Nanyun Peng· arXiv.org· 2 citations· ⚡1
There is potential to improve cross-lingual parametric knowledge transfer during post-training by providing the LLMs with the key entities of the questions in their source language and finding that this disproportionately improves cross-script questions.
Lucas Bandarkar, Alan Ansell, Trevor Cohn· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.