MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set, and it is found that models robustness in HRL setting do not necessarily translate to LRL.