Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\% pass rate, with many tasks unsolved by any model, highlighting that multilingual coding competence is a distinct and underexplored capability axis.
Yunsu Kim, Kaden Uhlig, Ashwin Purohit et al.
· 0 citations