Skip to content
#natural language processing Preprint Open access

Last Translation Benchmark

Vil\'em Zouhar Niyati Bafna Mukund Choudhary Maike Z\"ufle Sara Rajaee Pinzhen Chen Jannis Vamvas Sara Papi Ona de Gibert Bhavitvya Malik Eliya Habba Orfeas Menis Mastromichalakis Patr\'icia Schmidtov\'a Michelle Wastl Sheriff Issaka Leshem Choshen Stella Biderman Antonis Anastasopoulos Jan Niehues Rico Sennrich Mrinmaya Sachan Ond\v{r}ej Bojar Kenton Murray J\"org Tiedemann Alham Fikri Aji Philipp Koehn Christof Monz Alexandra Birch Sowmya Vajjala Chalamalasetti Kranti Cristina Espa\~na-Bonet Nobin Sarwar David Kacz\'er Shunta Asano Malik Marmonier Daban Q. Jaff Vaisakhi Mishra Hend Al- Khalifa Gabriele Sarti Sourajit Saha Nils Rehlinger Juan Daniel Cuervo Villa Jonathan Tonglet Saugata Purkayastha Dominik Mach\'a\v{c}ek Jagannathan Ramanujam Heejin Do Zuzana Nadova Fred Philippy Fabian Retkowski Maria Lymperaiou Silvia Casola Hanna Yukhymenko Shubhashis Roy Dipta Sangwon Ryu Andr\'es Jerez Ron Keinan Shuaib Shuaib Yusuf Avantica Vempati Maria Carmen Staiano Sukannya Purkayastha Adrian Cosma Vitalii Babenko Erivan Inan Aviral Nigam Wafa Aissa Fatima Haouari Venkata Prasanth Kumar Gummadi Mehdi Jafarzadeh Valentin Scourneau Lukas Edman Kaiser Sun Shaomu Tan Mohammad Sadegh Gholizadeh Johannes-Rudolf David Dipankar Srirag Javier Garc\'ia Gilabert Ruta Binkyte Manar Ali Ana-Maria Bucur Sabry E. Farrag Youssef Saber Yihong Liu Jean Maillard Cojocaru Nicoleta Xiaochuang Yuan Sina Ahmadi Philipp Mondorf Kaustubh Dhole Roman Wixinger Shenbin Qian Manuel Tuor Sergey Troshin Jonathan Yahav Fida Mohammad Thoker Amir Arsalan Rezapour Lance Calvin Lim Gamboa Manon Reusens K\"atriin Kukk Koel Dutta Chowdhury Giuseppe Gallipoli Christian Hoang Shaswati Saha Seth Aycock Jan Koco\'n Bo Chen Linh Vu Vatsal Venkatkrishna Arafat Ahsan Luan Thanh Nguyen Hassan Soliman Daryna Dementieva Theresia Veronika Rampisela Ngoc Quynh Tram Do Marius Huber Kazuki Egashira Azmine Toushik Wasi Vladislav Poritski Mike Zhang Deep Shah Paul Gavrikov Luis Frentzen Salim David Africa R. Damanhuri Bello Umar Bello Anumit Garg Gengyu Rao Pawan Sasanka Ammanamanchi Kamile Dementaviciute Andrianos Michail L D M S Sai Teja Dawei Zhu Yi Fan Wei Liu Farhan Farsi Elias Herranen Sankalan Pal Chowdhury Karen Sanchez Farzad Shami Ashok Urlana Zimu Wang Tomasz Limisiewicz Priyaranjan Pattnayak Marii Ojastu Hongbin Na Emilian Radoi Chenyi Zhao Carlos Hinojosa Andrea Gregor de Varda Zaid Alyafeai Reem Alzahrani Nehal Kathrotia Alex Fl\"uckiger Ulysses Sekai Tully Carr Jimson Paulo Layacan Guy Kaplan Ritwik Tiwari Rishit Dagli Oksana Volchek Isaac R Caswell Bowen Yi Blanka K\"ov\'er Amir Hossein Yari Aicha Chorana Zhengxiang Wang Selja Ker\"anen Samuel Simko Joy Olusanya Jenny Chim Enzo Doyen Vivek Harsha Lakkamaneni Sophia Conrad Pouya Sadeghi Panayiotis Panayiotou Luis Lara Jannatul Nayem Eran Yahav Debanshu Das Antonia Karamolegkou Anmol Goel Aishik Mandal Tommaso Cerruti Raoyuan Zhao Mykola Haltiuk Thura Aung Naser Almousa Amir Hossein Kargaran Rachel Bawden Qiaoyuan Zheng Mateusz Lango Beni Egressy Fidel Rodr\'iguez Vel\'asquez Natchapon Jongwiriyanurak Minh Ngoc Do Marco Gaido Lena Libon Dzmitry Kuzmin Badal Nyalang Antoine Taroni Andrei Niculae Abdulaziz Nura Kani Rushikesh Zawar Marek \v{S}uppa Beatrice Savoldi Andreas Simons Rayyan Merchant Ilai Yaron Levy Francesco Pinto Ziyi Yang Yolanda Xavier Samuel Frontull Muhammad Ravi Shulthan Habibi Kenneth Enevoldsen Harris Abdul Majid Francesca Padovani Tim Graf Tatiana Bielakova Sharifa Djurabaeva Shaoxiong Ji Raia Abu Ahmad Pavel Stepachev Jirui Qi Ayush Sunil Munot Alireza Pakniat Ayla Rigouts Terryn Yuxing Lu Yurii Paniv Xiyan Fu Tosin Adewumi Sunisth Kumar St\'ephane J. P. S. Thunus Shree Harsha Bokkahalli Satish Shayan Bali Prakhar Gupta Papa Abdou Karim Karou Diallo Matija Akrap Marko Culjak Krist\'yna Onderkov\'a Joseph Attieh Esrael Teferi Tensay Elisabeth Fittschen Beno\^it Sagot Jingwei Ni Yu Fan
Sep 2026
Natural Language Processing

Abstract

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18
#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#computer vision Open access Mar 2014

Happy software developers solve problems better: psychological measurements in empirical software engineering

A study with 42 participants investigates the relationship between the affective states, creativity, and analytical problem-solving skills of software developers and offers support for the claim that happy developers are indeed better problem solvers in terms of their analytical abilities.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 216 citations · ⚡13
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.