The WMT 2020–2024 shared-task lineage with an extended English–Malayalam resource is consolidated into INDICQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.
The WMT 2020-2024 shared-task lineage with an extended English-Malayalam resource is consolidated into IndicQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level tags and an error explanation, and a test set stratified over four difficulty axes.
Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al.· 0 citations
COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target i...
The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in which some forms of human references are available):...
Diptesh Kanojia, Chi-Kiu Lo, Archchana Sindhujan et al.· 0 citations
MultiGhostBench is introduced, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book, that supports evaluation under domain, author, and language shifts.
M. Greco, Anudeex Shetty, Andrea Tagarelli et al.· 0 citations
This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...
A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shi...