Skip to content

INDICQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

· 0 citations · 41 references

TL;DR

The WMT 2020–2024 shared-task lineage with an extended English–Malayalam resource is consolidated into INDICQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.

View source

Similar papers

#natural language process... Preprint Aug 2026

IndicQE-APE: A Consolidated Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

The WMT 2020-2024 shared-task lineage with an extended English-Malayalam resource is consolidated into IndicQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level tags and an error explanation, and a test set stratified over four difficulty axes.

Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al. · 0 citations
#machine learning Preprint Oct 2026

Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation

COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target i...

G. L. John Salvin, Swapnil Hingmire · 0 citations
#natural language process... Preprint Sep 2026

In the Blind: Building Pseudo-References for MT Evaluation

The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in which some forms of human references are available):...

Diptesh Kanojia, Chi-Kiu Lo, Archchana Sindhujan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

MultiGhostBench is introduced, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book, that supports evaluation under domain, author, and language shifts.

M. Greco, Anudeex Shetty, Andrea Tagarelli et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
#machine learning Review Oct 2026

Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score

A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shi...

Parth Kulshreshtha, Shivali Dalmia, Abhishek Mukherji · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.