Skip to content

Confident but Wrong: A Constrained Decoding Diagnostic for Low-Resource Automatic Post-Editing

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

A black-box, inference-time diagnostic that tells these two cases apart without retraining or annotation is introduced, with the pattern holding on English-Marathi and English-Tamil, with the failure modes tracking the post-edit distribution rather than MT quality or language family.

Abstract

Automatic Post-Editing (APE) for low-resource languages (LRLs) often fails to improve Machine Translation (MT), and the score alone cannot say why: whether more training would help, or whether the training data is too inconsistent to learn from. We introduce a black-box, inference-time diagnostic that tells these two cases apart without retraining or annotation. It varies an edit-distance penalty $\lambda$ that drives the model from free editing towards copying the MT, and reads two signals: (1) the shape of the Translation Edit Rate (TER)-vs-$\lambda$ curve, U-shaped if edits from the model reduce error and monotonically decreasing if none does; and (2) the ordering of constraint variants that trust model confidence to increasing degrees, which shows whether confidence tracks edit quality. Across decoder-only and encoder-decoder models on English-Sinhala, the diagnostic exposes two failure modes consistent with a heterogeneous post-edit signal as the underlying cause: Binary Collapse, where the model copies the MT or makes off-target edits, and Confident Miscalibration, where the confidence signals we test do not separate useful edits from unnecessary ones. The pattern holds on English-Marathi and English-Tamil, with the failure modes tracking the post-edit distribution rather than MT quality or language family. Beyond diagnosis, the curve shape prescribes a concrete next step for practitioners; in the favorable case, a static constraint yields a free inference-time accuracy gain. We release the first English-Sinhala (~66k) and a new English-Tamil (~39k) APE datasets with all code.

View source

Similar papers

Book Open access Aug 2026

Cost-Aware Human-LLM Collaboration for Post-OCR Corrections in Swiss Historical Newspapers

A regression-guided routing approach that prioritizes segments by predicted CER improvement, paired with a safeguard layer that detects harmful LLM corrections and routes uncertain segments to human review, and substantially outperforms standard confidence-based approaches is introduced.

Stergios Konstantinidis, Hayman Lotfy, Michalis Vlachos · 0 citations

INDICQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

The WMT 2020–2024 shared-task lineage with an extended English–Malayalam resource is consolidated into INDICQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.

Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al. · 0 citations
#natural language process... Preprint Sep 2026

In the Blind: Building Pseudo-References for MT Evaluation

The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in which some forms of human references are available):...

Diptesh Kanojia, Chi-Kiu Lo, Archchana Sindhujan et al. · 0 citations
Aug 2026

Augmenting Text to Increase Translation Difficulty

This work proposes augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator, and uses gradients from a combined difficulty and fluency objective to iteratively replace tokens in Adversarial Translation Optimization (AT...

William Kalikman, Šimon Sukup, Michal Tesnar et al. · 2 citations
Preprint Aug 2026

When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

It is found that, in this benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy and, under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement lear...

Siddharth Chauhan, Thomas Butler, Abhishek Singhania et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-...

Jin-Hao Zhang, Ze-Yu Liu, Zi-Cheng Yan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.