Skip to content

Author

M. Khapra

We have 4 of 174 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations

Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlled audio conditions, and pessimistic scores from overly rigid transcription standards that penalize valid linguistic variations. We introduce Vimarsha, a 100-hour benchma...

K. Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri et al. · 0 citations

INDICQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

The WMT 2020–2024 shared-task lineage with an extended English–Malayalam resource is consolidated into INDICQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.

Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al. · 0 citations
Jul 2026

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-correct...

Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi et al. · 0 citations
#natural language process... Preprint Aug 2026

IndicQE-APE: A Consolidated Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

The WMT 2020-2024 shared-task lineage with an extended English-Malayalam resource is consolidated into IndicQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level tags and an error explanation, and a test set stratified over four difficulty axes.

Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.