Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlled audio conditions, and pessimistic scores from overly rigid transcription standards that penalize valid linguistic variations. We introduce Vimarsha, a 100-hour benchma...
K. Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri et al.· 0 citations
The WMT 2020–2024 shared-task lineage with an extended English–Malayalam resource is consolidated into INDICQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.
Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al.· 0 citations
In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-correct...
The WMT 2020-2024 shared-task lineage with an extended English-Malayalam resource is consolidated into IndicQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level tags and an error explanation, and a test set stratified over four difficulty axes.
Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.