Skip to content

Author

Petr Motlícek

We have 6 of 28 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an...

Hritika Sharma, Thibault Bañeras-Roux, Alessandra Pinto et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state throu...

Hasindri Watawana, Sergio Gastón Burdisso, Esaú Villatoro-Tello et al. · 0 citations
Preprint Aug 2026

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

The results show that encoder-based metrics remain highly competitive, while generative LLMs perform strongly in hypothesis comparison and improve the interpretability of ASR evaluation.

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil et al. · 0 citations
Preprint Jul 2026

When Synthetic Speech Is All You Have: Better Call GRPO

This work shows that Group Relative Policy Optimization (GRPO) extracts far more from the same synthetic speech than SFT, and traces the gain to behavior rather than representation: GRPO reduces insertion errors by improving stopping calibration and speech-to-text alignment by better anchoring attention to audio, leavi...

Shashi Kumar, Yanis Labrak, Hasindri Watawana et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.