Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spannin...
Mario Sanz-Guerrero, M. Bui, Manuel Mager et al.· 1 citation
It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
Mario Sanz-Guerrero, Katharina von der Wense· 1 citation
It is concluded that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.
M. Bui, Mario Sanz-Guerrero, Abteen Ebrahimi et al.· 0 citations
The translationese content of MLLM generations is assessed and the key features that distinguish MLLM-generated text from typical translation-related interference are examined.
Maria R. Valentini, Téa Wright, Julisa Granados et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.