Skip to content

Author

Katharina von der Wense

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spannin...

Mario Sanz-Guerrero, M. Bui, Manuel Mager et al. · 1 citation
#natural language process... Preprint Sep 2026

Calibration as a First-Class Criterion in LLM Evaluation

It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.

Mario Sanz-Guerrero, Katharina von der Wense · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.