Skip to content

Author

Sebastian Padó

We have 3 of 10 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Aug 2026

Effects of Answer Format Variation on Gender Bias in Large Language Models

It is found that answer format does substantially alter measured outcomes, including reversals in order rankings, and the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment is highlighted.

K. Merzlyakova, Sebastian Padó, Franziska Weeber · 1 citation
Preprint Jul 2026

Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing

Large language models (LLMs) are known to be sensitive to prompt and input formulations. However, existing studies have focused on lexical realization and largely ignored constructional choice. This paper studies whether linguistic construction can systematically shift LLM decisions and where these shifts can be causally localized inside the model. We use political stance judgment as a meaning-sensitive case study and extend an English political statements dataset, resulting in six controlled linguistic rewrite types that preserve or invert the meaning of a statement. Experiments on four open-weight models show that stance instability affect both meaning-preserving and meaning-inversing rewrites. Because output shifts reveal that rewrites affect stance, but not where in the model, we apply activation patching, where activations from the original statement are substituted into the forward pass for the rewritten statement and measure which components recover the original stance distribution. The results show that mid-to-late decoder layers, especially block outputs at the final prompt position, provide the strongest restoration signal.

Langchen Huang, Sebastian Padó, Franziska Weeber · 0 citations

Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

It is concluded that current transformer models do not explain human morphosyntactic processing, and that evaluations of transformers as cognitive models must adopt rigorous, comprehensive experimental designs to avoid spurious generalizations from isolated syntactic configurations or individual models.

Titus von der Malsburg, Sebastian Padó · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.