Skip to content

Author

Francesco Locatello

We have 5 of 78 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Oct 2026

Emergent Structure in the Marginal Attention Space of Language Models

While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping t...

Valentino Maiorca, Walter Nelson, Francesco Locatello · 0 citations
#artificial intelligence Preprint Sep 2026

Counterfactual Predictions in Scientific Emulators Without Controlled Experiments

Many scientific questions require reasoning about what was never observed: What if the conditions, interventions, or history had been different? Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. A common remedy is to add controlled simulatio...

Ding-Ling Yao, Kahaan Gandhi, Valentin Duruisseaux et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models

Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $\sigma$, we treat the denoiser as a self-map and study its dynamics. For a...

C. Amado, Marco Fumero, Francesco Locatello · 0 citations
#machine learning Preprint Sep 2026

LocUS: Head Selection and Subspace Projection for Targeted Activation Steering

Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties...

Irene Tallini, Lorenzo Basile, Valentino Maiorca et al. · 0 citations
Preprint Aug 2026

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

ADAPT---Amortized Distillation Across Post-Trained LLMs is introduced, a framework for amortizing distillation across both axes of a model family: size and post-training variant, producing models for interpolated sizes across post-trained variants with a single distillation run.

Yan Zhou, Sara Kangaslahti, Jonathan Geuter et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.