While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping t...
Valentino Maiorca, Walter Nelson, Francesco Locatello· 0 citations
Many scientific questions require reasoning about what was never observed: What if the conditions, interventions, or history had been different? Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. A common remedy is to add controlled simulatio...
Ding-Ling Yao, Kahaan Gandhi, Valentin Duruisseaux et al.· 0 citations
Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $\sigma$, we treat the denoiser as a self-map and study its dynamics. For a...
C. Amado, Marco Fumero, Francesco Locatello· 0 citations
Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties...
Irene Tallini, Lorenzo Basile, Valentino Maiorca et al.· 0 citations
ADAPT---Amortized Distillation Across Post-Trained LLMs is introduced, a framework for amortizing distillation across both axes of a model family: size and post-training variant, producing models for interpolated sizes across post-trained variants with a single distillation run.
Yan Zhou, Sara Kangaslahti, Jonathan Geuter et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.