This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attribute steering vectors can preserve the intended steering effect of each attribut...
Hyunku Kang, Daniil Gurgurov, Tanja Baeumel et al.· 0 citations
It is argued that framing tokenization as output supervision provides a principled account of why tokenization consistently affects model performance, and that differences in task performance, training dynamics, and model internals are induced by output tokenization and largely invariant to input tokenization.
Tanja Baeumel, Josef van Genabith, Simon Ostermann· 0 citations
It is shown that translation is even more modular than previously assumed and that the output language production in translation processes is actually further separable into a syntax and a surface language process.
M. Sonkin, Tanja Baeumel, Daniil Gurgurov et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.