When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs
This work presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers to qualify the language-universality of safety alignment as architecture-dependent and offer a mechanistic account of multilingual safety interventions.