Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this st...
Shu-Yi Miao, Wangjie Qiu, Pengyang Shao et al.· 0 citations
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue,...
Xian-Hui Zhang, Jian Yu, Cheng-Yu Xie et al.· 0 citations
A neuron-level cross-dimensional safety alignment framework driven by modality- and language-shared safety neurons (MLS-Neurons) that significantly outperforms state-of-the-art approaches across diverse multilingual and multimodal safety benchmarks while preserving general utility.
Enyi Shi, Fei Shen, Chuancheng Shi et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.