2026· International Conference on Language Resources and Evaluation· pp. 4647-4657· 0 citations· 31 references
Computer Science
TL;DR
This paper presents a methodology for building safety evaluation datasets that comprehensively cover the full spectrum of sensitive topics relevant to LLM safety, and releases a public repository containing the list of categorized Italian Wikipedia pages, the automatically generated prompts, and the standard prompt template used for safety testing.
As large language models are deployed across multilingual environments, the benchmarks used to evaluate their safety remain largely designed for high-resource, English-dominant contexts. Trust & Safety systems increasingly rely on automated classifiers and generative models to moderate harmful content across dozens of...
This work investigates cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts to demonstrate superficial safety alignment.
Abigail Oppong, P SAM SAHIL, Tadesse Destaw Belay et al.· 1 citation
SurakshaEval is introduced, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu - along with English.
Debopriyo Banerjee, K. R. Kavitha, Angana Borah et al.· 0 citations
Assessment of an open‐source and proprietary large language models for analyzing stakeholder inputs from three Swiss pre‐parliamentary consultations highlights both the potential and limits of algorithmic support for participatory governance in linguistically and technically diverse environments.
Edgar Mathevet, Lauriane Cailleux· Policy Studies Journal· 0 citations
CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Malvina Nissim, Danilo Croce, V. Patti et al.· Italian Journal of Computati...· 0 citations
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue,...
Xian-Hui Zhang, Jian Yu, Cheng-Yu Xie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.