2026· SemEval@ACL· pp. 2729-2743· 0 citations· 16 references
Computer Science
TL;DR
A distribution-aware, LLM-augmented dataset was constructed by selectively paraphrasing minority-class instances to enhance class balance, and its performance was benchmarked against full, rebalanced, and undersampled training configurations.
Abstract
Detecting equivocation is essential, as indirect or evasive responses can shape public perception, influence political narratives, and undermine transparency in democratic discourse. To address the challenge of detecting evasive political responses on digital platforms, participation in the CLARITY SemEval-2026 Task was undertaken, which focuses on (i) clarity-level classification and (ii) fine-grained evasion-type classification in political question-answer contexts. This study introduces a data-centric framework that systematically examines the effects of class distribution and refinement strategies on the performance of Large Language Models (LLMs). A distribution-aware, LLM-augmented dataset was constructed by selectively paraphrasing minority-class instances to enhance class balance, and its performance was benchmarked against full, rebalanced, and undersampled training configurations. To comprehensively assess the proposed method, Qwen3-14B, Phi-4, Gemma-2 9B, and Mistral 7B were evaluated in in-context learning (ICL) settings (zero-shot and few-shot) and with LoRA fine-tuning. Experimental results indicate that fine-tuning Phi-4 with class rebalancing yields strong performance, achieving 74.77% on Subtask-1 and 51.55% on Subtask-2
Political evasion refers to responses that engage with a question while withholding the requested information. Recent NLP work frames political evasion as a classification task using a two-level taxonomy of response clarity and fine-grained evasion strategies. Existing work on response clarity and evasion classification is limited to English, leaving open whether the taxonomy and model behavior transfer across languages and political contexts. We introduce PolERo, a dataset of 3,574 human-annotated question-answer pairs extracted from official transcripts of five Romanian presidents. We evaluate multiple classification approaches on both datasets under matched conditions, including TF-IDF baselines, fine-tuned encoder models, a proposed sliding-window encoder, and zero/few-shot LLM prompting. We study cross-lingual transfer through joint bilingual training and machine-translation-based data augmentation. Our results indicate that fine-tuned encoders are competitive, cross-lingual transfer is asymmetric, and ambivalent evasion categories involving pragmatic cues remain the main challenge across all model families.
This work evaluates prompting strategies for subtask 2 of the GermEval 2025 Harmful Content Detection challenge, which involves classifying whether a tweet attacks the free democratic basic order and shows that techniques such as Chain-of-Thought, In-Context Learning or Task Decomposition outperform approaches like Task Description.
It is proposed that persona-based evaluation can serve as a scalable diagnostic of what generative systems value and prioritize when depicting humanity, and that persona generations are far from neutral.
N. Corrêa, Rafaela Weber Mallmann, David Kaczér et al.· Artificial Intelligence Revi...· 0 citations
This proposed approach combines cross-validation, structured aggregation and bias-aware evaluation to optimize the robustness–performance trade-off, and achieves 93.19% accuracy with a TCE of 3.13, yielding a strong combined score of 38.56 under the official evaluation metric.
The lack of high-quality labeled datasets remains a major challenge for sentiment analysis in low-resource languages such as Indonesian, particularly in specialized domains like fiscal policy. This study investigates the effectiveness of Large Language Models (LLMs) as automated annotators within a teacher-student knowledge distillation framework. Using social media data from X related to Indonesia's Coretax system, three training scenarios were evaluated: AI-labeled data, human-labeled data, and a hybrid approach. The results show that GPT-4o achieves substantial agreement with human annotators, with a Cohen's Kappa score of 0.61. Furthermore, the student model IndoBERT trained on the combined dataset outperforms other configurations, achieving a Macro F1-score of 0.64 and a Macro ROC-AUC of 0.84. These findings indicate that while LLMs cannot fully replace human judgment, they significantly enhance scalability and enable near real-time policy evaluation in low-resource settings through effective human-AI collaboration.
Novialdi Ashari, Ulfah Oktarida Sihaloho, Novi Aulia Sari· International Seminar on Int...· 0 citations
Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety alignment outside English. A natural hypothesis is that these compound. We test it directly. Using a deterministic IICL operator and a StrongREJECT-style rubric judge, we red-team two Google Gemini models on two benchmarks, a 30 general-harm behaviours from HarmBench and 30 financial-abuse behaviours from FinProof, each under a single-shot baseline and under IICL in four languages (English, Spanish, Hindi, Arabic). First, IICL generalizes to a second provider and is worse in finance: it lifts attack success from<=6.7% to 80-90% on HarmBench and 97-100% on FinProof, an order of magnitude above the<=24% its introducing study reported on OpenAI's GPT-5.4. Second, against the hypothesis, forcing the IICL output into a non-English language does not stack the two weaknesses, it attenuates the attack. Eleven of twelve non-English conditions score below their English baseline (sign test, p~0.003), the lone exception a ceiling tie near 100%; on the stronger model's financial set Arabic collapses from 100% to 33%. We attribute this to a relevance curse: once structure has unlocked compliance, the models produce lower-quality harmful content in lower-resource languages, which a substance-grading judge scores as partial. The pattern replicates under an independent non-Google judge (Cohen's kappa=0.86, 377 paired verdicts), and 76.6% of non-English responses were verified in-language. Jailbreak vulnerabilities are therefore not additive; the dominant residual risk is the English structural attack, most acute for financial abuse, not a multilingual one.
Tejasvi C. Addagada· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.