2026· SemEval@ACL· pp. 3035-3044· 1 citation· 20 references
Computer Science
TL;DR
This work presents NarSiL 1, a system for SemEval-2026 Task 4 Track A on Narrative Story Similarity that employs a Mixture-of-Experts initial classifier that also leverages supermajority voting across three large language models, followed by a structured three-pathway fallback for ambiguous cases.
Abstract
We present NarSiL 1 ( Nar rative Si milarity L earners), our system for SemEval-2026 Task 4 Track A on Narrative Story Similarity. NarSiL employs a two-stage architecture: a Mixture-of-Experts (MoE) initial classifier that also leverages supermajority voting across three large language models (Gemma-3-12B, GPT-3.5-turbo-instruct, and Gemini-2.5-Flash) over multiple runs, followed by a structured three-pathway fallback for ambiguous cases. The three pathways correspond directly to the task’s three core similarity components, abstract theme, narrative outcome, and course of action. Each path yields a similarity score corresponding to its respective component, and the scores are then combined through a weighted aggregation step. NarSiL achieves 64.25% accuracy on the official test set. An improved score of 70.25% is obtained by considering only the supermajority voting of GPT, followed by the previously described fallback.
Error analysis shows that a non-trivial fraction of failures are placeholder strings caused by API errors rather than incorrect generations, and that surface-level mismatches (verbosity, ortho-graphic variation) account for many of the remaining errors.
The hugang11 system addresses a practical trade-off in creative text generation: models that produce sharper and more stylized jokes often become less stable in output format, and builds a three-stage pipeline that combines chain-of-thought-augmented supervised fine-tuning (CoT-SFT), teacher-constructed direct preference optimization (DPO), and deterministic post-processing.
The lack of high-quality labeled datasets remains a major challenge for sentiment analysis in low-resource languages such as Indonesian, particularly in specialized domains like fiscal policy. This study investigates the effectiveness of Large Language Models (LLMs) as automated annotators within a teacher-student knowledge distillation framework. Using social media data from X related to Indonesia's Coretax system, three training scenarios were evaluated: AI-labeled data, human-labeled data, and a hybrid approach. The results show that GPT-4o achieves substantial agreement with human annotators, with a Cohen's Kappa score of 0.61. Furthermore, the student model IndoBERT trained on the combined dataset outperforms other configurations, achieving a Macro F1-score of 0.64 and a Macro ROC-AUC of 0.84. These findings indicate that while LLMs cannot fully replace human judgment, they significantly enhance scalability and enable near real-time policy evaluation in low-resource settings through effective human-AI collaboration.
Novialdi Ashari, Ulfah Oktarida Sihaloho, Novi Aulia Sari· International Seminar on Int...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.