2026· SemEval@ACL· pp. 2793-2799· 1 citation· 22 references
Computer Science
TL;DR
This paper addresses SemEval-2026 Task 4 Track A: Narrative Story Similarity by reformulating it as an instruction-following generation problem, employing parameter-efficient fine-tuning via LoRA to adapt pretrained large language models for triplet-based narrative comparison and incorporating synthetic triplet samples generated by a large language model for data augmentation.
Abstract
Narrative similarity assessment requires models to reason beyond surface-level lexical overlap and capture higher-level plot structures and thematic relationships. In this paper, we address SemEval-2026 Task 4 Track A: Narrative Story Similarity by reformulating it as an instruction-following generation problem. We employ parameter-efficient fine-tuning via LoRA to adapt pretrained large language models for triplet-based narrative comparison. To overcome the limitations imposed by the scarcity of human-annotated data, we further incorporate organizer-provided synthetic triplet samples generated by a large language model for data augmentation. Experimental results demonstrate that our fine-tuned Qwen2.5-7B model achieves slightly better performance than the zero-shot GPT-4o-mini base-line. These findings underscore the effectiveness of task-specific adaptation combined with synthetic data augmentation for narrative similarity modeling.
NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.
Yuheng Huang, Jianlang Chen, Jiayang Song et al.· 0 citations
This work investigates continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that is curate from millions of news articles and demonstrates the importance of targeted evaluation in the adaptation process.
Lukas Borggren, Jenny Kunz, Marco Kuhlmann· 0 citations
Findings support the complete framework on the evaluated corpus, while the controlled comparisons indicate a modest complementary contribution from the BiLSTM and do not identify it as the sole source of the performance gains.
Shweta Bansal, S. Yogarayan, S. F. Abdul Razak· Information· 0 citations
The exponential growth of scientific literature has intensified the demand for automated summarization systems capable of producing abstracts that are both linguistically fluent and factually reliable. Existing approaches face a fundamental trade-off: encoder-decoder models such as BART and T5 maintain strong factual grounding but produce rigid, extractive outputs, while decoder-only large language models (LLMs) such as Llama and Gemma generate highly fluent text yet remain susceptible to hallucination. This paper proposes a two-stage Synergistic Hybrid Ensemble framework designed to resolve this dichotomy. In Stage 1, a fine-tuned BART-Large model generates a factually grounded scaffold draft from a structured input representation comprising the document title, key sentences, method highlights, and results summary. In Stage 2, a QLoRA-adapted Llama-3.2-1B model performs coherent rewriting and stylistic polishing by conditioning on both the scaffold draft and the original source document. Experiments conducted on the arXiv Scientific Research Papers Dataset using BERTScore and entailment-based Factual Consistency metrics demonstrate that the proposed ensemble achieves a Factual Consistency metrics demonstrate that the proposed ensemble achieves a Factual Consistency score of 0.9140, substantially outperforming BART-Large (0.2890) and Llama-3.2-1B (0.6630) individually. Although the ensemble incurs a marginal reduction in BERTScore (0.8980) relative to Llama-3.2-1B (0.9555), this trade-off is justified given the critical importance of factual reliability in high-stakes scientific discourse. These findings confirm that anchoring the generative capacity of decoder-only LLMs to verified factual scaffolds effectively mitigates hallucination risk, offering a scalable and reproducible solution for high-fidelity scientific abstract generation.
Geoffrey Antonio Arifin, Andrew Widyanata, Henry Lucky et al.· International Conference on...· 0 citations
SemPOI-RL is proposed, a framework that aligns LLM semantic reasoning with structured sequence generation for interpretable OOT recommendation and consistently outperforms both traditional recommenders and direct LLM baselines, while providing interpretable style attribution across different phases of a trip.
Yunqi Liu, Yang Zhang, Ruixing Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.