Skip to content
Open access

Team HITS at SemEval-2026 Task 4: Enhancing narrative text embedding model training with hard negatives generation and self-distillation

2026 · SemEval@ACL · pp. 2679-2688 · 1 citation · ⚡ 1 influential · 16 references
Computer Science

TL;DR

A task-aligned system for SemEval-2026 Task 4 Track B that achieves the best result in the current training phase by introducing "soft label" via KL Divergence.

Abstract

Narrative text embedding is the basis for machines to understand and represent stories. However, it is challenging because it depends on similarities in theme, course of action, and outcomes. To target this challenge, we present a task-aligned system for SemEval-2026 Task 4 Track B. We first use Qwen2.5-32B-Instruct model to generate hard negatives from three narrative dimensions. We then train a Qwen3-Embedding-8B model with a multi-negative contrastive objective and use a teacher model that has the same architecture as the training model. The model achieves the best result in the current training phase by introducing "soft label" via KL Divergence.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

This work investigates continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that is curate from millions of news articles and demonstrates the importance of targeted evaluation in the adaptation process.

Lukas Borggren, Jenny Kunz, Marco Kuhlmann · 0 citations
Jul 2026

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

The Maskability Index (MI) is introduced, a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot generation, providing a principled measure of objective-template alignment.

Ahmad Pouramini, Mahsa Afsharzadeh · 0 citations
Preprint Jul 2026

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 102M Chinese words. The models were evaluated on three tracks: natural language understanding, cognitive alignment, and Hanzi knowledge. There were no restrictions on tokenizers, model architectures, or the number of training epochs. Eighteen teams submitted 28 distinct models, generating 74 result files. The overall-winning team used a DeBERTa-v2 architecture and introduced an auxiliary pinyin-prediction objective during pretraining. Several submissions also explored curriculum-learning strategies and architectural innovations. Overall, the challenge provides a benchmark for advancing data-efficient and cognitively plausible approaches to Chinese language modeling.

Siyuan Song, Zhiheng Qian, Yunhao Zhang et al. · 0 citations
Review Jul 2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check, and human review to remove items solvable from text alone. The 3000 items that pass form the ADQA-Bench evaluation set, spanning music, speech, and environmental audio. The inaugural edition draws 14 teams and 36 submissions across two tracks defined by total parameter count (up to 100B and under 10B). A Chung-Ang University ensemble of MOSS-Audio-8B-Thinking and Qwen3-Omni-30B reaches the top overall accuracy at \pct{58.33}, and a MOSS-only configuration from the same team leads the sub-10B track at \pct{57.30}. Across the 30 submissions with a comparable development score, evaluation accuracy falls by 11.91 percentage points (pp) on average (median 10.91\,pp) on the hidden evaluation split, which is designed to be harder than the development split. The most common building blocks are: the MOSS-Audio-8B-Thinking backbone (13 of 36 submissions), Low-Rank Adaptation (LoRA) fine-tuning on AudioMCQ-StrongAC, and preference or reinforcement-learning objectives -- Group Relative Policy Optimization (GRPO) in five teams, Group reward-Decoupled Normalization Policy Optimization (GDPO) in two. At test time, prompt engineering is near-universal, and majority or choice-permutation voting is common. Every system misses the same set of 233 evaluation items.

Haolin He, Renhe Sun, Zheqi Dai et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.