Skip to content

Category

natural language processing

6,429 papers

#natural language process... Preprint Open access Oct 2026

Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLMs

Multilingual LLMs answer the same multi-hop reasoning question across languages, but we lack a mechanistic account of whether they share an internal circuit. We identify Bridge Routing Heads (BRH) in two large multilingual LLMs through a three-stage pipeline. The resulting language-specific head sets exhibit near-compl...

Seunghan Kim, Minyeong Choe, Hyunil Kim et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

SpikingVLA: Asynchronous Spiking Vision-Language-Action Models

ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference l...

Jingya Wang, Dehao Zhang, Shuai Wang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain

Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations acro...

Linqi Zhang, Chong Qi, Yan Cheng et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

SAPD: Step-Aligned Privileged Distillation

On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training tra...

Tianle Wang, Jiayu Liu, Ruizhi Zhao et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic acti...

Zhifan Sun, Sebastian Gombert, Jannik Lossjew et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring

Automatic Short Answer Scoring (ASAS) requires models that can score student responses against question-specific criteria while remaining efficient and transferable across rubric sets. We propose RUSPAN, a rubric-conditioned ASAS framework that treats rubric descriptions as semantic label representations. RUSPAN serial...

Zhifan Sun, Sebastian Gombert, Fabian Zehner et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

On-Policy Distillation Teaches New Skills but Not New Knowledge

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently...

Yixuan Tang, Yi Yang · 0 citations
#natural language process... Preprint Open access Oct 2026

RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models

Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the...

Shivam Shukla, Jihye Kim, Shubham Gaur et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers

There are more than 2,000 listed companies on the UK's London Stock Exchange, divided into 11 sectors who are required to communicate their financial results at least twice in a single financial year. UK annual reports are very lengthy documents with around 80 pages on average. In this study, we aim to benchmark a vari...

Nadhem Zmandar, Mo El-Haj, Paul Rayson · 0 citations

Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness

Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definitio...

Yi-Han Li, Han-Yi Zhang, Xiao-Xi Jiang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech

Political discourse often uses rhetorical and persuasive language to frame narratives, influence public opinion, and mobilize audiences. While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and pers...

Rohit Kumar Sen, Anik Chowdhury · 0 citations
#natural language process... Preprint Open access Oct 2026

Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy

When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 juris...

Jiayu Feng · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.