Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Evaluating VQA in Vision Language Models using Cooperative Principles

We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatG...

Monika Shah, Sudarshan Balaji, Somdeb Sarkhel et al. · 0 citations
#natural language process... Preprint Oct 2026

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psych...

Longwei Cong, Sonja Hahn, Sebastian Gombert et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, cre...

Baohang Li, Xiaocheng Feng, Yichong Huang et al. · 0 citations

How Robust Is Multimodal Claim Verification to LLM Rewriting?

LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. W...

Yun-Ang Wu, Xanh Ho, André Greiner-Petter et al. · 0 citations
#natural language process... Preprint Oct 2026

Text-Centric Post-Training for Omni-Modal Reasoning

Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of pe...

Zi-Yang Cheng, Yu-Hao Wang, Hong-Cheng Liu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Automatic Evaluation of Mental Health Stigma in Online Communication

Mental health stigma has profoundly harmful impacts but its complexity makes it difficult to evaluate. Stigma may involve explicit derogation, but also subtler forms of blame, fear, paternalistic pity, social distancing, structural exclusion, and discrimination. We introduce a theory-grounded benchmark for automatic ev...

Naomi Baes, Jemima Kang, Nick Haslam et al. · 0 citations
#natural language process... Preprint Oct 2026

AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration

Document databases such as MongoDB are core infrastructure for modern applications, and natural-language interfaces to them---text-to-MQL---would let non-experts query complex, semi-structured data without mastering the query language. Progress on this task depends on high-quality benchmarks, which are most practically...

Hy Nguyen, Nabi Rezvani, Robin Vujanic · 0 citations
#natural language process... Preprint Oct 2026

EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models

Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quan...

Zeeshan Memon, Yi-Qi Su, Kai Shu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Beyond Correctness: Resolving Underspecification in Agentic Text-to-SQL

Agentic Text-to-SQL systems can interact with users to clarify underspecified queries before generating SQL. However, a correct execution result does not necessarily imply that the agent has adequately resolved the underlying underspecification: the agent may silently make unverified assumptions that happen to match th...

Wen-Zhi Li, Yue Gong, Konstantinos Kanellis et al. · 0 citations
#natural language process... Preprint Oct 2026

Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise

Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in"the capita...

Zi-Yi Ni, Peng-Jia Zou · 0 citations
#natural language process... Preprint Oct 2026

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translatio...

Hieu Hoang, Amittai Axelrod · 0 citations
#natural language process... Preprint Open access Oct 2026

Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We eval...

Fahmid Shahriar Iqbal, Ritam Dutt, Soumitra Das et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.