Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Oct 2026

Grounding Probes: Generator-Independent Hallucination Detection from Observer Model Hidden States

Detecting responses that retrieval-augmented generation does not ground in its context trades speed against accuracy: surface checks miss paraphrased fabrication, sampling-based methods cost extra generations. Hidden-state probes sit between the two, but every existing one reads the generating model's own activations,...

Michael Rathmayr, Adam Kovacs, Gábor Recski · 0 citations
#natural language process... Preprint Oct 2026

Stance Drift: How AI-mediated Communication Distorts Our Message

Large language models (LLMs) increasingly mediate human communication, from drafting emails to summarizing scientific reports, yet whether they faithfully preserve a speaker's position remains largely untested. We model AI-mediated communication as a two-step generation-extraction pipeline: one LLM produces an argument...

Ling-Chong Liu, Yan-Fei Zhou, Jacob Bien et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Correctness Is a Direction: Geometric Answer Selection in Language Models

Answer correctness is encoded as a recoverable geometric direction in the hidden states of language models. We show that the mean displacement from incorrect to correct answer representations, computed at approximately 70\% of model depth from fifty labeled examples with no parameter updates, yields a scoring direction...

Marcus Armstrong, Navid Ayoobi, Pradham Mummaleti et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Emoji-Emotion Ranking System Using Twitter Data

Nowadays, emojis are often replacing words. Yet computational systems still oversimplify them. Most existing approaches treat emojis as static sentiment indicators and overlook their emotional distributions. In this study, we propose an emoji-aware emotion analysis framework based on a Twitter (X) dataset of 100,000 em...

Danila Khlebokazov, Nurkhan Tashimov, Pakizar Shamoi · 0 citations
#natural language process... Preprint Open access Oct 2026

Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media

Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduce...

Safaruzzaman Shovo, Monowar Islam, Asif Hossain et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Boundaries Agree, Labels Do Not: Intra-Annotator Dynamics as a Kind of Training Data

Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combines annotators into a "ground truth" and measures how well t...

Marharyta Shvets · 0 citations
#natural language process... Preprint Open access Oct 2026

Evidence and Intervention: A Coupled Active-Inference Extension of Rational Speech Act Models

Identical utterance choices can arise from different communicative causes, and identical interpretations can leave different traces in what a listener learns. Rational Speech Act (RSA) models treat interpretation as inference over speaker meaning, but standard one-shot RSA does not intrinsically distinguish these causa...

Yonghyeon Gwon, Elliot Murphy, Chun Kee Chung · 0 citations
#natural language process... Preprint Open access Oct 2026

Evaluating Modeling Approaches for Experience-Level Classification in Job Description

This paper investigates the task of predicting job experience levels in recruitment texts, aiming to automatically identify the qualifications required for positions. Unlike traditional text classification, recruitment texts typically possess explicit internal structures, with different paragraphs playing disproportion...

Celia Liang, Eddie Wu, Shiqi Wang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams

Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the model treats a trailing user comment as part of the pasted te...

Sugam Panthi, Muhaiminul Yeamin, Rabab Abdelfattah · 0 citations
#natural language process... Preprint Oct 2026

Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?

Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned once rather than the name mentioned twice. Pythia models pass...

T. Dahiya, Cole Blondin · 0 citations
#natural language process... Preprint Oct 2026

LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?

Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social rea...

Xin-Yi Liu, R. Khaziev, Dilek Hakkani-Tur et al. · 0 citations
#natural language process... Preprint Oct 2026

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an an...

Jia-Rui Liu, Ren-Jie Tao, Yi-Wei Liao et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.