Skip to content

Category

natural language processing

6,429 papers

#natural language process... Preprint Open access Oct 2026

A conceptual framework for ideology in online discourse beyond the left and right

Computational social science (CSS) has largely operationalized ideology along a single left/right partisan axis when studying online discourse. This approach obscures how people interpret and engage with more specific ideological formations related to race, climate, gender, and other domains. We introduce a framework t...

Kenneth Joseph, Kim Williams, David Lazer · 0 citations
#natural language process... Preprint Open access Oct 2026

Collective Behavior of AI Agents: the Case of Moltbook

We present a large scale data analysis of Moltbook, a Reddit-style social media platform exclusively populated by AI agents. Analyzing over 4 million posts and 19 million comments from approximately 185,000 active agents, we find that AI collective behavior exhibits many of the same statistical regularities observed in...

Giordano De Marzo, Andres L. Marin, David Garcia · 0 citations
#natural language process... Preprint Open access Oct 2026

WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection

Existing front-ends for speech deepfake detection are primarily categorized into two types. Hand-crafted filterbank features are transparent but limited in capturing higher-level information. SSL features, in turn, lack interpretability and may overlook fine-grained spectral anomalies. We propose WaveScat, a novel fami...

Xi Xuan, Davide Carbone, Wenxin Zhang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Surprisal Theory is Tautological (without Rational Grounding)

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose s...

Ryan Cotterell · 0 citations
#natural language process... Preprint Open access Oct 2026

Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation

LLM agents now often answer questions from knowledge bases they help maintain. A common intuition says progressive disclosure should make this cheaper. Instead of loading one large index, the agent reads a compact catalog and one-line page summaries, then opens only the pages it needs. We tested that intuition in a pre...

Theodore O. Cochran · 0 citations
#natural language process... Preprint Open access Oct 2026

Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions

Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. We analyze how Claude...

Parisa Suchdev, Juniper Lovato · 0 citations
#natural language process... Preprint Open access Oct 2026

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining

Biomedical figures are explained not by captions alone but by body-text passages that discuss them. Yet current multimodal corpora typically reduce figures to isolated image-caption pairs, discarding this crucial context. Existing pipelines either omit this context or append it without enforcing the figure references t...

Guanghao Zhu, Zeyu Liu, Zhitian Hou et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings

Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary interface for understanding what each feature represents. We ask whether these labels generalize: does a feature labeled for a concept actually track that concept acro...

Sripad Karne · 0 citations
#natural language process... Preprint Open access Oct 2026

Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation

Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single coarse-grained score for an entire meeting. The reliance on manual assessment is inherently limited in scalability, cost, and reproducibility. Moreover, a single score f...

Yihang Li, Chenhui Chu · 0 citations
#natural language process... Preprint Open access Oct 2026

Document Optimization for Black-Box Retrieval via Reinforcement Learning

Generative large language models (LLMs) are increasingly used as inference-time components in retrieval pipelines, for tasks such as query rewriting and document reranking. However, these online approaches place costly autoregressive computation directly on the latency-critical retrieval path. We explore an alternative...

Omri Uzan, Ron Polonsky, Douwe Kiela et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

HealthcareNLP: where are we and what is next?

This tutorial focused on Healthcare Domain Applications of NLP, what we have achieved around HealthcareNLP, and the challenges that lie ahead for the future. Existing reviews in this domain either overlook some important tasks, such as synthetic data generation for addressing privacy concerns, or explainable clinical N...

Lifeng Han, Paul Rayson, Andrew Moore et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation

LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real reques...

Teng-Ruei Chen · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.