Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy

When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 juris...

Jiayu Feng · 0 citations
#natural language process... Preprint Open access Oct 2026

ARCS: Towards Precise Text-to-SQL via Structured Disambiguation

As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed...

Yihao Hu, Yanlin Feng, Naoki Otani et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch to...

Radha Gulhane, Quentin Anthony, Beren Millidge · 0 citations
#natural language process... Preprint Open access Oct 2026

Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation

Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a...

Shunsuke Mitsumori, Tomoya Mizumoto, Yusuke Fujita · 0 citations
#natural language process... Preprint Oct 2026

Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model

Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned...

Pulak Kuli · 0 citations
#natural language process... Preprint Open access Oct 2026

Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev

Jev offers an alternative interface for language understanding: given an input and predefined questions, it returns probabilistic decisions rather than free-form responses. Whether this interface can support effective reasoning for text classification against leading LLMs remains an open questions. We introduce Emo-Jev...

Yazhou Zhang, Junhao Yu · 0 citations
#natural language process... Preprint Open access Oct 2026

When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal

Small-data adaptation can improve speech detection while degrading speaker attribution. We study this discrepancy in a released streaming diarizer adapted on 7.5 h of two-party conversation and evaluated across six corpora. Adaptation substantially improves in-domain diarization performance and transfers to an independ...

Mo Yu, Yang Liu, Jing Qian · 0 citations
#machine learning Preprint Open access Oct 2026

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data. Recent work relies on advanced LLMs to synthesize training data, including rationales, labels, and hallucinated claims. However, these methods treat the generator as a static componen...

Shiping Yang, Shining Liang, Weihao Liu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across dive...

Jin Huang, Yutong Xie, Wanli Song et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Latent Performance Profiling of Large Language Models

Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. Evaluating open-source LLMs on leaderboards faces persistent issues such as data contamination, a narrow task scope, and poor alignment with real-world reliabilit...

Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators

Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation. However, industrial recommender systems often benefit from restricting the output space to a constrained subset of items based on business logic (e.g. enforcing content freshness or product category), which standard autoregressive dec...

Zhengyang Su, Isay Katsman, Yueqi Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise...

Ryan Solgi, Parsa Madinei, Jiayi Tian et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.