Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Expanding LLM Reasoning

Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and measure it at every eligible step for nine models on six bench...

Rian Atri, Evan Luo · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style,...

Abdul Rehman, Jian-Jun Zhang, Xiaosong Yang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential

Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive...

Swadesh Swain, Sanghamitra Dutta · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects

Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, wi...

Bowen Xu, Boyu Chen · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation

Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network...

Murad Hossen, Tasneem Selim, Gurur Gamgam et al. · 0 citations
#artificial intelligence Preprint Oct 2026

When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword...

Qi Zhan, Seoyeon Jang, Zi-Han Dong et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory

As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that l...

Hao Jiang, Zhanyu Guo, Chenwei Wu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-wr...

Cheng-Zhi Luo, Bing Li, Bernard Ghanem · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AraYoungVoices: A Diverse L1/L2 Corpus of Arabic Child and Adolescent Speech

State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes...

Shammur Absar Chowdhury, Zien Sheikh Ali, Houssam Eddine-Othman Lachemat et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Communication Shapes Collective Inference in Self-Adapting LLM Societies: Evidence from Mafia

When does communication help a group identify hidden adversaries, and how does its value change as the group adapts? In Mafia, an informed minority hides inside an uninformed majority whose only evidence is open play. The zero-information game, where each day's vote eliminates a random player, is exactly solved and sco...

Hao-Nan Huang, Joey Xiao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Scaling Verifiable Environments for Long-horizon Work Agents

Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis...

Jiazheng Zhang, Long Ma, Yunxian Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese

This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations...

Hongao Zhu (Department of Linguistics, University of California San Diego), Muxiaoqiao Xu (School of Foreign Languages et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.