Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility

Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the output vocabulary via the model's input-output Jacobian, have b...

Mohammad Mosafer · 0 citations
#artificial intelligence Review Oct 2026

Language-Conditioned Token and Reasoning Efficiency in Large Language Models: A Paired Cross-Lingual Study Protocol

Large language models incur language-dependent representation and inference costs, but existing comparisons often conflate input language, assigned observable-trace language, and answer realization. We specify a prospective paired study that separates these interfaces while holding the semantic item, checkpoint, and an...

Gen-Liang Zhu, Chu Wang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Language Model Activations Inhabit Privileged Error-Correcting Basins

Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward "good"...

Matthew Finlayson, Francisco Pernice, Eric Todd et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

InvestigationWorlds: An Agentic Environment for Legal Investigation

We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a court-adopted hypothesis that is treated as ground truth for th...

Albert Yu Sun, Andrew Benard, Sil Hamilton et al. · 0 citations
#artificial intelligence Review Oct 2026

Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas

Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals'decisions. We examine what information helps synthetic respondents predict each individual's later choices, usi...

Khashayar Pourtaheri, Ahmad Zareei · 0 citations
#artificial intelligence Review Oct 2026

Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities

A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier...

Yi-Kai Zhao, Saurabh Pandey, Pradeep Kumar Misra · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skil...

Peter Chen, Wotao Yin · 0 citations
#computer vision Preprint Open access Oct 2026

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression

Multimodal Large Language Models (MLLMs) achieve strong vision-language reasoning but incur large KV caches and high decoding latency with long visual contexts. Existing compression methods rely on observation window attention for stable token importance estimation, yet this aggregation can dilute sparse critical evide...

Tianhao Chen, Yuheng Wu, Kelu Yao et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

LLM Anonymization Against Agentic Re-Identification

Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-identification, yet those same details also carry downstream analytic value of the text. Existing defenses either remove explicit identifiers, perturb text for formal privacy,...

Ziwen Li, Jianing Wen, Tianshi Li · 0 citations
#natural language process... Preprint Open access Oct 2026

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

Ensuring aligned agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. We show that a single aligned agent can propagate cooperative behaviors to unmodified agents purely through natural-language interaction, a phenomenon we term Alig...

Nicole Summer Hsing, Asuka Yuxi Zheng, Yi Zhao et al. · 0 citations
#computer vision Preprint Open access Oct 2026

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-s...

Issa Sugiura, Shuhei Kurita, Yusuke Oda et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or whether the model has the relevant information but follows the wrong inference path. We study this phenomenon as inference misalignment: a mis...

Yangfan Hu, Xuhan Tong, Haoyue Bai et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.