Skip to content

Category

natural language processing

6,429 papers

#machine learning Preprint Open access Oct 2026

PatchBench: Measuring Collateral Damage in Activation Patching

An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmfu...

Alexi Canesse, Mathis Le Bail, Ma\"el Jenny et al. · 0 citations
#machine learning Preprint Open access Oct 2026

YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory

Long-horizon reasoning demands access to earlier information at a manageable generation cost. Full-history attention incurs growing storage and computation, while recurrent compression can lose precise details. Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors...

Huishan Ji, Hua Xu, Weiming Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

When Rank Rises as LLMs Degrade

Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B wit...

Zhaohui Geoffrey Wang · 0 citations
#machine learning Preprint Open access Oct 2026

CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering

Geometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases. In this paper, we propose CHASE (Channel-Aligned Structure Exploitation) to use these structures in practical model design. CHASE covers six...

Wei Wang, Wei Jiang, Ziran Liu · 0 citations
#machine learning Preprint Open access Oct 2026

Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models

We revisit a standard accepted practice in the continuous diffusion language model literature of fixing conditioning prompt tokens clean during training. We make a very simple modification: also noise the conditioning prompt tokens during training. We demonstrate that under this modified training objective, we ac...

Justin Jung · 0 citations
#machine learning Review Oct 2026

Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models

Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as la...

Sourabh Kasliwal, Shubhranshu Singh · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Walk fast but be careful: Understanding Parallel Sampling in Masked Diffusion

In this paper, we use random walks on graphs as a verifiable sandbox for studying parallel sampling strategies in masked diffusion models (MDMs). We train an MDM on random walk samples from a fixed graph. The graph and transition kernel are never shown to the model and serve as latent structure that is both controllabl...

Vansh Bansal, Cholyeon Cho, Syamantak Kumar et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation

Applying pre-trained language models to new domains through full fine-tuning is computationally expensive and prone to catastrophic forgetting. To address this limitation, we introduce a novel parameter-efficient strategy for unsupervised domain adaptation that combines a custom PEFT architecture with mixed-objective t...

Mohammed Rawhani, Dervi\c{s} Karabo\u{g}a, \"Ozkan Ufuk Nalbanto\u{g}lu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

Do large language models' (LLMs') answers to self-report questionnaires predict how they behave? Prior work finds they do not, but it uses human personality inventories, so the gap could reflect borrowed human constructs rather than LLM self-report itself. We test this with a self-report instrument built from LLM-speci...

Juan Manuel Contreras · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute valid actions. A training trajectory records this process as an interleaved sequence of user messages, agent responses, tool calls, etc. Synthesizing sufficiently complex tr...

Hengrui Gu, Xiaotian Han, Kaixiong Zhou · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Handle with CARE: Can LLMs Reproduce How Online Communities React?

Large language models (LLMs) are increasingly used as proxies for computational social analysis, yet faithfully representing the "thick descriptions" (Geertz, 1973) of human communities remains a critical challenge. Current evaluations often reduce social identity to static labels, sidelining how real-world groups navi...

Nuan Wen, Chanbin Lim, Xuezhe Ma · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models

Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive models. Controlled generation methods for DLMs, imported from autoregressive models, apply uniform intervention at every denoising step. We show this uniform schedule is...

Hanhan Zhou, Shamik Roy, Rashmi Gangadharaiah · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.