Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing....

Dani Roytburg, Matthew Bozoukov, Matthew Nguyen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Boosting Large Language Models with Mask Fine-Tuning

The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrati...

Mingyuan Zhang, Yue Bai, Huan Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifi...

Jaturong Kongmanee, Smile Thanapattheerakul · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Reinforcement Learning over Predictive Distributions for LLM Regression

Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression objectives optimize predictions independently, often yielding poor calibration. We introduce Distribution-Aware Reward (DAR), an on-policy reinforcement learni...

Jungsoo Park, Hyungjoo Chae, Ethan Mendes et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks

Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously corrupts both observation channels with a...

Haoyu Liu, Dingcheng Li, Lukas Rutishauser et al. · 0 citations
#machine learning Preprint Open access Oct 2026

RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, w...

Kaicheng Xiao, Haotian Li, Liran Dong et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we iden...

Chenliang Li, Adel Elmahdy, Alex Boyd et al. · 0 citations
#machine learning Preprint Open access Oct 2026

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning

Tasks on complex systems require high-precision numerical computation to support decisions. However, current large language models (LLMs), even with enhanced reasoning capabilities, cannot integrate such computations as an intrinsic and interpretable capability with existing architectures. To this end, we propose Physi...

Jingyuan Fan, Purui Liu, Hengbo Xiao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the...

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations and diffusion/flow spaces. In this work, we introduce Hierarchical Continuous Diffusion La...

Mathias Ollu, Nikos Komodakis · 0 citations
#machine learning Preprint Open access Oct 2026

When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting

Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only t...

Vedant Palit, Florent Draye, Nicolas Zucchet et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SquidAgent: Parallelize Wisely, Coordinate Efficiently

LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden co...

Yexiong Lin, Shanshan Ye, Yu Yao et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.