Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Introducing Code-Switched Contexts to Cognitively-Inspired Bilingual Model Training

During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While intr...

Zhuojing Huang, Luise Pohlmann, Lisa Beinborn · 0 citations
#natural language process... Preprint Open access Oct 2026

From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a tas...

Quanyu Long, Xiao Chen, Jianda Chen et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

TrustMI: Causally controlling how assistants trust their users

Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply with harmful requests or act on malicious instructions encoun...

Th\'eo Lasnier, Romain Froger, Maxence Lasbordes et al. · 0 citations
#natural language process... Preprint Oct 2026

D-Loop: Looped Diffusion Drafting for Speculative Decoding

Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighbor...

Ke-Cheng Chen, Yu-Yang He, Cheng Gong et al. · 0 citations
#natural language process... Preprint Oct 2026

Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models

With the rapid development of artificial intelligence, the emergence of various Large Language Models (LLMs) has created a rich model ecosystem. However, this also brings a key challenge: how to select the optimal model for a specific user query. LLM routing addresses this need by dynamically assigning queries to the m...

Yao Lu, Zhai-Yuan Ji, Ya-Xin Gao et al. · 0 citations
#natural language process... Preprint Oct 2026

Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can...

Jie Wang, Shi-Wei Luo, Qi Zhang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Can Language Models Learn to Reject Their Own Bad Reasoning Steps?

Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's recoverability as the probability that the frozen generator ca...

Siheng Xiong, Xiaoze Liu, Yiqiao Jin et al. · 0 citations
#natural language process... Preprint Oct 2026

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and...

Sarim Hashmi, Mukul Ranjan, Abdelrahman W. A. Elsayed et al. · 0 citations
#natural language process... Preprint Oct 2026

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fix...

Zhuo-Kun Chen, Xi Lin, Xi-Yu Wu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Plan Canvas: Fixed Reasoning Regions for Continuous Language Flows

Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start is therefore unknown during denoising, and the model has to d...

Miaohe Niu, Pengxiang Li, Jingbo Zhu et al. · 0 citations
#natural language process... Preprint Oct 2026

Adaptive Utilization of Low-Rank Adaptation via Conditioned Gating

Low-Rank Adaptation (LoRA) achieves parameter-efficient fine-tuning by constraining model updates to a low-rank subspace and has been widely used in practice. However, LoRA typically employs a shared low-rank update across tokens, which limits its ability to fully exploit the adaptation subspace for tokens from differe...

Guang Yang, Chang-Hao Guan, Chao Huang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

MedicalHarness: A Controlled Evaluation of LLMs and Agent Harnesses on Medical Tasks

LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is therefore a property of a model--harness pair. For medical agen...

Ziqing Wang, Lili Zhao, Kaize Ding · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.