Skip to content

Category

natural language processing

6,613 papers

#natural language process... Preprint Open access Oct 2026

Adaptive Information Control for Search-Augmented LLM Reasoning

Search-augmented reasoning agents interleave multi-step reasoning with external retrieval, but uncontrolled retrieval can introduce redundant evidence, saturate the context, and destabilize reinforcement learning (RL). Existing outcome-based RL methods provide only sparse terminal rewards, offering limited guidance for...

Siheng Xiong, Oguzhan Gungordu, James C. Kerce et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Distilling Token-Trained Models into Byte-Level Models

Byte Language Models (BLMs) have emerged as a promising direction for scaling language models beyond tokenization. However, existing BLMs typically require training from scratch on trillions of bytes, making them prohibitively expensive. In this paper, we propose an efficient distillation recipe that converts existing...

Zishuo Bao, Jiaqi Leng, Junxiong Wang et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning

Automated Essay Scoring systems disproportionately penalize high-proficiency English as a Second Language (ESL) learners. We propose Contrastive Learning with Matched Essay Pairs (CL-MEP), a bi-directional alignment strategy. CL-MEP reduces this scoring bias by 39.9% while improving overall accuracy, successfully disen...

Kevin Fan, Eric Yun · 0 citations
#natural language process... Preprint Open access Oct 2026

Mitigating Social Desirability Bias in Random Silicon Sampling

Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social...

Sashank Chapala, Maksym Mironov, Songgaojun Deng · 0 citations
#natural language process... Preprint Open access Oct 2026

Cross-Lingual Summarization as a Black-Box Watermark Removal Attack

Watermarking has been proposed as a lightweight mechanism to identify AI-generated text, with schemes typically relying on perturbations to token distributions. While prior work shows that paraphrasing can weaken such signals, these attacks remain partially detectable or degrade text quality. We demonstrate that cross-...

Gokul Ganesan · 0 citations
#natural language process... Preprint Open access Oct 2026

Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages

Understanding the multilingual mechanisms of large language models (LLMs) provides insight into how they process different languages, yet this remains challenging. Existing studies often focus on individual neurons, but their polysemantic nature makes it difficult to isolate language-specific units from cross-lingual r...

Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Automating MD simulations for Proteins using Large language Models: NAMD-Agent

Molecular dynamics (MD) simulations are essential for understanding protein structure, dynamics, and function, but preparing, running, and analyzing simulations remains time-consuming and error-prone. We present an automated pipeline that combines large language model (LLM) agents with Python scripting and HTMD MCP too...

Omid Barati Farimani, Achuth Chandrasekhar, Amir Barati Farimani · 0 citations
#natural language process... Preprint Open access Oct 2026

Towards Probabilistic Question Answering Over Tabular Data

Progress in question answering (QA) over tabular data has enabled reliable factual retrieval from relational tables. However, many real-world questions are probabilistic, requiring reasoning under uncertainty and latent conditional dependencies that are not directly stored in individual cells. We introduce LUCARIO, a l...

Chen Shen, Estevam Hruschka · 0 citations
#natural language process... Preprint Open access Oct 2026

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implementation. While recent benchmarks primarily focus on converting visual designs to code, we present FullFront, a benchmark designed to evaluate Multimodal Large Language Mod...

Haoyu Sun, Huichen Will Wang, Jiawei Gu et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Rethinking the Relationship between the Power Law and Hierarchical Structures

Statistical analysis of corpora provides an approach to quantitatively investigate natural languages. This approach has revealed that several power laws consistently emerge across different corpora and languages, suggesting universal mechanisms underlying languages. In particular, the power-law decay of correlations ha...

Kai Nakaishi, Ryo Yoshida, Kohei Kajikawa et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

Three tiers of computation in transformers and in brain architectures

Human language and logic abilities are computationally quantified within the well-studied grammar-automata hierarchy. We identify three hierarchical tiers and two corresponding transitions and show their correspondence to specific abilities in transformer-based language models (LMs). These emergent abilities have often...

E Graham, R Granger · 0 citations
#natural language process... Preprint Open access Oct 2026

TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity

Language models based on deep neural networks are vulnerable to textual adversarial attacks. While rich-resource languages like English are receiving focused attention, Tibetan, a cross-border language, is gradually being studied due to its abundant ancient literature and critical language strategy. Currently, there ar...

Xi Cao, Quzong Gesang, Yuan Sun et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.