Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Oct 2026

Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval

Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-k...

Nickil Maveli, Antonio Vergari, Shay B. Cohen · 0 citations
#artificial intelligence Preprint Oct 2026

Lexicographic Multi-Objective On-Policy Distillation

Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward...

Doseok Jang, Jon Ander Campos, You-Ran Qi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

IntentCoding: Amplifying User Intent in Code Generation

Large Language Models (LLMs) have shown strong capabilities in code generation, but their adherence to fine-grained user intent with multiple constraints remains a significant challenge. Our empirical analysis reveals two key observations: 1) Model performance deteriorates quickly as the number of constraints in the us...

Zheng Fang, Yihong Dong, Lili Mou et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control

Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low monitor readout does not identify whether such an intervention controls behavior. In a code-generation environment whose dominant expl...

Zhe Zhou, Tianhua Tao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Benchmarking Candidate Coverage in Typed Decision Models

Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initi...

Jiawen Lu, Tongtong Wu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Multilingual GSM-Symbolic: What determines capability transfer across languages?

We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all langu...

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard et al. · 0 citations
#artificial intelligence Preprint Oct 2026

KV$^2$: A Self-Refining KV Cache

The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accu...

Johannes Wesch, Danni Liu, Jan Niehues · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through...

Han Cui, Jianhao Yan, Yun Luo et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Peer Influence across Heterogeneous AI Models

When two AI agents disagree, who persuades whom? As multi-agent systems increasingly combine language models of different families and sizes, the answer can determine which judgments survive interaction. Measuring persuasion as the probabilistic shift in an agent's decision after a single exchange with a dissenting pee...

Frida Nøhr Laustsen, Marie Haahr Petersen, Victoria Popa et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Verifiable, Articulable, and Tacit Components of Preference

What makes a short story gripping; a news article newsworthy; or a math proof elegant? These constructs resist articulation or verification; their meaning is at least partially tacit. However, modern AI models are improved primarily via articulated constitutions, rubrics and verifiers (i.e. in RLAIF and RLVR); tacit co...

Alexander Spangher, Sheldon Huang, Andreas Haupt et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel...

David Nadrchal, Monorama Swain, Florian Schmid et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Continual Graph Memory for Mathematical Research Agents

Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generating an enormous volume...

Jun-Yi Zhang, Jin-Xi Yu, E. Jiang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.