Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

Language models can notice an impossible engineering problem yet still report it as solved

Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key...

Shaoliang Yang, Jun Wang · 0 citations
#artificial intelligence Preprint Oct 2026

Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model...

Sheng-Tao Wen, Xiang Chen, Yu Tian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential dec...

Zhongxiang Sun, Jiahao Yan, Hongkang Zhao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers

LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective p...

Pedro Tabacof, Sagar Joglekar · 0 citations
#artificial intelligence Preprint Oct 2026

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero...

Chubin Zhang, Zheng-Lin Wan, Xing-Rui Yu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge

Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and repo...

Abdul Basit, Muhammad Abdullah Hanif, Muhammad Shafique · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation

Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing th...

Alireza Jafari, Arman Adibi, Mohammad Ghavamzadeh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Mining Agent Skills from Production Traces

Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whether a run has succeeded may be unavailable. We study how the...

Yue Ran Kang, Colton Mikolajczyk, Chhaya Methani et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate

Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variatio...

Khanh K. Nguyen, Van Dai Do, Tien-Thuy Nguyen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DelegationBench: Measuring When AI Agents Should Ask Before Acting

AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test...

Shiva Pochampally · 0 citations
#artificial intelligence Preprint Oct 2026

TeleTune: Evolving Agent Skills From Offline Telemetry

Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) N...

J. Chen, Elias Stengel-Eskin, Yan Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Look Before You Leap: Thermodynamic Arbitration of Parametric and Non-Parametric Knowledge in LLM Agents via Self-Regulating Memory Architectures

The architecture of modern LLMs consists of a profound cognitive polarization. LLMs possess implicit intuition encoded in their parameters, yet rely on a disconnected, explicit mechanism to access the outside world. Agentic frameworks have not bridged this gap; instead, models are often compelled into pathological "ind...

Akash Das, Ishan Roy · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.