Skip to content

Category

natural language processing

6,613 papers

#artificial intelligence Preprint Open access Oct 2026

InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers

Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt...

Jonathan Ivey, Aimee Liang, Arthur Y. S. Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation

Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility. While recent LLM-b...

Mengdi Liu, Wenjue Chen, Wenyue Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on th...

Haoyu Zhao, Zhengxu Yu, Zhiyuan He et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretab...

Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models

Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that ge...

Peter Devine, Nick Ryan, Benjamin Sirb et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. W...

Yuan Tian, Bing Hu, Hao Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and...

Xing Zhang, Guanghui Wang, Yanwei Cui et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

UniData: Universal Multimodal Instruction Generation Pipeline

Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they...

Jiaqi Tang, Yi-Feng Wu, Yuting Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-awar...

Gukhyeon Lee, SangKeun Lee · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GameCommBench: A Unified Benchmark and Type-Aware Evaluation for AI-Generated Game Commentary

Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the fun...

Qirui Zheng, Zhengteng Lin, Yunyi Xiao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-...

Daksh Raghuvanshi, Ved Vedere, Yifan Wang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Plan-and-Patch: Diffusion Language Models for Agentic Planning

Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agen...

Syamantak Kumar, Jiang Guo, Hassan Hamad et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.