Skip to content

Category

artificial intelligence

14,233 papers

#artificial intelligence Preprint Open access Oct 2026

LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses li...

Jun Zhao, Leiming Fu, Yanbo Wen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents

Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations:...

Wenjie Liao, Liangjie Zhao, Zehong Cao · 0 citations
#artificial intelligence Preprint Oct 2026

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

Memory-augmented reinforcement learning strengthens LLM agents'ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries ac...

Yuyao Ge, Yi-Wei Wang, Yu-Chen He et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Expert-Guided Proof Search to Automated Open-Problem Solving

Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-read...

Adri\'an Z\'ame\v{c}n\'ik, Mat\v{e}j Kripner, Martin Kouteck\'y et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers

The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be selective (e.g., reject l...

Evgenia Ilia, Wilker Aziz · 0 citations
#artificial intelligence Preprint Open access Oct 2026

System Switch: When Should a Fast Decision Model Stop and Think?

Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a...

Gian Luca Bailo · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Cost-Efficient Theorem Proving via Agent Orchestration in Program Verification

Program verification establishes software correctness through machine-checkable proofs constructed in theorem provers. It's a guarantee especially valuable for code generated by large language models (LLMs), which is fluent but carries no assurance of correctness. Almost all existing provers, however, pursue pass rates...

Shuangjie Yao, Nikolaus Holzer, Mark Paul Santolucito et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Shared and structured inputs undermine collective random choice by reasoning AI agents

Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems. Behavioural tests across six reasoning models uncovered threshold and divisibility rules used in identifier-based choices. For threshold-following GPT-6 Sol and Gemini 3.8 Flash, single-ag...

Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari · 0 citations
#artificial intelligence Preprint Oct 2026

MeshSIPP: Efficient Lattice Planning in Dynamic Environment

Autonomous navigation in dynamic environments requires computing spatiotemporal trajectories that satisfy non-holonomic motion constraints. When the trajectories of the moving obstacles are predictable or known, a promising approach is to rely on the combination of state lattices constructed from precomputed feasible m...

Marat Agranovskiy, Konstantin S. Yakovlev · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Automatically Building and Updating a Knowledge Graph of MLIP Models

Complementing the many efforts in providing semantic representations of concepts, notions, and entities in materials science, we report and illustrate a process by which we can automatically build a knowledge graph of the fast evolving field of machine learning applied to the prediction of material properties, focusing...

Alexis Beer, Liudmyla Klochko, Mathieu d'Aquin · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request ac...

Xijie Gong, Tingxu Han, Jiahao Zhang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment...

Miao Yu, Hao-Hao Huang, Luiza S. B. Yuan et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.