Skip to content

Category

artificial intelligence

14,158 papers

#artificial intelligence Preprint Open access Oct 2026

DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identif...

Yuxuan Lou, Kai Yang, Geng Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction

Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications. We study the broader problem of...

Yulin Fu (Beijing University of Posts and Telecommunications), Junren Wang (West China Hospital, Sichuan Provincial Engineering Research Center of Intelligent Diagnosis and Treatment of Breast Diseases) et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation. A duplex model supports continuous listening and speaking, but complex reasoning and tool u...

Yingda Shen, Yuxiang Wang, Kunyu Feng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ORDO: Operation-level Round-aware Dynamic Ordering for MIP Presolve

Presolve strongly affects mixed-integer programming (MIP) performance, yet learning-based methods only optimize parameter configurations and cannot express the non-commutative temporal dependencies among actions, whose default order is nearly unique on most domains, yet functionally necessary: artificially shuffling th...

Zehuan Chen, Chunhe Song · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Open-ended Scientific Discovery with Possibilistic Reasoning

Autonomous scientific discovery with LLMs requires generating and testing hypotheses adaptively as evidence accumulates while maintaining statistical validity. Existing anytime-valid methods can handle data-dependent hypotheses, but open-ended discovery poses a deeper challenge: the best discovered hypothesis may still...

Anita Yang, Siu Lun Chau, Tomoya Wakayama et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TaReD: Tool-Aware Recursive Decomposition for Long-Horizon Tasks

Agents combine reasoning with tools to interact with external systems and complete real-world tasks. Early agents typically interleave reasoning and actions along a single execution chain. On complex tasks, this chain becomes unreliable because growing histories obscure intermediate dependencies and allow early plannin...

Wei-Xiang Mao, Zhi-Kai Chen, De-Chuan Zhan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models

Large language model agents are being increasingly deployed as autonomous scientists, designing experiments and inferring mechanistic world models with minimal human oversight. Yet identifiability is often overlooked: when a plateau is reached, the agent needs to know whether it is not yet capable enough or the model s...

Surya Shetty, Ulisses Braga-Neto · 0 citations
#artificial intelligence Preprint Oct 2026

Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?

Small vision-language models may be able to read external evidence yet struggle to obtain it. We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness. A large teacher uses student execution traces to revise reusable content and con...

Min-Hao Fan, Yin-Yi Liu, Jiayu Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade mode...

Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AliO: Output Alignment Matters in Long-Term Time Series Forecasing

Long-term Time Series Forecasting (LTSF) tasks, which leverage the current data sequence as input to predict the future sequence, have become increasingly crucial in real-world applications such as weather forecasting and planning of electricity consumption. However, state-of-the-art LTSF models often fail to achieve p...

Kwangryeol Park, Jaeho Kim, Seulki Lee · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What to Admit and How to Present: Governing Persistent Memory in LLM Agents

Persistent memory can improve personalization in LLM agents but can also induce sycophancy and cross-domain leakage. We distinguish two governance decisions: admission, which determines what recalled information enters the working context, and presentation, which determines how admitted information is expressed. We imp...

Chang Liu, Deliang Ding · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Social Pain Disrupts Emotion-Action Brain-State Dynamics in Adolescents with Non-Suicidal Self-Injury

Non-suicidal self-injury (NSSI) is prevalent among adolescents with depression, but the rapid brain-state dynamics linking social distress to maladaptive behavior remain unclear. We combine an experimental pain paradigm, electroencephalography (EEG) microstate analysis, and interpretable deep sequence modeling to inves...

Ying Xu, Xiaojun Liang, Li Zhang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.