Skip to content

Category

artificial intelligence

14,237 papers

#artificial intelligence Preprint Open access Oct 2026

Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation

Simulating individual behavior requires large language models (LLMs) to preserve persona traits while adapting to dynamic social contexts. However, general-purpose LLMs often flatten distinct personas, while task-specific tuning suffers from fragmentation and generalization. To overcome these challenges, we organize in...

Liang Wang, Wenxuan Xie, Xinyi Mou et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns...

Xingang Guo, Jing Gu, Brian Jang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate...

Wenyu Du, Stephen Chung · 0 citations
#artificial intelligence Preprint Oct 2026

AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails

Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We prese...

Melissa Kazemi Rad, Si-Hui Dai, Isha Slavin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Agent Plasticity: Measuring Self-Improvement Through Experience

AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future perfor...

Harman Singh, Anton Bakhtin, Rulin Shao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Desp...

Tun-Yu Zhang, Zi-Hao Zhao, Yu-Song Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Humanize: Judgement Engineering for Agentic Coding

Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the...

Sihao Liu, Ligeng Zhu, Zijian Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk

This article develops a failure-mode framework for analyzing how advanced artificial intelligence could contribute to human extinction, irreversible civilizational collapse, or permanent human disempowerment. The central thesis is that catastrophic AI risk does not require consciousness, hostility, or an explicit inten...

Miko{\l}aj Sienicki, Krzysztof Sienicki · 0 citations
#artificial intelligence Review Oct 2026

An Empirical Study of Agent Skills'Downstream Utility

Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Sk...

Yu Cheng, De-Hai Zhao, Zhong-Xin Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning

Compositional generalization remains challenging when language models must combine familiar reasoning operations in unfamiliar ways. The Scenario-Based Commonsense Reasoning Evaluation (SCoRE) 2026 tests this ability on three mixed domains absent from training and requires models to identify the complete set of correct...

Xinchen Xiao · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Accelerating Floating-Point Satisfiability Solving via Gradient Normalization

Satisfiability Modulo Theories (SMT) solvers are foundational to software verification, program analysis, and compiler testing, particularly over the theory of Quantifier-Free Floating-Point (QF_FP). While recent optimization-based SMT solvers have successfully applied gradient descent to continuous relaxations of logi...

Yuanzhuo Zhang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation

Enterprise systems increasingly rely on automated workflows, yet many AI-driven solutions remain brittle under non-stationary conditions, evolving policies, and delayed operational feedback. While reinforcement learning and large language model (LLM) agents offer partial adaptability, they do not by themselves provide...

Sreedevi Pandiyath Viswambaran · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.