Skip to content

Category

artificial intelligence

14,234 papers

#artificial intelligence Preprint Oct 2026

Efficient Reasoning with Flow Language Models

Flow Language Models (FLMs) have emerged as a continuous-state alternative to discrete diffusion language models, yet the role of their continuous representations in reasoning remains unclear. We investigate this question by comparing the reasoning efficiency of FLMs and discrete diffusion models, measured by solution...

Han-Ru Bai, Faissal Izermine, Oscar Davis et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Let the Library Speak: Self-Advertised Method Selection for Formal Proving

LLM-based formal provers can retrieve relevant lemmas and prior proofs, but relevance alone does not say whether a mathematical method can be used on the current theorem. A method has prerequisites, a target, an intended action, and obligations that its use leaves to prove. Methods that look equally related to a theore...

Xiaopeng Yuan, Suijin Wang, Yanli Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Plausible Hierarchies to Useful Taxonomies: Evaluating Agentic Harnesses on Customer Feedback

Taxonomies are the symbolic representations through which AI systems organize evidence, aggregate patterns, and answer questions over large document collections. Over customer feedback, the category tree decides how every record is counted and routed, which problems get seen, and which team owns them. Agentic harnesses...

Prabhath Chellingi, Raviraja G, Viraj Bagal · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis

Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the e...

Yizhi Song, Hang Ni, Weijia Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing Grap...

Ruochi Li, Jianzhe Lin, Haoxuan Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranke...

Yufeng Li, Shuxin Li, Zhenhua Xu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models

Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instru...

Ya-Tai Ji, Zheng-Qiu Zhu, Yong Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The AI Evaluation Ecosystem

AI evaluation shapes the decisions of model providers, users, funders, and regulators. We argue that designing valid benchmarks requires contextualizing design choices in the dynamics of this ecosystem of actors. We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic act...

Yash Dave, Sang T. Truong, Serena Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures can cause human injury and costly hardware damage. Beyond selecting safe actions, embodie...

Tianruo Rose Xu, Jiawei Ren, Yichi Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents

Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two forms of agentic system, Workflows and Age...

Kefan Liu, Fengning Ou, Yelin Luo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is r...

Litao Hu, Yutong Tang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Trajectory Abstraction for the Science of Language Agent Behavior

Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constraine...

Tianqiang Yan · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.