Skip to content

Category

artificial intelligence

14,192 papers

#artificial intelligence Preprint Open access Oct 2026

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversa...

Arkajyoti Chakraborty, Aryan Tayal, Ishika Agarwal et al. · 0 citations
#artificial intelligence Preprint Oct 2026

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first pa...

Marek Suppa, Ivan Vykopal, Andrej Ridzik et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations

Autonomous and agentic clients can distribute reconnaissance across many identities so that each request remains valid, low-rate, and benign-looking while the population collectively acquires broad system knowledge. We formalize this threat as Distributed Collective Reconnaissance (DCR) and present SwarmReconGuard, a r...

Vahid Tavakkoli, Kabeh Mohsenzadegan, Kyandoghere Kyamakya · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata

Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block. In multidimensional data, serialization breaks this assumption by scattering spatial neighbors across distant positions in the...

Kimia Kazemian, Menghan Xu, John Thickstun et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Convex-Concave Reinforcement Learning

Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it:...

Shripad Deshmukh, Yaswanth Chittepu, Dhawal Gupta et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models

State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning informati...

Syed Ibrahim Omer, Ginny Y. Wong. Xiangyu Zhao · 0 citations
#artificial intelligence Preprint Oct 2026

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent...

Tobias Braun, Nils Loose, Alexander Herzog et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Talking with Language Models

When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We introduce the artifactual stance, a framework that reconceives human-AI interaction as artifact...

James Ravi Kirkpatrick, Alexandru Radulescu, Rachel Katharine Sterken · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking

Human videos provide rich motion targets for humanoid learning, yet visually plausible references can still produce persistent failures under physics-based execution. These failures reveal where training supervision should change. We present MimicX, a policy-in-the-loop framework that uses execution feedback to refine...

Shuaijun Liu, Chenglong Zhang, Xuhao Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making

In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively quer...

Chenyu Zhang, Rachel Luo, Boyi Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness...

Pavan Maddula · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Visual Memory Attacks Can Persist Through The KV Cache

Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tok...

David Dobre, Leo Schwinn, Gauthier Gidel et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.