Skip to content

Category

artificial intelligence

14,192 papers

#artificial intelligence Preprint Open access Oct 2026

Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages

Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage...

Ben Gubler · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RoboJEPA: Scaling Robotic Latent World Models

Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a wo...

Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SciExam for ENSO: Can AI Agents Build Climate Models?

Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for...

Yinling Zhang, Langchen Liu, Dongbin Xiu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing

Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the e...

Yilun Hao, Krishna Sayana, Isabella Ye et al. · 0 citations
#artificial intelligence Review Oct 2026

Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a f...

D. Alexander, C. Kling · 0 citations
#artificial intelligence Preprint Oct 2026

EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt ar...

Python Song, Zhi-Xuan Liang, Kelsey Fu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation requ...

Tan Yu, Alexander Bukharin, Khushi Bhardwaj et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what...

Jinheon Baek, Soyeong Jeong, Yumin Choi et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions

As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the a...

Yi-Zhen Xie, Meng-Yang Liu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations t...

Maverick Morales, Tom\'a\v{s} Dominik, Vermut Gao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning

As multimodal large language models (MLLMs) become more capable and widely deployed, concerns about privacy and safety have become increasingly pressing. Machine unlearning offers one approach to addressing these concerns by removing designated information from trained models while preserving unrelated capabilities. Ho...

Junkai Chen, Yuhao He, Qianshan Wei et al. · 0 citations
#artificial intelligence Preprint Oct 2026

The Handover Problem: Governing Autonomy Transitions in Human-AI Collaboration

Human-machine systems rarely operate at a fixed level of AI autonomy. As operators and AI systems collaborate over time, control must shift: the AI can take on more responsibility when collaboration is stable, maintain its current role when evidence is ambiguous, or return control to the human when conditions deteriora...

Vicente Pelechano, Antoni Mestre, Manoli Albert et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.