Skip to content

Category

artificial intelligence

14,234 papers

#artificial intelligence Preprint Open access Oct 2026

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replac...

Rabimba Karanjai, Qun Gu, Hemanth Hegadehalli Madhavarao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms

LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is diffi...

Hao-Ran Li, Z. Ge, Xiao-Ming Yuan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AGAR: a reinforcement learning substrate for LLM program evolution

Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remem...

Haoran Li, Zengle Ge, Xiaomin Yuan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Few Bits, One Law: Toward W2A4KV2

Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating sou...

Kai Yi, Tarek Elgamal, Sruthikesh Surineni et al. · 0 citations
#artificial intelligence Preprint Oct 2026

FreeEvolve: Learning to Evolve Beyond Fixed Loops

Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this proces...

Lecheng Kong, Li-Ke Hui, Nikos Kanakaris et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the thre...

Egor Pakhomov, Erik Nijkamp · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev

Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev's judgment inside Stockfish sear...

Mohamad Yazan Sadoun, Sarah Sharif, Yaser Mike Banad · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Training Language Models To Be Coherent Decision-Makers

Reliable decision-making requires more than accurate prediction: a model must preserve its beliefs, apply the relevant utilities, and recognize when the information needed to justify an action is missing. We study whether language models can learn this decision procedure from supervised fine-tuning and generalize it ac...

Khurram Yamin, Xavier Fernandes, Paul Koch et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SpecGuard: Proving a Task Is Broken Before the Agent Cheats

As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent behavior is critical to address. Prior work has shown that agents given such tasks rarely flag the conflict and instead cheat, editing tests or hard-coding expecte...

Param Biyani, Krishnamurthy Dvijotham · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve. This is especially concerning in medicine, where new clinical evidence, updated guidelines, and new therapies can change established practice. Fine-tuning can update the mo...

Yexiao He, Yucheng Tang, Pengfei Guo et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies

Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation modul...

Kai Feng, Guoheng Sun, Zi-Yao Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Finding Blind Spots in AppWorld and WorkArena Task Verifiers

Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value....

Richard Abrich · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.