Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and aut...

Jingyao Zhang, Yun Li, Lu Han · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Think Before You Paint: Recursive Latent Reasoning for Diffusion Models

Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles. We ask how such...

Pawe{\l} Skier\'s, Ma{\l}gorzata Grzanka, Wojciech Masarczyk et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request ac...

Xijie Gong, Tingxu Han, Jiahao Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Plausible Hierarchies to Useful Taxonomies: Evaluating Agentic Harnesses on Customer Feedback

Taxonomies are the symbolic representations through which AI systems organize evidence, aggregate patterns, and answer questions over large document collections. Over customer feedback, the category tree decides how every record is counted and routed, which problems get seen, and which team owns them. Agentic harnesses...

Prabhath Chellingi, Raviraja G, Viraj Bagal · 0 citations
#artificial intelligence Preprint Oct 2026

SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models

Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instru...

Ya-Tai Ji, Zheng-Qiu Zhu, Yong Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures can cause human injury and costly hardware damage. Beyond selecting safe actions, embodie...

Tianruo Rose Xu, Jiawei Ren, Yichi Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is r...

Litao Hu, Yutong Tang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replac...

Rabimba Karanjai, Qun Gu, Hemanth Hegadehalli Madhavarao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Few Bits, One Law: Toward W2A4KV2

Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating sou...

Kai Yi, Tarek Elgamal, Sruthikesh Surineni et al. · 0 citations
#artificial intelligence Preprint Oct 2026

FreeEvolve: Learning to Evolve Beyond Fixed Loops

Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this proces...

Lecheng Kong, Li-Ke Hui, Nikos Kanakaris et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use

Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordina...

Gennadiy Savrasov, Maksim Elistratov, Nikita Gavrilov et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Uncertainty to Action: Learning to Steer LLM Agents

Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and...

Hanwen Li, Jinhao Duan, Guanhua Zhu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.