Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation

Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior. Because the same analytics request can be expressed in two distinct target l...

Bhawani Shankar Leelar, Pawan Chorasiya, Davin Hill et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning

Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients. For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a \emph{phase memory}, these records take several times more memory than the model itself. We...

Ahmed Nebli · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking

In post-deployment time, inputs to deep learning models may or may not be adversarially patched. Patch robustness certification on such inputs within a patch bound can verify their label benignity and should retain high prediction accuracy. However, existing smoothing-based and masking-based recovery defenders cannot a...

Qilin Zhou, Zhengyuan Wei, Haipeng Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Strategic Governance of AI Models in Earth Science

AI foundation models pretrained on weather and climate data are increasingly fine-tuned to Earth science tasks well beyond weather forecasting. Their development and adoption are outpacing the scientific community's ability to evaluate them. These models are judged almost entirely by benchmark skill metrics, which meas...

Makoto Kelp, Amirhossein Arzani, Patricia Castellanos et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Freeze the Decoder, Heal the Encoder: Parameter-Efficient Adaptation for SVD-Based KV-Cache Compression

Comparing parameter-efficient fine-tuning recipes under a single, shared learning rate is a common but flawed practice: when the arms being compared have very different trainable-parameter counts, a shared rate can simultaneously depress the larger arms' means and inflate their variance, manufacturing a large, seemingl...

Yufeng Wang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning

In self-supervised contrastive learning, negative pairs are typically constructed using an anchor image and a sample drawn from the entire dataset, excluding the anchor. However, this approach can result in the creation of negative pairs with similar semantics, referred to as "false negatives", leading to their embeddi...

Vicente Balmaseda, Bokun Wang, Ching-Long Lin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling

Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals em...

Qun Dai, Liangjian Wen, Jiang Duan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport

Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior o...

Babak Barazandeh, Connor Swanson, Chinmay Kulkarni et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution

Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel not...

Haohui Wang, Jiahao Xu, Wangzhi Zhan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Prior or Feedback? What an LLM Uses When Adapting Neural Operators

Do LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback? We study this question in neural operator adaptation, where a large language model (LLM) selects fine-tuning configurations under a limited trial budget. Across transfers within and b...

Julian Chan, Javier Mora Jimenez · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models

We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly tra...

Lars Simon, Holger Eble, Manuel Radons · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks

Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on th...

Haoyu Zhao, Zhengxu Yu, Zhiyuan He et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.