Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Gated Memory: Admission-Controlled Memory Formation for Conversational AI

Personalized conversational AI relies on long-term memory systems that extract facts from user utterances and store them in persistent vector stores. Despite progress in retrieval, deduplication, and lifecycle management, the formation stage, the moment a fact is first written to storage has received almost no principl...

Preeti Saraswat, Divya Neelagiri, Ajay Manoj · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Neuro-Memory Fuzzy Inference System for Mimicking Human-like Car Following Behavior

This study presents the Neuro-Memory Fuzzy Inference System (NeMeFIS), a hierarchical machine learning architecture that asymmetrically models acceleration and deceleration in car following behavior by integrating five human memory types procedural, working, episodic, semantic, and declarative. By linking external vari...

Nazmul Haque, Md Asif Raihan. Md. Hadiuzzaman · 0 citations
#artificial intelligence Preprint Oct 2026

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement lea...

Heng Li, Yong Zhang, Ning Cheng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis

Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-langu...

Weiwei Ma, Xiaobing Yu, Peijie Qiu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning

Supervised fine-tuning (SFT) equips large language models (LLMs) with specialized capabilities, but often comes at the cost of forgetting the general capabilities of their parent models (i.e., the pretrained models before fine-tuning). This trade-off is especially limiting for queries that require both specialized and...

Ke-Nan Tang, An-Dong Hua, Cheng-Xuan Qian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Stability-Plasticity Balance via Singular-Vector Selection in LLM Continual Learning

Domain-specific continual adaptation of LLMs risks catastrophic forgetting, creating a fundamental tension between acquiring new capabilities and preserving those learned during pretraining. PEFT mitigates this problem by restricting the number of trainable parameters, but existing methods lack a principled unit for de...

Lingxiang Wang, Hainan Zhang, Liang Pang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Emergent Inverse-Depth Scaling From Nonlinearity In Attention

Scaling laws describe power-law improvements in model performance with dataset size and parameter count, yet their underlying mechanisms are not fully understood. To explain the parameter count scaling, existing theory posits power-law scaling with model depth. In linear-attention models, this scaling is tied to a powe...

Zirui Peng, Yizhou Liu, Ziming Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents

Tool-using LLM agents violate the policies they are deployed to enforce, often silently. Prior defenses hand-write rules, query an LLM verifier per action, or compile policies through heavyweight formal machinery. Naive compilation fails: extracted rules block the tool satisfying their own precondition, or read argumen...

Min-Young Yu, Tony Kim, Jang Won Choi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Mid-Training Language Models on Raw Video

Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded in...

Jaedong Hwang, Xiaoqian Shen, Ernie Chang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Optimizing Large Language Models with Chained LMOs

Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chai...

Sungyoon Kim, Kaan Ozkara, Youngsuk Park · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large lan...

Biao Xiang, Ali Eshragh, Yuexing Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TRACE: A Governance Framework for Measuring Explainability Debt in Production AI Systems

Production AI systems deployed in high-stakes domains accumulate a governance liability that existing monitoring frameworks fail to detect: the progressive inability to explain individual decisions when regulators, auditors, or affected individuals demand accountability. We introduce TRACE (Transparency, Risk, Accounta...

Harish Kant Pathak · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.