Skip to content

Category

machine learning

12,457 papers

#artificial intelligence Preprint Open access Oct 2026

Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment

We introduce Score Broadcast and Decorrelation (SBD), a principled framework for broadcast-based credit assignment for general families of differentiable losses. Error broadcast is a biologically plausible alternative to backpropagation that sends output information to hidden layers without weight transport. The Error...

Mustafa Uzun, Mete Erdogan, Cengiz Pehlevan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation

Knowledge distillation (KD) transfers a single scalar prediction from a large foundation model (FM) to compact vertical models (VMs), suffering from diminishing transfer ratio -- the fraction of FM improvement captured by the VM -- as a single scalar cannot convey the rich intermediate knowledge that larger FMs learn....

Hua Zheng, Shali Jiang, Boyang Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

Text-to-image diffusion models generate images by iterative denoising, so their internal layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable features, but most approaches analyze in...

Calvin Yeung, Prathyush Poduval, Ali Zakeri et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or pa...

Subba Reddy Oota, Anant Khandelwal, Khushbu Pahwa et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Constrained latent state modeling: A unifying perspective on representation learning under competing constraints

Learning latent representations from temporal, multimodal, and partially observed data requires specifying what information a latent state should retain, discard, and organize. Existing approaches encode these requirements through heterogeneous objectives, making methods difficult to compare and learned representations...

Gwenol\'e Quellec · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training

Model providers increasingly release the weights of large language models. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. A growing class of defenses aims to make alignment robust to such malicious fine-tuning, but these defenses are typica...

Itay Zloczower, Eyal Lenga, Gilad Gressel et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models

Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive models. Controlled generation methods for DLMs, imported from autoregressive models, apply uniform intervention at every denoising step. We show this uniform schedule is...

Hanhan Zhou, Shamik Roy, Rashmi Gangadharaiah · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which...

Giacomo Spigler · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Metagame of Interpretability and Meta-Attributions

How can an arbitrary attribution method be generalized from first principles to capture interactions? We answer this with the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. We cast the attribution value $\phi_i$ of feature $i$ as a cooperative game among the oth...

Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy

Machine Learning (ML) has transformed many scientific fields, yet key applications still lack standardized benchmarks. Raman spectroscopy, a widely used technique for non-invasive molecular analysis, is one such field where progress is limited by fragmented datasets, inconsistent evaluation, and models that fail to cap...

Mario Koddenbrock, Christoph Lange, Robin Legner et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Packets to Patterns: Interpreting Encrypted Network Traffic as Longitudinal Behavioral Signals

Human behavior is difficult to observe continuously at scale, yet it leaves measurable traces in everyday device use. We test whether encrypted smartphone network traffic---a ubiquitous, always-on, passive sensing modality---can passively capture behavioral patterns related to sleep, stress, and loneliness. We model sh...

Rameen Mahmood, Omar El Shahawy, Souptik Barua et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Clinical Point Cloud Paradigm for In-Hospital Mortality Prediction from Multi-Level Incomplete Multimodal EHRs

Deep learning-based modeling of multimodal Electronic Health Records (EHRs) has become an important approach for clinical diagnosis and risk prediction. However, due to diverse clinical workflows and privacy constraints, raw EHRs are inherently multi-level incomplete, including irregular sampling, missing modalities, a...

Bohao Li, Tao Zou, Junchen Ye et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.