The growing demand for on-device large language model (LLM) services on mobile edge devices has driven the adoption of Mixture-of-Experts (MoE) architectures, which scale model capacity with limited computation. Since fine-tuning MoE-based LLMs relies on privacy-sensitive local data, federated learning (FL) offers a na...
Zihan Fang, Qianru Wang, Haonan An et al.· 0 citations
Multimodal web agents that process both screenshots and accessibility trees are increasingly deployed to interact with web interfaces, yet their dual-stream architecture opens an underexplored attack surface: an adversary who injects content into the webpage DOM simultaneously corrupts both observation channels with a...
Haoyu Liu, Dingcheng Li, Lukas Rutishauser et al.· 0 citations
Embedding tables are critical components of large-scale recommendation systems, facilitating the efficient mapping of high-cardinality categorical features into dense vector representations. However, as the volume of unique IDs expands, traditional hash-based indexing methods suffer from collisions that degrade model p...
Ziliang Zhao, Bi Xue, Emma Lin et al.· 0 citations
Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, w...
Kaicheng Xiao, Haotian Li, Liran Dong et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
We study the mean-field Langevin descent-ascent (MFL-DA), a coupled optimization dynamics on the space of probability measures for entropically regularized two-player zero-sum games, together with its associated interacting particle system. For general nonconvex-nonconcave payoffs, Wang and Chizat (COLT 2024) asked whe...
Geuntaek Seo, Minseop Shin, Pierre Monmarch\'e et al.· 0 citations
Asynchronous primal-dual methods for decentralized non-smooth convex optimization often require each node to maintain $\mathcal{O}(d)$ auxiliary variables, where $d$ is its degree. This dependence on degree increases memory requirements and can amplify the effects of stale information, especially in dense networks. Mot...
Anna van Elst, Olivier Fercoq, Igor Colin et al.· 0 citations
The deployment of Artificial Intelligence in high-risk domains, such as finance and healthcare, necessitates models that are both fair and transparent. While regulatory frameworks, including the EU's AI Act, mandate bias mitigation, they are deliberately vague about the definition of bias. In line with existing researc...
Ji\v{r}\'i N\v{e}me\v{c}ek, Mark Kozdoba, Illia Kryvoviaz et al.· 0 citations
Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets. However, recent state-of-the-art self-supervised learning approaches like Masked Autoencoding (MAE), despite their strong representation learning capabilities, lack temporal awareness. In...
Taha Emre, Arunava Chakravarty, Thomas Pinetz et al.· 0 citations
Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tu...
Mingyuan Zhang, Yue Bai, Yifan Wang et al.· 0 citations
Inferring interaction structure from steady-state observations is a central inverse problem when transient trajectories are unavailable. Here we formulate this problem as simultaneous compatibility of a single interaction operator with equilibrium constraints generated by heterogeneous perturbations. We introduce a var...
Optimal transport (OT) and Gromov-Wasserstein (GW) alignment provide interpretable geometric frameworks for comparing, transforming, and aggregating heterogeneous datasets---tasks ubiquitous in data science and machine learning. Because these frameworks are computationally expensive, large-scale applications often rely...
Sanjit Dandapanthula, Aleksandr Podkopaev, Shiva Prasad Kasiviswanathan et al.· 0 citations
Eddy-covariance (EC) flux towers provide in situ measurements of $CO_2$ flux and serve as the ground-truth data for predictive `upscaling' models derived from satellite products. However, many satellites now resolve spatial scales smaller than an EC tower's footprint. We show theoretically that upscaling models trained...
Jacob Searcy, Anish Dulal, Courtney Mathers et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026