Skip to content

Category

artificial intelligence

14,237 papers

#artificial intelligence Preprint Open access Oct 2026

Memory-Efficient Expert Routing for Distributed MoE Training

As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routin...

Arnab Kanti Tarafder, Jaume Guasch-Mart\'i, Gokcen Kestor et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Polar: LLM-Powered Synthesis of Real-World Cyber Evidence for Prioritization and Mitigation

Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources. Turning these fragmented observations into timely decisions requires models to connect technical severity with evolving exploitation evidence and available defensive act...

Luoxi Tang, Yuqiao Meng, Ankita Patra et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs

Modern edge Systems-on-Chip (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, each with distinct performance and energy characteristics. Deploying AI inference workloads on them under real-time latency and energy constraints requires jointly mapping workloads to PUs and configuring eac...

Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Aggregating User Preferences while Ensuring Equity, Diversity, and Inclusion using Graph Summarization

Aggregating the preferences of diverse user groups into a collective outcome raises fundamental challenges of equity, diversity, and inclusion (EDI): classical aggregation rules such as Borda and Condorcet have no mechanism to prevent results from systematically favoring majority groups, collapsing onto homogeneous ite...

Adji Marieme Sita Ciss\'e, Malek Mouhoub · 0 citations
#artificial intelligence Preprint Oct 2026

Beyond Successor Accuracy: State Retention for Recursive Self-Improvement in Recommendation

Recommendation recursive self-improvement (Rec-RSI) feeds recommender outputs into subsequent training. Evaluating each round solely through its latest model assumes that the successor consolidates the update, although pre- and post-update models may retain complementary ranking decisions. We term this \emph{distribute...

Jin-Feng Xu, Zhe-Yu Chen, Zi-Yue Peng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator

Large transformer-based models increasingly depend on multi-GPU execution, which requires frequent collective communication among GPUs. Existing communication libraries often rely on many GPU threads to achieve high bandwidth or low latency, resulting in a large streaming multiprocessor (SM)-side resource footprint. Th...

Keyvan Dadashzadeh, Yuehong Zhou, Minyu Cui et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Graph-Based Recognition of Simulated Train-Driver States From Facial and Upper-Body Keypoints

Driver fatigue poses a significant challenge to railway safety, with traditional systems like the dead-man switch offering limited and basic alertness checks. This study presents a vision-based monitoring system that relies solely on a single front-facing RGB camera and a graph neural network to classify simulated trai...

Olivia Nocentini, Marta Lagomarsino, Gokhan Solak et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Demo: Vision-Language Model-Guided Online Calibration of an Electromagnetic Digital Twin

An electromagnetic (EM) digital twin gives mobile robots wireless situational awareness but depends on material conductivities that change with the environment. Online calibration faces initialization sensitivity and measurement travel costs. We demonstrate a vision-language model (VLM)-guided framework using a Unitree...

Zerui Kang, Yishen Lim, Zhouyou Gu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

An Empirical Fault Vulnerability Exploration of ReRAM-based Process-in-Memory CNN Accelerators

Resistive random-access memory (ReRAM)-based Processing-in-Memory (PIM) accelerator is a promising platform for processing massively memory intensive matrix-vector multiplications of neural networks in parallel domain, due to its capability of analog computation, ultra-high density, near-zero leakage current, and non-v...

Aniseh Dorostkar, Hamed Farbeh, Hamid R. Zarandi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Anchor and Adapt: Asymmetric Prompt Adaptation for Few-Shot Industrial Anomaly Detection

In few-shot industrial anomaly detection, the few normal target images provide no direct defect supervision, making anomaly prompts difficult to learn from these samples alone. Some vision-language methods therefore use manually specified descriptions to supply explicit anomaly semantics. However, constructing these de...

Mengyang Zhao, Teng Fu, Haiyang Yu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs

Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP...

Ruizhi Xu, Wei Xu, Sibo Zhu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CrystalJev: thinking fast and slow with atomistic foundation models for materials discovery

Atomistic foundation models triage millions of hypothetical materials but are used as slow simulators, their thresholded energies taken at face value. They are better read as fast decision-makers. CrystalJev queries a frozen interatomic potential once per unrelaxed structure and answers typed questions with calibrated...

Peng Kang, Zhen Li, Yu Liu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.