Skip to content

Category

computer vision

2,913 papers

#machine learning Preprint Open access Oct 2026

For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance

The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect t...

Bj{\o}rn Leth M{\o}ller, Bulat Ibragimov, Christian Igel · 0 citations
#machine learning Preprint Open access Oct 2026

HCPN-GCN: Scaling Hierarchical Prototype Networks with Cone Geometry for Continual Graph Learning

Continual Graph Learning (CGL) aims to incrementally learn from graph-structured data while preserving knowledge acquired from previous tasks. A major challenge in this setting is catastrophic forgetting, where learning new tasks degrades performance on previously learned ones. Hierarchical Prototype Networks (HPNs) ad...

Sammuel R. Silva, Vander L. S. Freitas, Gladston Moreira et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs

This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators--global, hand, and head--each guide a corresponding expert br...

Ding-Zhan Nong, Zhi-Hao Ren, Ziqi Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ED3R: Energy-Aware Distributed Disaster Detection via Cooperative Agents in Robotic Systems

Robotics are expected to support environmental monitoring and disaster detection, where decisions must be made under uncertainty, resource limitations, and strict operational constraints. In critical missions, such as wildfires, robots must not only identify hazardous events with sufficient confidence, but also manage...

Lina Magoula, Nikolaos Koursioumpas, Nancy Alonistioti et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Synthetic Benchmarks Overstate Forward-Forward Scaling: Real-Data Limits of Layer-Local Training

Forward-Forward (FF) learning [Hinton, 2022] replaces backpropagation with strictly layer-local goodness updates. Recent FF-CNN work has narrowed the gap to BP on 32x32 benchmarks, raising the question of whether layer-local training is becoming a viable alternative at realistic scale. To probe this rigorously, we deve...

Yucheng Chen · 0 citations
#artificial intelligence Preprint Open access Oct 2026

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evaluation and improvement typically rely on nominal imaginations, which can miss high-impact outcomes of...

Junwon Seo, Sushant Veer, Ran Tian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

Text-to-image diffusion models generate images by iterative denoising, so their internal layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable features, but most approaches analyze in...

Calvin Yeung, Prathyush Poduval, Ali Zakeri et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

Active vision -- where a policy controls its own gaze during manipulation -- has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which...

Giacomo Spigler · 0 citations
#artificial intelligence Preprint Open access Oct 2026

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics

Storyboarding is a core skill in visual storytelling for film, animation, and games. However, automating this process requires a system to achieve two properties that current approaches rarely satisfy simultaneously: inter-shot consistency and explicit editability. While 2D diffusion-based generators produce vivid imag...

Bingliang Li, Zhenhong Sun, Jiaming Bian et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Attention at Rest Stays at Rest: Breaking Visual Inertia to Mitigate Relation Hallucinations

While multimodal large language models demonstrate strong entity-level perception, faithfully grounding relational interactions between objects remains a persistent challenge. Although conventional visual grounding techniques attempt to resolve hallucinations by amplifying visual attention, strengthening overall visual...

Boyang Gong, Yu Zheng, Fanye Kong et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Bi-CamoDiffusion: A Boundary-informed Diffusion Approach for Camouflaged Object Detection

Bi-CamoDiffusion is introduced, an evolution of the CamoDiffusion framework for camouflaged object detection. It integrates edge priors into early-stage embeddings via a parameter-free injection process, enhancing boundary sharpness and preventing structural ambiguity. An optimization objective that unifies spatial acc...

Patricia L. Suarez, Leo Thomas Ramos, Angel D. Sappa · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Parameter-efficient Convolutional Approach for Camouflaged Weed Detection in Multispectral Aerial Imagery

We introduce FCBNet, an efficient model designed for camouflaged weed detection. The architecture is based on a fully frozen ConvNeXt backbone, the proposed Feature Correction Block (FCB), which leverages efficient convolutions for feature refinement, and a lightweight decoder. FCBNet is evaluated on the WeedBananaCOD...

Leo Thomas Ramos, Angel D. Sappa · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.