Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Sep 2026

FedMAD: Modulation-Aware Directional Aggregation for Federated Learning in Remote Sensing Image Classification

Federated learning (FL) has recently attracted increasing attention in remote sensing (RS) since it enables collaborative model training across decentralized RS image archives without requiring direct access to local data. However, FL performance significantly degrades when the data distributions between clients are he...

Barış Büyüktaş, B. Demir · 0 citations
#machine learning Preprint Open access Oct 2026

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly t...

Mikhail Dereviannykh, Vikram Voleti, Simon Donne et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Multi-Resolution Feature Fusion U-Net for Magnetic Resonance Imaging Segmentation

The segmentation of anatomical structures in medical images and particularly in MRI scans, is essential for clinical diagnosis and monitoring disease progression. While Deep Learning (DL) architectures, such as U-Net and its extensions are very effective in medical image segmentation tasks, they often struggle with pre...

Eirini Cholopoulou, Dimitrios E. Diamantis, Dimitris K. Iakovidis · 0 citations
#machine learning Preprint Sep 2026

Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active f...

Gen-Pei Zhang · 0 citations
#machine learning Preprint Open access Oct 2026

Emergent Object Binding Has a Finite Spatial Horizon

Pretrained Vision Transformers encode whether two image patches belong to the same object. This IsSameObject signal is decodable from frozen patch embeddings at high accuracy, which suggests that object binding emerges from self-supervised pretraining alone. We show that this single accuracy number hides the structure...

Mayank Singal · 0 citations
#machine learning Preprint Oct 2026

SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning

The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained i...

Si-Qi Goh, Cap Dang Xuan Kiet, Tat-Jen Cham et al. · 0 citations
#machine learning Preprint Oct 2026

Two Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize

Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspe...

Yuan Huang, Zi-Han Chen, Run-Bin Zhang et al. · 0 citations
#machine learning Preprint Oct 2026

Dataset Identity, Not Novelty: The Source of an Inflated OOD Detection Gain

A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and returns a score. It fits that score on in-distribution data, and the benchmarks that evaluate it supply a second piece of OOD data for the fitting itself. Some detectors tune a constant on it. Others fit a direction in featu...

Donghoon Lee, Shinjin Kang · 0 citations
#machine learning Preprint Open access Oct 2026

Platonic Task Arithmetic

Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so weight-space task arithmetic stays confined to a single model and cannot cross architectures without a structural correspondence. Drawing on Plato's allegory of the cave, we...

Junghwan Park, Woojin Cho · 0 citations
#machine learning Preprint Open access Oct 2026

Towards Fast and Disentangled Counterfactuals for Visual Foundation Models

Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen f...

Sidney Bender, Benedikt Kunz, Ahmed Zeid et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Signal-Noise Factorization Isolates Nuisance Variation into Removable Subspaces

Recent theoretical work identified fundamental properties of representation geometry that shape inference ability of deep neural networks. These include signal-noise factorization (SNF), the ability to segregate signal from noise, and signal-signal factorization (SSF), the ability to segregate task-specific and task-ir...

Sakin Kirti, Joel Zylberberg · 0 citations
#machine learning Preprint Sep 2026

Curvature Under Attack in hZACH-ViT: Gauge Symmetry, Boundary Saturation, and Adversarial Failure

Curvature is often treated as an intrinsic property of a representation, although its empirical effect also depends on coordinate scale, learned logit temperature, and numerical safeguards. We study this interaction in hZACH-ViT, a compact Vision Transformer with Euclidean, Poincare, and spherical prototype heads. The...

Athanasios Angelakis, M. Gomez-Barrero · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.