Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology

In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models. In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction. If...

Hyun Do Jung, Jungwon Choi, Soojung Choi et al. · 0 citations
#machine learning Preprint Open access Oct 2026

RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing

Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive...

Wanning He, Yuyao Zhang, Yu-Wing Tai · 0 citations
#machine learning Preprint Open access Oct 2026

REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with prac...

Sheir A. Zaheer, Jihwan Moon, Chan Y. Park · 0 citations
#machine learning Preprint Open access Oct 2026

CETUS: How Far Do Representations Trained on Earth Transfer to Cassini SAR of Titan?

Cassini synthetic aperture radar (SAR) images reveal the dunes, plains, and lake basins of Titan, providing an instance of representations learned from Earth imagery for planetary terrain classification. Cross-domain Evaluation of Earth-to-Titan Transfer Using SAR (CETUS) compares features from DINOv2, DOFA and CROMA w...

Kevin Lee · 0 citations
#machine learning Preprint Open access Oct 2026

Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched...

Xi Ding, Naichen Shi, Jiawei Zhang · 0 citations
#machine learning Preprint Open access Oct 2026

MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation

Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot p...

Suzannah Wistreich, Stephen Tian, Isabella Huang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Identity-Conditioned Score Fusion for Open-Set Person Re-Identification

Robust person re-identification often combines complementary cues such as face, gait, and body shape. While adaptive fusion typically targets query quality, model strength also varies across identities. We introduce identity-conditioned score fusion, a framework that tailors weights to each gallery identity without tra...

Manyi Yao, Jurijs Nazarovs, Eunji Chong et al. · 0 citations
#machine learning Preprint Open access Oct 2026

What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization

Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study...

Ayesh Abu Lehyeh, Jay Hwasung Jung, Safwan Wshah · 0 citations
#machine learning Preprint Open access Oct 2026

Hybrid Cross-Modal Attention Network for Early Breast Cancer Detection in Low-Resource Clinical Settings

Breast cancer is the leading cause of cancer-related mortality among women in Sub-Saharan Africa, where delayed diagnosis results from limited radiology expertise and fragmented clinical data systems. Although deep learning models have demonstrated strong performance in mammographic analysis, most rely solely on imagin...

Simon Hadush Nrea (Mekelle University, Mekelle, Ethiopia) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

On Color Alignment in VAE Latent Spaces and Its Applications

Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance a...

Julian D. Santamaria, Kai Wang, Jes\'us Malo et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Hierarchy-GBP: Accelerating Factor Graph Inference via Abstraction and Recovery

Gaussian Belief Propagation (GBP) is a distributed inference algorithm that passes messages in graphical models, making it attractive for scalable spatial intelligence. However, we find GBP most effective locally: it rapidly smooths message errors that vary sharply between neighbor variables, but corrects global errors...

Yuzhou Cheng, Tom Yates, Ignacio Alzugaray et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Anchor Divergence for Semantic Geometry in Contrastive Learning

This paper concerns how semantic context determines geometry in learned vector representations. Similarity is typically measured using cosine similarity, which provides a single fixed geometry. Semantic similarity, however, is inherently context dependent: two images may be similar because they depict the same object,...

Akash Kannan, Kiho Park, Victor Veitch · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.