Skip to content

Category

computer vision

2,913 papers

#machine learning Preprint Open access Oct 2026

Temporal Visuo-Tactile Learning for Dexterous Grasp Stability

Humans can grasp everyday objects with almost perfect success rates using fingertip tactile feedback, yet much of the robotic grasping literature emphasizes vision-based grasp selection with parallel grippers. In this work, we systematically investigate how high-resolution, dynamic tactile sensing contributes to grasp...

Ken Nakahara, Aleksei Buvailik, Prokhor Kotov et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Global Average Precision for Representation Learning

Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surro...

Bill Psomas, Mohammad Mahdi, Michalis Thomas et al. · 0 citations
#machine learning Preprint Open access Oct 2026

DeepTopoClustering: Unsupervised Derivation of Surface Process Taxonomy from 4D Point Clouds for Topographic Monitoring

4D point clouds acquired by permanent laser scanning (PLS) enable accurate high-frequency monitoring of surface change in dynamic topographic environments. However, existing methods remain limited in organizing detected surface activities into meaningful process types. We propose DeepTopoClustering (DTC), an unsupervis...

Jiapan Wang, Daan Hulskemper, Mathilde Letard et al. · 0 citations
#machine learning Preprint Open access Oct 2026

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet thes...

Arshia Hemmat, Amirhossein Vahidi, Amitis Shidani et al. · 0 citations
#machine learning Preprint Open access Oct 2026

A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods

Self-supervised anomaly detection is a promising paradigm for medical ultrasound, as normal images are often easier to obtain than exhaustive annotations of all possible pathologies. However, most existing evaluations are limited to a single anatomy or task, making it unclear whether models learn a robust notion of nor...

Marco Riedenauer, Daniel Kienzle, Pratik Mayekar et al. · 0 citations
#machine learning Preprint Open access Oct 2026

It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. V...

Samuel Tetteh, Cody Fleming · 0 citations
#machine learning Preprint Open access Oct 2026

TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment

Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference d...

Tianxingjian Ding, Mubarak Shah, Yu Tian · 0 citations
#machine learning Preprint Open access Oct 2026

An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Fragment Adjacency Prediction

This paper addresses the prediction of adjacency between pairs of 2D fragments based on their contours. We improved the two-stage architecture proposed in Beaulac's thesis, in which a rotation-equivariant Siamese convolutional neural network scores pairs of local image windows along the two contours of two fragments. T...

Guillaume Brouillette (Universit\'e du Qu\'ebec \`a Trois-Rivi\`eres, Trois-Rivi\`eres, Canada) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Closing the Loop on Contrail Avoidance with Satellite Verification

Contrails are the thin ice clouds that aircraft leave behind. They cause a large share of aviation's warming, and rerouting the few flights that produce them could avoid much of it. However, an avoided contrail only counts if a satellite can confirm that it never formed, and this check is hard: contrails are one to two...

Spandan Ghose Chowdhury · 0 citations
#machine learning Preprint Open access Oct 2026

emg2face: Expressive Facial Animation with High-Density Surface EMG

Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concern...

Ganidhu Abey, Wendy Greening, Ashika Kamboj et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with e...

Weitian Wang, Shubham Rai, Cecilia De La Parra et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention

Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that ena...

Samrajya Thapa, Daniel J. Quest, Timothy L. Kline et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.