Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

PRUE: A Practical Recipe for Field Boundary Segmentation at Scale

Large-scale maps of field boundaries are essential for agricultural monitoring tasks. Existing deep learning approaches for satellite-based field mapping are sensitive to illumination, spatial scale, and changes in geographic location. We conduct the first systematic evaluation of segmentation and geospatial foundation...

Gedeon Muhawenayo, Caleb Robinson, Subash Khanal et al. · 0 citations
#machine learning Preprint Open access Oct 2026

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechan...

Kelly Cui, Nikhil Prakash, Shoval Messica et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Improving Mixup Calibration with Wasserstein Distributionally Robust Optimization

In many real-world applications, ensuring the robustness and stability of deep neural networks (DNNs) is crucial, particularly for image classification tasks that encounter various input perturbations. While Mixup-based data augmentation techniques have been widely adopted to enhance the resilience of trained models ag...

Jiaming Hu, Yeping Jin, Debarghya Mukherjee et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework

Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure. A recently proposed developmental, gradient-free learning framework addresses these limitations by learning a discrete, topological model of its inputs through local variation and selecti...

Zeki Doruk Erden · 0 citations
#machine learning Preprint Open access Oct 2026

Stochastic Siamese MAE Pretraining for Longitudinal Medical Images

Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets. However, recent state-of-the-art self-supervised learning approaches like Masked Autoencoding (MAE), despite their strong representation learning capabilities, lack temporal awareness. In...

Taha Emre, Arunava Chakravarty, Thomas Pinetz et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tu...

Mingyuan Zhang, Yue Bai, Yifan Wang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Co-Evolving Paths and Flows via Path-Flow Alignment

We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path l...

Zeyu Michael Li, William Xingxu Chen, Xiang Cheng · 0 citations
#artificial intelligence Preprint Oct 2026

FedDermaSeg: Federated Learning for Dermatological Image Segmentation

Skin cancer is a major global health concern, and early detection and accurate lesion delineation are important for effective diagnosis and treatment planning. Automated skin lesion analysis can assist dermatologists, with lesion segmentation serving as a fundamental step in computer-aided diagnostic systems. Conventio...

Anabik Pal, Ganesh Patidar, Bikash Santra · 0 citations
#machine learning Preprint Open access Oct 2026

Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-...

Dan Ben-Ami, Kobi Cohen, Chaim Baskin · 0 citations
#machine learning Preprint Open access Oct 2026

Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations

Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the p...

Yongsheng Luo, Wengan He, Yu Li et al. · 0 citations
#machine learning Preprint Oct 2026

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on ou...

Shuo Yang, Chang-Bai Li, Lin-Lin Yang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data

Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actual...

Robin Young · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.