Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Sep 2026

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In th...

Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini · 0 citations
#machine learning Preprint Sep 2026

SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation

Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two...

Sara Ghazanfari, Siddharth Garg, P. Krishnamurthy et al. · 0 citations
#machine learning Preprint Sep 2026

ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning

Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adapta...

Tao Hu, Zhi-Nuo Zhou, Xia-Liang Tong et al. · 0 citations
#machine learning Preprint Sep 2026

Visual Branch is What You Need for CLIP-based Class-Incremental Learning

Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name...

Tao Hu, Zheng-He Xie, Jing-Cai Guo et al. · 0 citations
#machine learning Preprint Open access Sep 2026

FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators

Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distributi...

Zhiqi Li, Bo Zhu · 0 citations
#machine learning Preprint Sep 2026

Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study

Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifie...

Bhanu Prakash Vangala, S. Guda, Latha Peddi et al. · 0 citations
#machine learning Preprint Sep 2026

Planetary Feature Fields are Scalable Earth Representations

Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Field...

Arjun Rao, Sebastian Loeschcke, A. Fuller et al. · 0 citations
#machine learning Preprint Open access Sep 2026

The Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns

Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, foc...

Sebastian R\"uckerl · 0 citations
#machine learning Preprint Open access Sep 2026

Are In-Context Images Worth 10 Dimensions?

There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. Howe...

Adhemar de Senneville, Xavier Bou, J\'er\'emy Anger et al. · 0 citations
#machine learning Preprint Sep 2026

TomoTransformer: Towards a Foundation Model for CT Reconstruction

Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detec...

AmirEhsan Khorashadizadeh, Benjamín Béjar · 0 citations
#machine learning Preprint Sep 2026

Principled MAP estimation for inverse problems: bridging the gap between convergence and performance

Pretrained denoisers provide a powerful way to incorporate image priors into restoration algorithms. Plug-and-Play and RED approaches exploit fixed-noise-level denoisers within first-order optimization schemes, with convergence guarantees, but often struggle to achieve high-quality reconstruction on severely ill-posed...

Alexandre Lagier, Valentine Tosel, Anne Gagneux et al. · 0 citations
#machine learning Preprint Sep 2026

Physics-Guided Flow-Map Matching for Precipitation Nowcasting

Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and rare high-intensity structures that matter most. Determinist...

Shunya Nagashima, Takumi Bannai, Makoto Misaizu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.