Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

Why Convolution Still Matters: Evaluating Inductive Biases in Cryospheric Image Classification

Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unclear. In this work, we revisit a benchmark Greenland Ice She...

Chhaya Kulkarni, Emam Hossain · 0 citations
#machine learning Preprint Open access Oct 2026

Evaluating Zone-Guided Front Extraction for Glacier Calving-Front Delineation in SAR Imagery

Automatic calving-front delineation from synthetic aperture radar imagery is challenging because the front is a thin and often ambiguous boundary between glacier ice, ocean, and surrounding rock or terrain. The CAlving Fronts and where to Find thEm (CaFFe) dataset provides both binary calving-front masks and broader se...

Chhaya Kulkarni, Emam Hossain · 0 citations
#machine learning Preprint Open access Oct 2026

Dynamic Quadtree Tokenization and Transformer for Adaptive Mesh PDE Forecasting

The quadratic attention cost of Vision Transformers (ViTs) forces a trade-off between spatial resolution and rollout horizon, particularly for fine-scale PDEs where shocks, reaction fronts, and material interfaces occupy small, evolving regions of the domain. Conventional neural surrogates also lack mechanisms to adapt...

Yilin Zhuang, Noah Zambrano, Karthik Duraisamy · 0 citations
#machine learning Preprint Oct 2026

SUAVE: Unified Video-Action Models via Masked Diffusion

Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion...

Rhythm Syed, Jean-Pierre Mercat, Sedrick Scott Keh et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Masked Privileged-Information Distillation for Multimodal Skin Lesion Classification Under Missing Clinical Metadata

Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings additionally require computational efficiency. We address these chal...

Anirban Barua, Md Mahir Abrar Khan, Ayman Iktidar et al. · 0 citations
#machine learning Preprint Oct 2026

StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search

Recovering executable CAD programs from 3D meshes is challenging due to the compositional nature of CAD construction and the interaction between discrete modeling choices and continuous parameters. Many learning-based methods predict complete programs in a single pass and rely predominantly on sketch-extrude representa...

Ghadi Nehme, Faez Ahmed · 0 citations
#machine learning Preprint Open access Oct 2026

LoRA Direction Extraction for Controllable Light Toggling in FLUX.1 Kontext

We propose a fine-tuning method for flow-matching diffusion models aimed at realistic artificial light modeling without the need for a large training dataset. We address the task of controllable interior image editing, where the goal is to turn artificial light sources on or off while preserving the scene geometry, obj...

Petr Golenderov, Dmitry Mazyar, Natalia Sovpel et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Llama 3 Herd of Models

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a...

Aaron Grattafiori (Jack), Abhimanyu Dubey (Jack), Abhinav Jauhri (Jack) et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Broaden Your Views for Self-Supervised Video Learning

Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views are similarly extracted by cropping and augmenting the resulting crop. However, these methods miss a...

Adri\`a Recasens, Pauline Luc, Jean-Baptiste Alayrac et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Detecting Nighttime Anomalies from NASA Black Marble Using a Generalized Spatio-Temporally Robust Framework of Machine Leaning Ensembles

Nighttime lights from NASA's Black Marble product suite capture thermal and light emission signals from anomalous events including fires, volcanic eruptions, and gas flaring. Existing detection approaches rely primarily on thermal bands, limiting sensitivity to weaker signals. We propose a novel machine learning framew...

Srija Chakraborty · 0 citations
#machine learning Preprint Open access Oct 2026

Improving Proactive AI Assistance with Hierarchical Procedural Understanding

Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporti...

Jin-Seop Lee, TaeYeon Won, SeongJun Jung et al. · 0 citations
#machine learning Preprint Oct 2026

LeAVJEPA: A Minimalist Architecture for Audio-Visual Self-Supervised Learning

Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes...

B. Robson, Santeri Mentu, Wen-Shuai Zhao et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.