Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Sep 2026

Image Classifiers are Efficient Self-Supervised Video Representation Learners

We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing v...

Owais Iqbal, Sudipta Sarkar, Shyam Marjit et al. · 0 citations
#machine learning Preprint Sep 2026

Looped Diffusion Transformer

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par...

Yong Xien Chng, Tian-Yi Chen, Wen-Wen Tong et al. · 0 citations
#machine learning Preprint Sep 2026

Spherical Interpolation for Backward-Compatible Multimodal Representations

Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces...

Simone Ricci, Niccoló Biondi, F. Pernici · 0 citations
#machine learning Preprint Open access Oct 2026

Unapologetically Distributed: A Call for Decentralized Document Analysis

Privacy has become an increasingly important concern in the Document Analysis community, to the extent that in many environments such as archives, governmental institutions, and local businesses, the adoption of automation is restricted by legal and policy constraints. While federated learning has often been regarded a...

Adri\`a Molina, Oriol Ramos Terrades, Josep Llad\'os · 0 citations
#machine learning Preprint Sep 2026

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such a...

Shuang Liang, Le-Jun Liao, Shi-Yuan Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages...

Aravindh Mahendran, Michael King, Matthew Koichi Grimes et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Comparative study of adapting pre-trained models for driving behavior video captioning

This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of di...

Sayak Mallick, Philipp Geiger, Augustin Kelava · 0 citations
#machine learning Preprint Sep 2026

MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies

Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared g...

Jing-Qi Wang, Yan Wang · 0 citations
#machine learning Preprint Sep 2026

Fiber-Resolved Microstructure Quantification from Multi-Shell Diffusion MRI using Detection Transformers

Fiber orientation and compartmental microstructure are central to the characterization of white matter tissue in diffusion MRI, yet existing methods either resolve fiber orientations without quantifying microstructure, or quantify microstructure while assuming a fixed number of compartments and a single fiber direction...

Sebastian Endt, Marcus Wirth, Johannes Reinhold Schlund et al. · 0 citations
#machine learning Preprint Open access Oct 2026

BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers

Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG mod...

Zhihang Wu, Zhongqi Wang, Jie Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

DCM-SAM: Defect-Conditioned Mixture of LoRA Experts for NPU-Deployed AM Defect Segmentation

Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone...

Md Mushfiqur Rahaman, Md Mahedi Hasan, Imtiaz Ahmed et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Recovering Off-Policy Supervision for Speculative Decoding

Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preservi...

Jungseob Lee, Chanjun Park, Sugyeong Eo et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.