Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Open access Oct 2026

Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential t...

Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe · 0 citations
#machine learning Preprint Sep 2026

Uncertainty-Aware RL-Controlled Adaptive 3D Mapping

Voxel-based volumetric mapping is fundamental to 3D reconstruction, yet fixed-resolution grids remain inherently inefficient - wasting memory in uniform regions and losing detail in complex ones. Existing adaptive methods, such as MAP-ADAPT, partially address this by varying resolution based on geometry and user-define...

Alpay Ozkan, T. Aydın, M. Pollefeys et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

FuncBridge: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning

While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones -- a gap we formalize as functional generalization. Functionally equivalent tools share visually recognizable functional intent, such as where contact can occur an...

Chuhao Zhou, Liquan Wang, Shuxin Cao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Vision-language models for chest radiography do not always need the image

Vision-language models that answer questions about chest radiographs are evaluated by their accuracy on labels derived from radiology reports. High benchmark accuracy is often interpreted as evidence that the model uses the image. A model that answers from the finding named in the question can score as well as a model...

Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DynGhost: Temporally-Modelled Transformer for Dynamic Ghost Imagings

Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensity measurements. While deep learning approaches have achieved promising results on static scenes, two critical limitations remain unaddressed: existing architectures fail...

Vittorio Palladino, Ahmet Enis Cetin · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression kn...

Roy Xie, Dan Friedman, Donghan Yu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is...

A. Sophia Koepke, Daniil Zverev, Shiry Ginosar et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning

Imitation learning has enabled robots to acquire complex visuomotor manipulation skills from demonstrations, but deployment failures remain a major obstacle, especially for long-horizon action-chunked policies. Once execution drifts off the demonstration manifold, these policies often continue producing locally plausib...

Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Do Vision Language Models Understand Human Engagement in Games?

Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vision--language models (VLMs) can infer such latent psychological states from visual cues alone. Using the GameVibe Few-Shot dataset across nine first-person shooter games, we eval...

Ziyi Wang, Qizan Guo, Rishitosh Singh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ProtoDCS: Towards Robust and Efficient Open-Set Test-Time Adaptation for Vision-Language Models

Large-scale Vision-Language Models (VLMs) exhibit strong zero-shot recognition, yet their real-world deployment is challenged by distribution shifts. While Test-Time Adaptation (TTA) can mitigate this, existing VLM-based TTA methods operate under a closed-set assumption, failing in open-set scenarios where test streams...

Wei Luo, Yangfan Ou, Jin Deng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Modeling The Object Representations Underlying Human Physical Reasoning

Humans appear to represent objects when reasoning about physics with coarse, volumetric "bodies" that smooth concavities, trading fine visual detail for efficient physical predictions. Yet, the structure of these representations remains largely unknown. Segmentation models, in contrast, are trained for pixel-accurate m...

Andrey Gizdov, Andrea Procopio, Lorenzo Caputi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly...

Yang Xing, Jiong Wu, Savas Ozdemir et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.