Skip to content

Category

computer vision

2,913 papers

#artificial intelligence Preprint Open access Oct 2026

OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception

Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping approaches enable high-fidelity geometry and efficient open-vocabulary perception, but typically represent semantics as unstructured feature fields that limit object-centr...

Binh Long Nguyen, Kien Nguyen, Clinton Fookes et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation

Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creat...

Tzu-Hsin Hsieh, Ricardo Marroquim · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Graph-Based Recognition of Simulated Train-Driver States From Facial and Upper-Body Keypoints

Driver fatigue poses a significant challenge to railway safety, with traditional systems like the dead-man switch offering limited and basic alertness checks. This study presents a vision-based monitoring system that relies solely on a single front-facing RGB camera and a graph neural network to classify simulated trai...

Olivia Nocentini, Marta Lagomarsino, Gokhan Solak et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Anchor and Adapt: Asymmetric Prompt Adaptation for Few-Shot Industrial Anomaly Detection

In few-shot industrial anomaly detection, the few normal target images provide no direct defect supervision, making anomaly prompts difficult to learn from these samples alone. Some vision-language methods therefore use manually specified descriptions to supply explicit anomaly semantics. However, constructing these de...

Mengyang Zhao, Teng Fu, Haiyang Yu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS...

Xiaojun Jia, Simeng Qin, Yiming Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Energy-Conditioned Noise Schedule and Whitening for Spectral Diffusion

This paper introduces an energy-adaptive noise scheduling and whitening strategy for transform-domain diffusion models. Existing spectral diffusion methods account for the non-uniform statistics of transform coefficients through coefficient scaling, normalization, or frequency prioritization, while the forward diffusio...

Bata Vasic, Bane Vasic · 0 citations
#artificial intelligence Preprint Oct 2026

PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for eval...

Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini et al. · 0 citations
#computer vision Preprint Oct 2026

WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin,...

Turhan Can Kargin, Piotr Kubaty, Ekaterina Rostovskaya et al. · 0 citations
#computer vision Preprint Oct 2026

Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration

Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters is proposed, which achieves the highest mean accuracy among state-of-the-art backpropagates nor updates any model parameters.

Hyeong-Tae Cha, Young D. Kwon, Sung-Ju Lee · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RoboCap: A New Platform for Egocentric Robot Learning

Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrating ergonomic hardware with centimeter-precise 3D algorithms, at a precision that has not been publicly demonstrated. To address this gap, we introduce RoboCap, a 250\,g si...

Grounded Superintelligence, BitRobot · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage

This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters a...

Yi Xia, Ibrahim Khan, Mury Fajar Dewantoro et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

Spatial reasoning benchmarks typically evaluate whether vision-language models can derive the correct answer from a visual observation. Yet in real 3D environments, the observation itself may be unreliable: occlusion can remove task-relevant evidence, while perspective can make visible geometry misleading. Reliable spa...

Yue Zhang, Zun Wang, Han Lin et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.