Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Sep 2026

From Concept Erasure to Style Purification: Contrastive Eigenbases for Artist Style Protection

Text-to-image diffusion models can reproduce specific artists visual styles at extremely low cost, raising copyright and deployment safety concerns about unauthorized style mimicry. Existing model-side protection methods generally follow ordinary concept erasure, emphasizing aggressive deletion or redirection of target...

Tong Zhang, Ru Zhang, Jianyi Liu · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Science

Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the application of tokens to physical simulation remains nascent....

Levi Lingsch, Georgios Kissas, Johannes Jakubik et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Training-Free Global Geometric Association for 4D LiDAR Panoptic Segmentation

Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds or design dedicated modules for instance association. However, these approaches perform redundant point processing and consequently become computationally expensive, yet still ov...

Gyeongrok Oh, Youngdong Jang, Jonghyun Choi et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Hybrid Approach for Enhancing Lesion Segmentation in Fundus Images

Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to improving survival rates, but misdiagnosis or delayed diagnosis can lead to poor outcomes. Despite advancements in AI-based image analysis, diagnosing choroidal nevi in colour f...

Mohammadmahdi Eshragh, Emad A. Mohammed, Behrouz Far et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Geometrically Constrained and Token-Based Probabilistic Spatial Transformers

Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high stakes settings. A model should stay robust under such transformations, expose why a correction was applied, and signal when its input is...

Johann Schmidt, Tom Siegl, Martin Becker et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models

Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly understood. We empirically observe visual elements from different i...

Yeji Park, Minyoung Lee, Sanghyuk Chun et al. · 0 citations
#artificial intelligence Preprint Jun 2025

Gondola: Grounded Vision Language Planning for Robotic Manipulation

A modular manipulation framework that separates high-level planning from low-level control and coupling grounded plan generation with a 3D-based execution policy, this framework achieves state-of-the-art performance on the challenging GemBench benchmark and demonstrates promising transfer to real robots.

Shi-Zhe Chen, Ricardo Garcia, Paul Pacaud et al. · 2 citations
#artificial intelligence Preprint Open access Sep 2026

MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing

Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers a frozen video-diffusion denoising trajectory through concep...

Tong Zhang, Victor Escorcia, Juan C Leon Alcazar et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Beyond Pixels: A Vector-to-Graph Framework for Reliable Schematic Auditing

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the-art models fail to capture topology and symbolic logic in engineering schematics, as their pixel-driven paradigm discards the explicit vect...

Chengwei Ma, Zhen Tian, Zhou Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across...

Hui-Hui Ren, Lei Fan, Henry Pao et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.