Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

AnyGroundBench: A Multi-Domain Adaptation Benchmark for Video Grounding in VLMs

Vision-Language Models (VLMs) have shown strong performance in Spatio-Temporal Video Grounding (STVG), yet they are still evaluated mostly in a zero-shot manner on general-purpose benchmarks of everyday scenes. This creates a critical disconnect from real-world applications in specialized domains, where models inevitab...

Rintaro Otsubo, Ryo Fujii, Reina Ishikawa et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

Vision-language models (VLMs) are increasingly proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction obscured by binary safety benchmark...

Siddhant Panpatil, Arth Singh, Mijin Koo et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on Video Question Answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the faithfulness of predicted evidence supporting those answers remains insufficiently eva...

Linpeng Huang, Weixing Chen, Zexin Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Gefen: Optimized Stochastic Optimizer

AdamW is a default optimizer for deep learning, but its moment states add two parameter-sized buffers to training memory, increasing the cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first momen...

Nadav Benedek, Tomer Koren, Ohad Fried · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CultureScore: Evaluating Cultural Faithfulness in Video Generation Models

As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for assessing cultural faithfulness. Consequently, a model that r...

Anku Rani, Wei Dai, Shravan Nayak et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

Synthetic histopathology image generation addresses patient-privacy concerns and the growing data demands of foundation models. Existing state-of-the-art histopathology generative models use pretrained Vision Foundation Models (VFMs) as conditioning signals. We show this yields conditioning-dominated diversity: on TCGA...

Won June Cho, Daeky Jeong, Hyeongyeol Lim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features...

Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Brain-IT-VQA: From Brain Signals to Answers

Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual question answering (VQA) from fMRI, performance remains limited. Moreover, although...

Roman Beliy, Matias Cosarinsky, Oliver Heinimann et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

vSV-ViT: Variable-size SuperVertex Vision Transformer for Cortical Surface Learning in Alzheimer's Disease

Learning from 3D meshes is challenging because the data reside on non-Euclidean surfaces embedded in 3D space. This is particularly evident in domains such as brain cortical surface analysis, where existing models typically rely on ROI-agnostic, face-based, or fixed-size patches. Such patches can duplicate boundary ver...

Geonwoo Baek, Ikbeom Jang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These attacks, nevertheless, do not assume a realistic threat model, i.e. they either assume access to the...

Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DrawVideo: Grounded and Faithful Multi-Shot Video Generation from Storyboard Keyframe Sketches

Long video generation requires high-fidelity visual synthesis, coherent narrative organization, shot-level structure, and explicit user control. Existing text-to-video methods typically generate videos from a single long-form prompt, making it difficult for creators to directly control character pose, camera compositio...

Chuanzhi Xu, Huiqi Liang, Bang Shi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The TIME Machine: On The Power of Motion for Efficient Perception

Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of self-supervised models trained on next-frame prediction. While these factors have pushed the boundaries of what video models can do, they also introduce th...

Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.