Skip to content

Category

computer vision

3,022 papers

Rethinking Object-Centric Representations for Video Dynamics Modeling

Unified Slots (Unified Slots), an unsupervised framework for learning robust and disentangled object-centric representations from videos, is introduced, achieving state-of-the-art performance in unsupervised object-centric video decomposition and tracking.

Amaury Wei, Ismail Nejjar, Olga Fink · 1 citation

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

This work aggregates patch features from all views onto a single equirectangular panoramic canvas, and introduces a spatial pretraining curriculum by procedurally placing patch features of objects at chosen 3D world positions on an otherwise empty canvas, generating on-the-fly supervision spanning a broad range of spat...

Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner · 0 citations

What Should a Streaming Video Model Remember?

This work proposes SelectStream, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget.

Haonan Ge, Yi-Wei Wang, Hang Wu et al. · 8 citations · ⚡1

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

This work introduces Sci-Rho (Science Rhobustness), a dynamic benchmark for visually-grounded STEM problems spanning five subjects and seven languages, comprising 4,242 problem templates crafted by domain experts, including Olympiad medalists.

Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra et al. · 0 citations
#artificial intelligence Preprint Jun 2026

Quantifying and Mitigating Domain Shift in Peach Leaf Damage Classification: Attention Mechanisms and Fine-Tuning Strategies

A transferability baseline for peach leaf diagnosis is established and the adaptation cost of moving from public benchmarks to operational orchards is quantified, establish a transferability baseline for peach leaf diagnosis and quantify the adaptation cost of moving from public benchmarks to operational orchards.

Adrián Cánovas-Rodriguez, Miguel A. González-Illán, Maria Fernanda García-Cruz et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Knowledge-Intensive Video Generation

Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness. We introduce knowledge-intensive video generation (KIVI), where models generate videos from short information-seeking prompts that ask for explanations, procedures, or demonstrations. To...

Chenxu Wang, Mingda Chen · 0 citations

One-Forcing: Towards Stable One-Step Autoregressive Video Generation

One-Forcing is proposed, a simple yet effective approach that augments the DMD objective with an auxiliary GAN loss for high-quality and efficient one-step video generation, and finds that framewise autoregression stabilizes adversarial training, enabling higher-quality generation with substantially fewer training iter...

Jia-Qi Feng, Justin Cui, Yuan-Hao Ban et al. · 13 citations · ⚡3
#artificial intelligence Preprint Open access Sep 2026

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Video large language models (Video LLMs) can achieve strong video-QA accuracy without reliably tracking spatiotemporal dynamics. A model may answer a motion question from static cues, for example, and give the same prediction even after the underlying motion is reversed. Correctness-based reinforcement learning does no...

Dazhao Du, Jian Liu, Jialong Qin et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs) understand not only what happens but also when it happens. Although modern MLLMs describe video content fluently, their timestamp prediction...

Dazhao Du, Liao Duan, Jian Liu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Reducing Hallucination in Multimodal Large Language Models through Hard Grounding Preference Supervision

Hallucination remains a major challenge in vision-language models (VLMs), particularly when linguistically plausible responses are unsupported by visual evidence. We study whether multimodal hallucination can be reduced by concentrating post-training supervision on hard grounding boundaries, where preferred and rejecte...

Qinwu Xu · 0 citations
#artificial intelligence Preprint Open access Sep 2026

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

We propose EverAnimate, an efficient post-training method for long-horizon animated video generation that preserves visual quality and character identity. Long-form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation...

Wuyang Li, Yang Gao, Mariam Hassan et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Cross Modality Image Translation In Medical Imaging Using Generative Frameworks

Magnetic Resonance Imaging (MRI), Computed Tomography (CT), and Positron Emission Tomography (PET) provide complementary information about tissues. Medical image-to-image (I2I) translation enables virtual scanning by synthesizing a target modality from a source one without requiring an additional acquisition. Despite g...

Giulia Romoli, Filippo Ruffini, Francesco Di Feola et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.