Unified Slots (Unified Slots), an unsupervised framework for learning robust and disentangled object-centric representations from videos, is introduced, achieving state-of-the-art performance in unsupervised object-centric video decomposition and tracking.
Amaury Wei, Ismail Nejjar, Olga Fink· arXiv.org· 1 citation
This work aggregates patch features from all views onto a single equirectangular panoramic canvas, and introduces a spatial pretraining curriculum by procedurally placing patch features of objects at chosen 3D world positions on an otherwise empty canvas, generating on-the-fly supervision spanning a broad range of spat...
Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner· arXiv.org· 0 citations
This work proposes SelectStream, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget.
Haonan Ge, Yi-Wei Wang, Hang Wu et al.· arXiv.org· 8 citations· ⚡1
This work introduces Sci-Rho (Science Rhobustness), a dynamic benchmark for visually-grounded STEM problems spanning five subjects and seven languages, comprising 4,242 problem templates crafted by domain experts, including Olympiad medalists.
Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A transferability baseline for peach leaf diagnosis is established and the adaptation cost of moving from public benchmarks to operational orchards is quantified, establish a transferability baseline for peach leaf diagnosis and quantify the adaptation cost of moving from public benchmarks to operational orchards.
Adrián Cánovas-Rodriguez, Miguel A. González-Illán, Maria Fernanda García-Cruz et al.· 0 citations
Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness. We introduce knowledge-intensive video generation (KIVI), where models generate videos from short information-seeking prompts that ask for explanations, procedures, or demonstrations. To...
One-Forcing is proposed, a simple yet effective approach that augments the DMD objective with an auxiliary GAN loss for high-quality and efficient one-step video generation, and finds that framewise autoregression stabilizes adversarial training, enabling higher-quality generation with substantially fewer training iter...
Jia-Qi Feng, Justin Cui, Yuan-Hao Ban et al.· arXiv.org· 13 citations· ⚡3
Video large language models (Video LLMs) can achieve strong video-QA accuracy without reliably tracking spatiotemporal dynamics. A model may answer a motion question from static cues, for example, and give the same prediction even after the underlying motion is reversed. Correctness-based reinforcement learning does no...
Dazhao Du, Jian Liu, Jialong Qin et al.· 0 citations
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs) understand not only what happens but also when it happens. Although modern MLLMs describe video content fluently, their timestamp prediction...
Dazhao Du, Liao Duan, Jian Liu et al.· 0 citations
Hallucination remains a major challenge in vision-language models (VLMs), particularly when linguistically plausible responses are unsupported by visual evidence. We study whether multimodal hallucination can be reduced by concentrating post-training supervision on hard grounding boundaries, where preferred and rejecte...
We propose EverAnimate, an efficient post-training method for long-horizon animated video generation that preserves visual quality and character identity. Long-form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation...
Wuyang Li, Yang Gao, Mariam Hassan et al.· 0 citations
Magnetic Resonance Imaging (MRI), Computed Tomography (CT), and Positron Emission Tomography (PET) provide complementary information about tissues. Medical image-to-image (I2I) translation enables virtual scanning by synthesizing a target modality from a source one without requiring an additional acquisition. Despite g...
Giulia Romoli, Filippo Ruffini, Francesco Di Feola et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.