Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close th...
Jaewoo Jung, Hyeonseo Yu, Honggyu An et al.· 0 citations
Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexpl...
Jie Yang, Jia-Jun Chen, Jia-Zheng Zhou et al.· 0 citations
Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, h...
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper...
Ze-Xiao Wang, Zi-Hao Zhang, Xu-Dong Wang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for d...
Wenbin Shen, Guoxuan Qin, Guangxu Yao et al.· 0 citations
Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracit...
Wen-Bin Shen, Guo-Xuan Qin, Guang-Xu Yao et al.· 0 citations
Authoring with a text-to-motion generator needs sparse anchors: chosen joints, at chosen frames, at given positions. Meeting them currently costs
a conditioning branch trained for the task, or hundreds of per-clip optimisation steps in the architecture's native variables. In any generator
that decodes a continuous...
Pengcheng Fang, Tengjiao Sun, Xiaoyu Zhan et al.· 0 citations
Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual i...
Xiaoze Liu, Ruowang Zhang, Weichen Yu et al.· 0 citations
Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant content while compressing less informative areas, all while pre...
Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra et al.· 0 citations
Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, without requiring explicit bounding-box annotations. Given their reliance on classification objectives, traditional WSOL methods, like class activation mapping, tend to focus on the...
Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre et al.· 0 citations
Incomplete multi-view clustering (IMVC) is typically evaluated by retraining separate models under different missing-view configurations. Evaluations indexed only by nominal missing rate can overlook differences in observation structure across missing-view protocols. We show that missing-data protocols with identical n...
Haolu Liu, Xiyue Wang, Xuanting Xie et al.· 0 citations
Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the broader sampling design problem, specifically solver selection and scheduling,...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.