Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes int...
Accurate assessment of Schistosoma japonicum-associated liver fibrosis is essential for disease management and long-term follow-up in endemic regions. Ultrasound provides non-invasive imaging, but complex local echogenic patterns and anatomical structures make fine-grained grading challenging. Existing deep learning me...
The USAI-Quant benchmark is developed, the first benchmark designed to quantitatively evaluate VLM's reasoning capabilities on built environment metrics via remote sensing imagery, and reveals that current state-of-the-art models consistently fall short on numeric reasoning tasks.
Dong-Dong Wang, Q. Song, Yu-Zhou Chen et al.· 0 citations
Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student, consistently enhances multimodal understanding and reasoning capabilities, and improves both visual understanding and image generation.
Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
OmniMoE-VL, a sparse VLM with a coupled visual-depth routed projector, enables question-dependent visual access while preserving the native visual-token sequence, and complements token-level expert routing in the vision and language stacks.
Long Qian, Bing-Ke Zhu, Jia-Qi Wei et al.· 0 citations
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choic...
Mino Nakura, Sriram Krishna, Yufei Wang et al.· 0 citations
Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hi...
LT-OPD, a training framework for extreme visual-token reduction, is proposed and it is shown that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
Junxian Li, Rui-Xuan Yang, Tian-Ao Zhang et al.· 0 citations
Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter...
Biprodip Pal, Kaushik Roy, Yan-Ming Zhu et al.· 0 citations
Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that l...
Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri et al.· 0 citations
Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minim...
Wen-Hao Zhang, Zhong-Liang Zhou, Shi-Yuan Zhang et al.· 0 citations
Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.