Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is s...
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving da...
Parthib Roy, Yashpal Tandon, Marcus Blennemann et al.· 0 citations
This work is the first to unveil the critical role of the null space and harness it for model optimization, and uses a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
Hong-Bo Ma, Sansheng Cao, Jia-Jun Fan et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring...
Yue-Dong Tan, Lei-Tao Qi, Yu Liu et al.· 0 citations
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rol...
Chen-Jian Gao, Zhi-Hao Hu, Jian-Qi Ma et al.· 0 citations
MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.
Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem et al.· 0 citations
Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline rec...
Fan Zhou, Shuairan Chen, Mengying Zhang et al.· 0 citations
This work considers pixel-level attention schemes and shows that the resulting models generally outperform those relying on larger patch sizes, and provides practical guidance for designing models for pixel-level regression tasks on medium-resolution satellite imagery.
Sven Ligensa, Jan Pauls, Karsten Schrödter et al.· 0 citations
Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet cont...
Kelly McConvey, Sajad Ebrahimi, Nima Jamali et al.· 0 citations
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficul...
Xuan-Yu Zhu, Yan Bai, Yang Shi et al.· 0 citations
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedic...
Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.