World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future...
Zi-Yao Zeng, Xiatao Sun, Hao Wang et al.· 0 citations
This paper proposes a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget, exceeding strong feed-forward trackers and a companion analysis explains why several published trackers lose most of their accuracy under this budget.
Masahiro Ogawa, Qi An, Atsushi Yamashita· 0 citations
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction...
Tongtong Liang, Siqi Kou, Ziqiao Xi et al.· 0 citations
Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task.
Ji-Fan Li, Ning Ning· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Diffusion models are now widely used in Bayesian inverse problems in imaging as priors, where latent diffusion models are often used for larger scale problems to keep the computational complexity and model-size manageable. Unfortunately, the auto-encoder based compression results in loss of spatial detail. In addition,...
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can...
Sen Nie, Jie Zhang, Zhong Ling Wang et al.· 0 citations
Constrained Edit Fields (CEF) achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment.
Jing-Xuan Kang, Yin-Song Wang, Che Liu et al.· 0 citations
PARA is proposed, which retains the frozen prompt prediction as a support-invariant semantic anchor and incorporates a visual prediction learned from the support set through an anchor-relative residual and achieves state-of-the-art performance in both few-shot classification and base-to-novel generalization.
Jing-Xuan Kang, Qian-Ying Yue, Che Liu et al.· 0 citations
Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extractio...
Yoav Evron, Michal Bar-Asher Siegal, Michael Fire· 0 citations
Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share this penalty. Because...
Adham M. Alkhadrawi, Mohammed A. B. Mahmoud· 0 citations
This work introduces feature-space routing: a topology-aware state-space operator embedded directly in the forecasting dynamics, which preserves the physical connectivity of the river system while allowing the propagated state itself to be learned end-to-end.
Mohamad Hakam Shams Eddin, M. L. Taccari, Yi-Kui Zhang et al.· 0 citations
Predicting spatial gene expression from routine H&E histology offers a scalable route toward spatial molecular profiling. Recent work has pursued increasingly sophisticated architectures to capture spatial context and richer expression structure. At the same time, simple estimators have shown strong performance in seve...
Duc T. Nguyen, Thanh Ha Do, Phuong M. Cao et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.