This framework minimizes maximum mean discrepancy in frozen self-supervised video representation spaces in frozen self-supervised video representation spaces, using a hybrid Nystr\"om--Monte Carlo estimator to balance approximation bias and sampling variance.
Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al.· 1 citation
Spectral super-resolution of multispectral satellite images can enable high temporal- and spatial-resolution hyperspectral satellite imagery at a modest cost, significantly increasing the applicability of hyperspectral remote sensing. This task is inherently ill-posed, making it well-suited for deep learning-based meth...
Eval-unlearn is an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models, integrating twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention.
Mansi, Nikhil Raghavan, Zi-Xia Huang et al.· 0 citations
Spatial Grafting is proposed, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry and injects them into the flow-matching action expert through cross-attention without modifying the host's perceptual pathway, so the host retains the full benefit of its pr...
Ding-Sheng Liu, Yang-Zheng Wu, Mahboubeh Asadi et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work presents Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets.
Hao-Xuan Ma, Yi-Hao Liu, Yu-Tao Sun et al.· 0 citations
Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-curated resource, th...
Marco Parola, Mario G. C. A. Cimino, Sabrina Senatore· 0 citations
Comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance, and that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule up...
Pankhuri Vanjani, M. Hatab, Can Mizrakli et al.· 0 citations
Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically distinct tumor subregi...
Md Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans et al.· 0 citations
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from select...
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded bench...
Hao Wang, Tao Yu, Liu-Zhou Zhang et al.· 0 citations
Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial in...
Qian-Kun Li, Yuechen Zhang, Bo-Wen Chen et al.· 0 citations
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using...
Tian-Xiang Chen, Zhentao Tan, Zi Ye et al.· IEEE Transactions on Pattern...· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.