Vision Transformers (ViTs) implement depth by stacking independently parameterized blocks, but it remains unclear how much of this parameterization is necessary and how much can be replaced by recurrent reuse. We study this question with bViT, a single-block recurrent ViT that repeatedly applies the same transformer bl...
Michal Byra, Pawel Olszowiec, Grzegorz Stefanski et al.· 0 citations
FraudBench is a multimodal benchmark for detecting AI-generated fraudulent refund evidence and shows that current MLLMs often recognize real-damaged evidence but fail on many fake-damaged subsets, with fake-damage detection rates far below the 50\% baseline on most generator subsets.
This work introduces LAGO (LAnguage-Guided adaptive Object-region focus), which reframes localized recognition as language-guided directed region discovery and addresses the circular dependence between recognizing a class and locating its supporting evidence, while preserving complementary local, contextual, and global...
Jun-Yi Hu, Qiji Zhou, Lei Zhang et al.· arXiv.org· 0 citations
Vision-language-action (VLA) models can use visual prediction to anticipate future states, but dense visual features make the generative sequence grow with the number of camera views, prediction horizon, and encoder resolution. Whether such dense representations are necessary for effective control remains unclear. We i...
Zuojin Tang, Shengchao Yuan, Xiaoxin Bai et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Human visual reasoning is governed by active vision, a process where meta-cognitive control drives top-down goal-directed attention, dynamically routing foveal focus toward task-relevant details while maintaining peripheral awareness of the global scene. In contrast, modern Vision-Language Models (VLMs) process visual...
Brown Ebouky, Gabriele Carrino, Niccolo Avogaro et al.· 0 citations
MMDG-Bench is introduced, the first unified and comprehensive benchmark for MMDG, which standardizes evaluation across six datasets spanning three diverse tasks: action recognition, mechanical fault diagnosis, and sentiment analysis.
Hao Dong, Hong-Zhao Li, Shu-Pan Li et al.· arXiv.org· 0 citations
Vision-Language Models (VLMs) hallucinate objects that are not present, and a growing line of work tries to curb this by feeding the model its own generated caption as auxiliary evidence -- assuming that a caption, once available, is something to consume. We show this fails: naively appending a caption can lower accura...
The DRAGON dataset contains 11,664 annotated question instances from six diagram QA datasets, with a 2,445-instance test set carrying human-verified evidence annotations and a standardized evaluation framework, which supports future research on models that ground their predictions in visual evidence.
Anirudh Iyengar, Tampu Ravi Kumar, Gaurav Najpande et al.· arXiv.org· 0 citations
SGP-SAM, a self-gated prompting framework for efficient and effective transfer to 3D lesion segmentation, and a Zoom Loss that up-weights lesion-focused supervision by combining Dice and a voxel-balanced focal term to address small-lesion learning.
Ze-Quan Yao, Zi-Xuan Tang, Jie Ma et al.· arXiv.org· 0 citations
Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability often outperforms their generation capability. This mismatch suggests that the model's rich internal knowledg...
Medical image generators trained on imbalanced data can fail at demographic intersections absent from training. We introduce CompDiff, which encodes age, sex and race separately and composes supervised demographic tokens alongside clinical text. Across chest radiographs and fundus images, CompDiff improves overall and...
Mahmoud K. Ibrahim, Bart Elen, Chang Sun et al.· 0 citations
A novel guided method is proposed by using the h-transform, a tool that can constrain stochastic processes (e.g., sampling process) under desired conditions and modify the transition probability at each sampling timestep by adding to the original differential equation with a drift function, which approximately steers t...
Yanghao Wang, Zi-Qi Jiang, Zhen Wang et al.· 3 citations· ⚡1
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.