This work introduces COMiT, a communication-inspired framework for learning structured discrete visual representations that substantially improves compositional generalization and relational reasoning over prior methods.
Aram Davtyan, Y. Şahin, Yasaman Haghighi et al.· 0 citations
Text-to-image diffusion models can reproduce specific artists visual styles at extremely low cost, raising copyright and deployment safety concerns about unauthorized style mimicry. Existing model-side protection methods generally follow ordinary concept erasure, emphasizing aggressive deletion or redirection of target...
Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the application of tokens to physical simulation remains nascent....
Levi Lingsch, Georgios Kissas, Johannes Jakubik et al.· 0 citations
Dominant paradigms for 4D LiDAR panoptic segmentation are usually required to train deep neural networks with large superimposed point clouds or design dedicated modules for instance association. However, these approaches perform redundant point processing and consequently become computationally expensive, yet still ov...
Gyeongrok Oh, Youngdong Jang, Jonghyun Choi et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Conformal Prediction Active TTA is proposed, which first brings principled, conformal uncertainty with coverage-aware online calibration into ATTA, and consistently outperforms the state-of-the-art ATTA methods by around 5% in accuracy.
Ting-Yu Shi, Fan Lyu, Hai-Hua Zhu et al.· 1 citation
Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to improving survival rates, but misdiagnosis or delayed diagnosis can lead to poor outcomes. Despite advancements in AI-based image analysis, diagnosing choroidal nevi in colour f...
Mohammadmahdi Eshragh, Emad A. Mohammed, Behrouz Far et al.· 0 citations
Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high stakes settings. A model should stay robust under such transformations, expose why a correction was applied, and signal when its input is...
Johann Schmidt, Tom Siegl, Martin Becker et al.· 0 citations
Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly understood. We empirically observe visual elements from different i...
Yeji Park, Minyoung Lee, Sanghyuk Chun et al.· 0 citations
A modular manipulation framework that separates high-level planning from low-level control and coupling grounded plan generation with a 3D-based execution policy, this framework achieves state-of-the-art performance on the challenging GemBench benchmark and demonstrates promising transfer to real robots.
Shi-Zhe Chen, Ricardo Garcia, Paul Pacaud et al.· 2 citations
Unlike traditional video editing or inpainting, video semantic mixing fuses a reference concept with a moving target entity to produce a hybrid while preserving the source video's motion and layout. We propose MoCA-Video, a training-free framework that steers a frozen video-diffusion denoising trajectory through concep...
Tong Zhang, Victor Escorcia, Juan C Leon Alcazar et al.· 0 citations
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the-art models fail to capture topology and symbolic logic in engineering schematics, as their pixel-driven paradigm discards the explicit vect...
Chengwei Ma, Zhen Tian, Zhou Zhou et al.· 0 citations
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across...
Hui-Hui Ren, Lei Fan, Henry Pao et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.