Coral reefs are on the brink of collapse, with climate change, ocean acidification, and pollution leading to a projected 70-90% loss of coral species within the next decade. Reef restoration is crucial, but its success hinges on introducing automation to upscale efforts. In this work, we present a highly configurable A...
Scarlett Raine, Emilio Olivastri, Benjamin Moshirian et al.· arXiv.org· 3 citations
Squeeze3D is a novel framework that leverages implicit prior knowledge learnt by existing pre-trained encoders and decoders to compress 3D data at extremely high compression ratios and can flexibly support different formats, including meshes, point clouds, and radiance fields.
Rishit Dagli, Yu-Shi Guan, Sankeerth Durvasula et al.· 0 citations
This work proposes a single-stage, EM-style framework for generative noisy-label learning that is direction-agnostic and avoids explicit image synthesis, and introduces Partial-Label Supervision (PLS), an instance specific prior over clean labels that balances coverage and uncertainty, improving data-dependent regulari...
Feng-Bei Liu, Chong Wang, Yuanhong Chen et al.· IEEE Transactions on Pattern...· 2 citations
Granularity alignment, not algorithm choice, localizes when routing helps, and what category supervision buys when deployed through routing is quantified.
Boyao Wang, Zhihan Lei· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregres...
This work introduces MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks, designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Mod...
Alan Arazi, Eilam Shapira, Shoham Grunblat et al.· arXiv.org· 3 citations
Targeted data selection seeks training examples from a candidate pool that improve downstream task performance. While trajectory-based selectors effectively guide this process, existing approaches typically construct reference states by warming up on the candidate pool itself, thereby inheriting pool-dependent distribu...
Huitao Yang, Hengzhi He, Tung Sum Thomas Kwok et al.· 0 citations
Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that this occurs when models rely more on textual cues and learned...
ICER is a black-box framework that addresses the gap in text-to-image models in harmful content generation through two components: an LLM-based rewriter that produces fluent, natural-language adversarial prompts, and in-context experience replay that accumulates successful jailbreaking patterns into a reusable prior.
Zhi-Yi Chin, Pin-Yu Chen, Wei-Chen Chiu et al.· 2 citations
This work introduces Projected Distribution Matching Distillation (PDMD) to filter critic errors, and shows how this projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures.
Zi-Mo Wang, Junkun Yuan, Ang-Tian Wang et al.· 0 citations
MT-OPSD is proposed, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations, and substantially improves long-horizon editing success and reduces multi-turn collapse.
Liang-Bing Zhao, Le Zhuo, Mohamed Elhoseiny· 0 citations
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions...
Carlos Garrido-Munoz, Jorge Calvo-Zaragoza· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.