Skip to content

Category

computer vision

3,022 papers

#machine learning Review Aug 2025

AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef

Coral reefs are on the brink of collapse, with climate change, ocean acidification, and pollution leading to a projected 70-90% loss of coral species within the next decade. Reef restoration is crucial, but its success hinges on introducing automation to upscale efforts. In this work, we present a highly configurable A...

Scarlett Raine, Emilio Olivastri, Benjamin Moshirian et al. · 3 citations
#machine learning Preprint Jun 2025

Squeeze3D: Extreme Neural Compression with Latent Space Bridging

Squeeze3D is a novel framework that leverages implicit prior knowledge learnt by existing pre-trained encoders and decoders to compress 3D data at extremely high compression ratios and can flexibly support different formats, including meshes, point clouds, and radiance fields.

Rishit Dagli, Yu-Shi Guan, Sankeerth Durvasula et al. · 0 citations

Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation.

This work proposes a single-stage, EM-style framework for generative noisy-label learning that is direction-agnostic and avoids explicit image synthesis, and introduces Partial-Label Supervision (PLS), an instance specific prior over clean labels that balances coverage and uncertainty, improving data-dependent regulari...

Feng-Bei Liu, Chong Wang, Yuanhong Chen et al. · 2 citations
#machine learning Preprint Open access Sep 2026

Fast LeWorldModel

Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregres...

Yuntian Gao, Xiangyu Xu · 0 citations

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

This work introduces MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks, designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Mod...

Alan Arazi, Eilam Shapira, Shoham Grunblat et al. · 3 citations
#machine learning Preprint Open access Sep 2026

Let the Target Select for Itself: Data Selection via Target-Aligned Paths

Targeted data selection seeks training examples from a candidate pool that improve downstream task performance. While trajectory-based selectors effectively guide this process, existing approaches typically construct reference states by warming up on the candidate pool itself, thereby inheriting pool-dependent distribu...

Huitao Yang, Hengzhi He, Tung Sum Thomas Kwok et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time

Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that this occurs when models rely more on textual cues and learned...

Itai Allouche, Joseph Keshet · 0 citations
#machine learning Preprint Nov 2024

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

ICER is a black-box framework that addresses the gap in text-to-image models in harmful content generation through two components: an LLM-based rewriter that produces fluent, natural-language adversarial prompts, and in-context experience replay that accumulates successful jailbreaking patterns into a reusable prior.

Zhi-Yi Chin, Pin-Yu Chen, Wei-Chen Chiu et al. · 2 citations
#machine learning Preprint Sep 2026

On-Policy Self-Distillation for Multi-Turn Image Editing

MT-OPSD is proposed, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations, and substantially improves long-horizon editing success and reduces multi-turn collapse.

Liang-Bing Zhao, Le Zhuo, Mohamed Elhoseiny · 0 citations
#machine learning Preprint Sep 2026

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions...

Carlos Garrido-Munoz, Jorge Calvo-Zaragoza · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.