Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Sep 2026

Learning Where and What to Restore for Composite Image Restoration

Real-world degraded images often contain multiple co-occurring degradation types, making composite image restoration fundamentally different from the single-degradation setting assumed by most existing all-in-one methods. These methods typically apply uniform spatial computation and single-label task conditioning, limi...

Jiachen Jiang, Tianyu Ding, Ke Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Localized time-frequency representation learning for bioacoustic classification in complex soundscapes

Prevailing bioacoustic classifiers assign species labels to fixed time-frequency windows rather than to individual vocalizations. When multiple vocalizations occur within the same window, predictions cannot be unambiguously linked to specific calls, which limits analyses at the level of individual vocalizations. This w...

Simen Hexeberg, Mandar Chitre, Matthias Hoffmann-Kuhnt et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

FacePlex: Toward Natural Full-Duplex Conversational Avatars

Natural human conversation is inherently a real-time interaction in which speech and facial behavior continuously evolve. Enabling such interaction requires a conversational avatar to jointly generate speech and facial motion in real time, prepare facial motion for upcoming speech before the corresponding audio is emit...

Habin Lim, Hah Min Lew, Jae-Ho Lee et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

GUI-GenBench: Evaluating Image Generation Models as Interactive GUI Environments

An image generation model used as a graphical user interface (GUI) environment must produce not only a plausible screen but also the correct response to an action. Existing visual-generation benchmarks do not directly assess this combination of functional correctness and temporal consistency in GUIs. We introduce GUI-G...

Haodong Li, Jingwei Wu, Quan Sun et al. · 0 citations

Selective Fine-Tuning for Targeted and Robust Concept Unlearning

TRUST (Targeted Robust Selective fine Tuning), a novel approach for dynamically estimating target concept neurons and unlearning them through selective finetuning, empowered by a Hessian based regularization, is proposed.

Mansi, Avinash Kori, Francesca Toni et al. · 2 citations
#artificial intelligence Preprint Feb 2026

Semantic Purification for Conditional Representation Learning

Semantic Purification for Conditional Representation Learning (SP-CRL) first decomposes the original text basis and performs curvature-based adaptive truncation on the resulting basis vectors to construct a purer conditional subspace, then identifies an appropriate noise subspace and projects image embeddings onto its...

Jia-Quan Wang, Y. Lyu, Chen Li et al. · 0 citations
#artificial intelligence Open access May 2025

Building Intelligent Agents with Neuro-Symbolic Concepts

A concept-centric framework for building agents that can learn continually and reason flexibly across multiple domains and offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.

Jia-Yuan Mao, Joshua B. Tenenbaum, Jia-Jun Wu · 14 citations · ⚡1
#artificial intelligence Preprint Sep 2026

FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets

FurE is presented, an efficient strand-based animal fur reconstruction method that recovers a per-strand, editable groom by optimizing a root-conditioned latent field, decoded into strand geometry via a PCA-based decoder, and shows that a PCA-based decoder learned from human-hair strand data can alleviate animal-data s...

Srinjay Sarkar, Prakhar Kaushik, Soumava Paul et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and th...

Yi-Jia Fan, Zi-Qi Huang, Zhongang Cai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and e...

Huai-Yuan Qin, Mu-Li Yang, Gabriel James Goenawan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation

Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has...

Yuta Oshima, Ku Onoda, Yusuke Iwasawa et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

SolveEdit: Benchmarking Visual Problem Solving in Generative Models

Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without di...

Wenjie Shu, Yexin Liu, Harold Haodong Chen et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.