Real-world degraded images often contain multiple co-occurring degradation types, making composite image restoration fundamentally different from the single-degradation setting assumed by most existing all-in-one methods. These methods typically apply uniform spatial computation and single-label task conditioning, limi...
Jiachen Jiang, Tianyu Ding, Ke Zhang et al.· 0 citations
Prevailing bioacoustic classifiers assign species labels to fixed time-frequency windows rather than to individual vocalizations. When multiple vocalizations occur within the same window, predictions cannot be unambiguously linked to specific calls, which limits analyses at the level of individual vocalizations. This w...
Simen Hexeberg, Mandar Chitre, Matthias Hoffmann-Kuhnt et al.· 0 citations
Natural human conversation is inherently a real-time interaction in which speech and facial behavior continuously evolve. Enabling such interaction requires a conversational avatar to jointly generate speech and facial motion in real time, prepare facial motion for upcoming speech before the corresponding audio is emit...
Habin Lim, Hah Min Lew, Jae-Ho Lee et al.· 0 citations
An image generation model used as a graphical user interface (GUI) environment must produce not only a plausible screen but also the correct response to an action. Existing visual-generation benchmarks do not directly assess this combination of functional correctness and temporal consistency in GUIs. We introduce GUI-G...
Haodong Li, Jingwei Wu, Quan Sun et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
TRUST (Targeted Robust Selective fine Tuning), a novel approach for dynamically estimating target concept neurons and unlearning them through selective finetuning, empowered by a Hessian based regularization, is proposed.
Mansi, Avinash Kori, Francesca Toni et al.· arXiv.org· 2 citations
Semantic Purification for Conditional Representation Learning (SP-CRL) first decomposes the original text basis and performs curvature-based adaptive truncation on the resulting basis vectors to construct a purer conditional subspace, then identifies an appropriate noise subspace and projects image embeddings onto its...
Jia-Quan Wang, Y. Lyu, Chen Li et al.· 0 citations
A concept-centric framework for building agents that can learn continually and reason flexibly across multiple domains and offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.
Jia-Yuan Mao, Joshua B. Tenenbaum, Jia-Jun Wu· Communications of the ACM· 14 citations· ⚡1
FurE is presented, an efficient strand-based animal fur reconstruction method that recovers a per-strand, editable groom by optimizing a root-conditioned latent field, decoded into strand geometry via a PCA-based decoder, and shows that a PCA-based decoder learned from human-hair strand data can alleviate animal-data s...
Srinjay Sarkar, Prakhar Kaushik, Soumava Paul et al.· 0 citations
UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and th...
Yi-Jia Fan, Zi-Qi Huang, Zhongang Cai et al.· 0 citations
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and e...
Huai-Yuan Qin, Mu-Li Yang, Gabriel James Goenawan et al.· 0 citations
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has...
Yuta Oshima, Ku Onoda, Yusuke Iwasawa et al.· 0 citations
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without di...
Wenjie Shu, Yexin Liu, Harold Haodong Chen et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.