IR4RL is introduced, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence, showing that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code model...
Omri Kaduri, Kate Feingold, Phillip Isola et al.· 0 citations
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-...
Ya-Jiao Xiong, You-Yu Luan, Xiao-Yu Zhou et al.· 0 citations
This work introduces ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection, and diagnostic measurements of these distinct failure m...
Gueter Josmy Faure, Min-Hung Chen, Hao-Ping Wang et al.· 0 citations
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simpl...
Hangyul Yoon, Hyungyung Lee, Edward Choi et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
VLAcneSeg is proposed, a multimodal framework for acne lesion segmentation that leverages CLIP and region-level text prompts to incorporate spatial priors, enabling lesions to be localized across the whole face.
Sukju Oh, Soo Ick Cho, D. Suh et al.· IEEE journal of biomedical a...· 0 citations
We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends world modeling from...
Xin-Yao Liao, Xian-Fang Zeng, Zhu Liang et al.· 0 citations
Due to the gigapixel-scale nature of whole-slide images (WSIs), weakly supervised WSI analysis is commonly formulated as a multiple instance learning (MIL) problem, where patch-level features are aggregated into slide-level representations. However, diagnostic and prognostic evidence often arises from spatially coheren...
Lei Wu, Jiashuai Liu, Di Zhang et al.· 0 citations
MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics, is proposed and established as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosys...
Bing Jiang, Li Luo, Zi-Chao Yu et al.· 0 citations
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in...
Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents sh...
Yan-Wei Huang, Min Zhu, Shu-Jie Li et al.· 0 citations
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typic...
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even whe...
Cheng-Yang Zhang, Wen-Chuan Zhang, Bo Li et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.