Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

COMIC: Agentic Sketch Comedy Generation

We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting from character references, the system employs a population of agents loosely modeled on roles in real production studios and structured to optimize the quality and diversity of ideas a...

Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RAC: Rectified Flow Auto Coder

In this paper, we propose a Rectified Flow Auto Coder (RAC) inspired by Rectified Flow to replace the traditional VAE: 1. It achieves multi-step decoding by applying the decoder to flow timesteps. Its decoding path is straight and correctable, enabling step-by-step refinement. 2. The model inherently supports bidirecti...

Sen Fang, Yalin Feng, Yanxin Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

De-occluding broadband metalens

Obstructions such as raindrops, fences, or dust degrade captured images, especially when mechanical cleaning is infeasible. Conventional solutions to obstructions rely on a bulky compound optics array or computational inpainting, which compromise compactness or fidelity. Metalenses composed of subwavelength meta-atoms...

Seungwoo Yoon, Dohyun Kang, Eunsue Choi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Universal Weight Subspace Hypothesis

We show that deep neural networks trained across diverse tasks exhibit remarkably similar low-dimensional parametric subspaces. We provide the first large-scale empirical evidence that demonstrates that neural networks systematically converge to shared spectral subspaces regardless of initialization, task, or domain. T...

Prakhar Kaushik, Shravan Chaudhari, Ankit Vaidya et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ZeST: an VLM-based Zero-Shot Traversability Navigation for Unknown Environments

The advancement of robotics and autonomous navigation systems hinges on the ability to accurately predict terrain traversability. Traditional methods for generating datasets to train these prediction models often involve putting robots into potentially hazardous environments, posing risks to equipment and safety. To so...

Shreya Gummadi, Mateus V. Gasparino, Gianluca Capezzuto et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Optimizing Breast Cancer Detection in Mammograms: A Comprehensive Study of Transfer Learning, Resolution Reduction, and Multi-View Classification

Mammography, an X-ray-based imaging technique, remains central to the early detection of breast cancer. Recent advances in artificial intelligence have enabled increasingly sophisticated computer-aided diagnostic methods, evolving from patch-based classifiers to whole-image approaches and then to multi-view architectur...

Daniel G. P. Petrini, Hae Yong Kim · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation

In this paper, we address the problem of human trajectory forecasting, which aims to predict the inherently multi-modal future movements of humans based on their past trajectories and other contextual cues. We propose a novel motion prediction conditional flow matching model, termed MoFlow, to predict K-shot future tra...

Yuxiang Fu, Qi Yan, Lele Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth...

Jiaming Bian, Bingliang Li, Yuehao Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulat...

Shih-Chen Tseng, Chih-Hsuan Chen, Ryan Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Learning to Read the Contextual Tokens in Diffusion Transformers

Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading t...

Omer Dahary, Etai Sella, Hadar Averbuch-Elor et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

Sliders provide an intuitive interface for continuous image editing. In current generative approaches, however, the slider is simply a rescaling of the method's strength parameter, such as an adapter coefficient, a prompt weight, or an interpolation factor. This strength relates poorly to perceptual change. The image c...

David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TAPDreamer: Transferable Adversarial Patches for World Action Models

World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models opti...

Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.