Skip to content
Preprint

Video Generation Models are General-Purpose Vision Learners

Jul 2026 · 6 citations · 86 references
Computer Science

TL;DR

GenCeption is introduced, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions, and suggests that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world.

Abstract

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io

View source

Similar papers

Jul 2026

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

This work introduces \method, a Multi-scale Adaptive Vision Encoder, a Multi-scale Adaptive Vision Encoder that uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure.

Sha Lei · 0 citations
Conference Jul 2026

Spatial Knowledge Distillation in Video Models via Vision-Language Guided Zero-Shot Pretraining

Video action recognition tasks require large-scale labeled data to achieve high performance when pretraining is not used. However, annotating video data is costly and time-consuming. To reduce this dependency on labeled data, transfer learning approaches are commonly employed. Vision–language models, which learn generalizable visual representations from large-scale data, are effective at capturing spatial information and thus serve as suitable teacher models for spatial knowledge distillation. In this work, we propose a knowledge distillation-based pretraining approach that leverages zero-shot predictions from vision–language models to initialize video models. In this framework, the resulting soft class distributions are used as supervisory signals and transferred to the student video model. Experimental results on the UCF101 and HMDB51 datasets demonstrate that the proposed method provides an effective weight initialization strategy and yields consistent performance improvements.

A. Çelik, Türkay Yildirim · 0 citations

Condensed Test-Time Adaptation of VLMs for Action Recognition

A novel training-free Condensed Dynamic Adapter C ON DA is proposed, which leverages vision-text alignment to guide vision-vision alignment and is compatible with arbitrary VLM and generalizes well across complex scenarios, such as long-term and egocentric scenarios.

Wenxuan Ge, Hongyu Qu, Rui Yan et al. · 0 citations
Preprint Jul 2026

Vision as Unified Multimodal Generation

Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.

Xiaoyang Han, Jianhua Li, Kewang Deng et al. · 1 citation
Jul 2026

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.

Senqiao Yang, Kaichen Zhang, Zhaoyang Jia et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.