Skip to content
Preprint

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

This work studies raw pixels, SD-VAE latents and DINOv2 as well as MAE representation-autoencoder features within a unified masked autoregressive rectified-flow model and shows that compression, reconstruction fidelity, token dimensionality, and visible semantic clustering do not individually predict generative behavior.

Abstract

The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumption is particularly questionable for continuous masked generators, which combine contextual inference from visible tokens with conditional modeling of each missing token. We study raw pixels, SD-VAE latents and DINOv2 as well as MAE representation-autoencoder features within a unified masked autoregressive rectified-flow model. Under a shared ImageNet training budget, these spaces exhibit distinct optimization and inference regimes. DINOv2 converges fastest in both iterations and computation but benefits strongly from a wider local denoiser and direct context fusion. Pixels optimize substantially more slowly and require a different prediction, masking, and guidance configuration. MAE reconstructs images more faithfully and exhibits clear semantic clustering, yet produces generations substantially worse than DINOv2. The representations also respond differently to classifier-free guidance and occupy distinct precision-recall trade-offs. Together, our results show that compression, reconstruction fidelity, token dimensionality, and visible semantic clustering do not individually predict generative behavior. Instead, target representations redistribute difficulty across contextual modeling, per-token denoising, and inference-time distributional control.

View source

Similar papers

Preprint Aug 2026

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.

Nikolai Bolik, Lennart Stöpler, Artur Andrzejak · 0 citations
Preprint Aug 2026

RA-ClipScore: Making Generative Model Evaluation More Interpretable

RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes, and aligns more closely with human perception of visual diversity than existing semantic metrics.

Yifan Lu, Taras Kucherenko, H. Kjellström et al. · 0 citations
Jul 2026

Twins: Learn to Predict Unified Representations with Focal Loss

Twin, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase is proposed, which performs competitively on multimodal understanding benchmarks and improves reconstruction fidelity.

Kaixiong Gong, Xin Cai, Bin Lin et al. · 0 citations
Jul 2026

Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations

A systematic evaluation on how static PTQ affects the interpretability / explainability of five widely used CNN architectures shows that architecture selection is as important as the quantization strategy, and shows that classification accuracy is not a reliable indicator of interpretability stability under reduced precision.

Kazi Kamruzzaman Rabbi, Md. Zami Al Zunaed Farabe, Mohammad Sohel Rahman · 0 citations
Preprint Aug 2026

PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization

Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image predictive subset sufficient for label prediction, while the remaining patches form non-essential complementary context that may correlate with the label. The theoretical analysis shows that restricting prediction to this oracle subset preserves the Bayes risk achievable by the full-patch representation while admitting a complexity bound that tightens with the oracle-subset size. Based on this view, we propose PatchGen, a text-free module that learns a sample-dependent soft predictive-subset mask as a task-driven proxy for the unobserved oracle subset mask. Specifically, histopathology visualizations suggest that PatchGen assigns higher scores to tumor-consistent regions than to some frequently co-occurring inflammatory context. Extensive experiments on natural and histopathological image benchmarks spanning all three shift settings show that PatchGen improves average performance over matched-backbone baselines in most evaluated configurations, enhances generalization to unknown classes, and remains competitive with vision-language methods without text supervision.

Zhaorui Tan, Weimiao Yu, Xi Yang · 0 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.