Skip to content

Toward Efficient Consistency Models via Variance-Reduced Distillation.

Jul 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP, pp. 1-10 · 0 citations
Medicine

TL;DR

This article introduces variance-reduced consistency learning (vrCL), a novel distillation technique that enables stable and efficient training of consistency models without relying on teacher model evaluations, resulting in high computational efficiency and significantly reduced training time.

Abstract

Diffusion models have demonstrated remarkable performance across a wide range of generative tasks; however, their high sampling cost remains a critical bottleneck. To address this, consistency distillation (CD) was proposed, offering a reduction in sampling cost by distilling a pretrained diffusion model. However, achieving generative quality comparable to diffusion models requires extensive training for the distillation process, posing a substantial computational challenge. In this article, we introduce variance-reduced consistency learning (vrCL), a novel distillation technique that enables stable and efficient training of consistency models without relying on teacher model evaluations. By leveraging a student-guided sample pair, vrCL ensures training stability while significantly reducing computational costs. This design eliminates the need for repeated teacher model evaluations during training, resulting in high computational efficiency and significantly reduced training time. Empirical results demonstrate that vrCL achieves competitive generative performance with high training efficiency, reaching strong results within just 100k training iterations.

View source

Similar papers

Conference 2026

Shortcut Diffusion Training With Cumulative Consistency Loss: An Optimal Control View

This paper forms few-step generation as a controlled base generative process, and shows that self-consistency loss can be understood through the lens of optimal control, and draws a connection between this approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generation.

Paribesh Regmi, S. Ghimire, Rui Li · 0 citations
Preprint Jul 2026

Diversify Diffusion with Temperature Sampling and Variance-Corrective Time Shifting

The correction turns simple temperature sampling into a practical diversity knob for pretrained diffusion and flow-matching backbones with no retraining, and consistent gains at minimal cost to sample quality and condition fidelity across DiT, Stable Diffusion and Motion Diffusion models are demonstrated.

Peizhuo Li, Emre Aksan, A. Ichim et al. · 0 citations
Preprint Jul 2026

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stages. Specifically, CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC), which respectively capture reliable semantic relations for correlation estimation and preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of the calibrated redundancy suppression objective according to its relative loss scale, reducing sensitivity to coefficient settings. Extensive experiments on CIFAR-100 and ImageNet-1K validate the effectiveness of CoCaRS in improving distillation performance and reducing sensitivity to coefficient settings. Code will be released soon.

Feng Yu, Haiwei Pan, Kejia Zhang et al. · 0 citations
Preprint Aug 2026

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. In this paper, we propose a compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across scaling model widths, and subsequently extrapolating to trillion-token horizons. First, we formulate a Maximal Update Parameterization ($\mu$P) adaptation for MoE architectures utilizing Multi-head Latent Attention (MLA) and the Muon optimizer, demonstrating that optimal learning rates transfer consistently across width-scaled models. Second, we extend this transferability along the token dimension by establishing a predictive scaling law. By applying linear regression to the optimal values derived from small proxy models on limited budgets, we successfully extrapolate the ideal learning rate to massive training horizons (e.g., 10 trillion tokens) with high fidelity ($R^2=0.95$). Consequently, this indicates that proxy training on small models is sufficient to determine the optimal learning rate for the extensive training of large-scale MoEs. We apply the proposed methodology to pretrain our foundation model (155B total, 17B active parameters) from scratch, and the stable training and evaluation results validate that optimal configurations for full-scale target models can be accurately predicted with minimal ablation costs.

Nayeon Kim, Hojin Lee, Yunju Bak et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.