Skip to content
Open access

Why diffusion models do not memorize: the role of implicit dynamical regularization in training

Aug 2026 · Journal of Statistical Mechanics: Theory and Experiment · Vol 2026 · 2 citations · 66 references
Physics

TL;DR

It is found that τmem increases linearly with the training set size n, whereas τgen remains constant, which creates a growing window of training times with n where models generalize effectively, despite showing strong memorization if training continues beyond it.

Abstract

Diffusion models have attained remarkable success across a wide range of generative tasks. A key challenge lies in understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time τgen at which models begin to generate high-quality samples and a later time τmem beyond which memorization emerges. Crucially, we found that τmem increases linearly with the training set size n, whereas τgen remains constant. This creates a growing window of training times with n where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when n becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allows to avoid memorization even in highly overparameterized settings. Our findings are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, alongside a theoretical analysis using a tractable random features model studied in the high-dimensional limit.44 https://github.com/tbonnair/Why-Diffusion-Models-Don-t-Memorize. https://github.com/tbonnair/Why-Diffusion-Models-Don-t-Memorize.

Read PDF

Similar papers

Preprint Aug 2026

Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statisti...

Lai Shun Chan, Xiaotian Zhang, Yue Shang et al. · 1 citation
Preprint Aug 2026

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

This work develops a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime by studying denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel.

Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian et al. · 1 citation · ⚡1
Jul 2026

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability comp...

Xiaotian Zhang, Lai Shun Chan, Yue Shang et al. · 0 citations
#machine learning Preprint Aug 2026

Canalization Before Generalization: Grokking as a Dynamical Probe

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-...

Yi-Min Lin · 1 citation
Preprint Aug 2026

Continual-learning rules shape representational drift

Together, these results link representational drift to the stability--plasticity trade-off: its magnitude is shaped by the mechanism that protects old knowledge, and suppressing it can restrict future learning.

Yi-Kai Si, Shanshan Qin · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.