Aug 2026· Journal of Statistical Mechanics: Theory and Experiment· Vol 2026· 2 citations· 66 references
Physics
TL;DR
It is found that τmem increases linearly with the training set size n, whereas τgen remains constant, which creates a growing window of training times with n where models generalize effectively, despite showing strong memorization if training continues beyond it.
Abstract
Diffusion models have attained remarkable success across a wide range of generative tasks. A key challenge lies in understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time τgen at which models begin to generate high-quality samples and a later time τmem beyond which memorization emerges. Crucially, we found that τmem increases linearly with the training set size n, whereas τgen remains constant. This creates a growing window of training times with n where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when n becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allows to avoid memorization even in highly overparameterized settings. Our findings are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, alongside a theoretical analysis using a tractable random features model studied in the high-dimensional limit.44 https://github.com/tbonnair/Why-Diffusion-Models-Don-t-Memorize. https://github.com/tbonnair/Why-Diffusion-Models-Don-t-Memorize.
Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statisti...
Lai Shun Chan, Xiaotian Zhang, Yue Shang et al.· 1 citation
This work develops a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime by studying denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel.
Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian et al.· 1 citation· ⚡1
While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability comp...
Xiaotian Zhang, Lai Shun Chan, Yue Shang et al.· arXiv.org· 0 citations
For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-...
Together, these results link representational drift to the stability--plasticity trade-off: its magnitude is shaped by the mechanism that protects old knowledge, and suppressing it can restrict future learning.
The evidence supports single-direction decodability for the studied task groups but challenges fixed-direction stability: the information persists while its geometric realization changes.
Shaheen Nabi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.