A theoretical analysis of a lazy, high-dimensional diffusion regime finds that a model can fit its training objective long before it memorizes its training examples. The analysis follows how implicit regularization—the built-in preferences of a training process—affects generalization, objective overfitting, sample memorization and the distribution produced by reverse diffusion. Along gradient-flow training, it identifies three phases: an early spectral estimator that generalizes, a pure-noise score that interpolates the objective, and an eventual empirical Bayes estimator—the data-fitting endpoint in this analysis—that memorizes the data.
The best performance on new data comes at an early stopping scale proportional to the data dimension, d. Training past that point overshoots toward a pure-noise score and progressively forgets learned covariance. Within this model, the training trajectory therefore has a distinct high point before prolonged optimization begins to undo part of what it learned.
The first overfit is not memorization
At polynomially longer training times, high-degree kernel components fit localized bumps around the training data and drive the empirical denoising loss to zero. This is not merely a matter of matching one target number per example: the interpolation covers an entire function on each of n Gaussian tubes rather than n scalar labels.
That distinction is central to the analysis. A score is the learned function used to guide the reverse process. The objective can be overfit while the reverse process still generates a Gaussian law, so a vanishing training loss is not, in this setting, the same thing as reproducing the training samples.
Memorization arrives later
Memorization is assigned to a far later, super-polynomial endpoint—on a time scale beyond any fixed polynomial in the dimension. At that point, gradient flow converges to the empirical Bayes denoiser, the population risk doubles, and the reverse process collapses onto the training set.
The same endpoint appears in the study’s treatment of an unrestricted learner. In a sufficiently rich function class, exact minimization of the empirical objective is identified with the empirical Bayes denoiser. The result links the most complete fit to the memorizing endpoint, rather than to the earlier generalizing phase.
Why the model tends toward Gaussian output
The route to the Gaussian result is tied to the kernel’s structure. At each noise level, the primary denoiser is trained as a separate problem in an RKHS, or kernel-based function space, by gradient flow from zero initialization. The nonlinear part is represented as self-induced regularization that averages over each noisy tube and decomposes the learned score into a linear map plus localized bumps.
Along typical reverse trajectories, the paths delocalize from the training samples, so the localized bumps rarely activate. Learned scores and their linearized surrogates then become asymptotically indistinguishable along those paths. The paper gives this pathwise result only over a restricted training-schedule range, making it a statement about the analyzed regime rather than every possible training duration.
In that same proportional high-dimensional limit, the generated distribution is asymptotically Gaussian. Its covariance is described by a deterministic equivalent, and the target distribution affects the limit only through its first two moments. The finding is a property of the specified lazy model, not a general prediction about all diffusion systems.
A narrow theory with a clear warning
The assumptions are restrictive. The data are centered, their covariance satisfies a log-Sobolev condition, and sample size and dimension grow proportionally. The setup also uses a separate denoiser at each noise level, rather than a single time-conditioned network that couples the levels. The conclusions are asymptotic statements about this theoretical construction, not finite-sample performance results for practical diffusion systems.
The paper’s proposed way past the Gaussian barrier is to leave this regime: use feature learning, add architectural anisotropy or work with genuinely low-dimensional structure. Those are routes suggested by the analysis, not findings established by this study.
One appendix calculation adds a sharper caution. In a special Gaussian-target construction, the coordinate test risk has an almost-surely unbounded limsup along gradient flow. That result belongs to the special construction and should not be read as a finite-sample claim about every diffusion model.
Paper data and sources
Original title: Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime
Authors: Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text