publications
Publications by categories in reversed chronological order. You can also find my articles on my Google Scholar profile and my arXiv profile.
2026
- Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regimeHugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, and Theodor Misiakiewicz2026
Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime n≍d, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.
@article{latourellevigeant2026generalizationmemorizationoverfittingdiffusion, title = {Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime}, author = {Latourelle-Vigeant, Hugo and Chewi, Sinho and Pooladian, Aram-Alexandre and Sous, John and Misiakiewicz, Theodor}, year = {2026}, eprint = {2608.23938}, archiveprefix = {arXiv}, primaryclass = {stat.ML}, url = {https://arxiv.org/abs/2608.23938}, } - Statistical-Computational Trade-offs in Learning Multi-Index Models via Harmonic AnalysisHugo Latourelle-Vigeant and Theodor Misiakiewicz2026
We study the problem of learning multi-index models (MIMs), where the label depends on the input x ∈\mathbbR^d only through an unknown s-dimensional projection W_*^⊤ x ∈\mathbbR^s. Exploiting the equivariance of this problem under the orthogonal group O_d, we obtain a sharp harmonic-analytic characterization of the learning complexity for MIMs with spherically symmetric inputs—which refines and generalizes previous Gaussian-specific analyses. Specifically, we derive statistical and computational complexity lower bounds within the Statistical Query (SQ) and Low-Degree Polynomial (LDP) frameworks. These bounds decompose naturally across spherical harmonic subspaces. Guided by this decomposition, we construct a family of spectral algorithms based on harmonic tensor unfolding that sequentially recover the latent directions and (nearly) achieve these SQ and LDP lower bounds. Depending on the choice of harmonic degree sequence, these estimators can realize a broad range of trade-offs between sample and runtime complexity. From a technical standpoint, our results build on the semisimple decomposition of the O_d-action on L^2(S^d-1) and the intertwining isomorphism between spherical harmonics and traceless symmetric tensors.
@article{latourelle2026statisticalcomputationaltradeoffslearningmultiindex, title = {Statistical-Computational Trade-offs in Learning Multi-Index Models via Harmonic Analysis}, author = {Latourelle-Vigeant, Hugo and Misiakiewicz, Theodor}, year = {2026}, eprint = {2602.09959}, archiveprefix = {arXiv}, primaryclass = {math.ST}, url = {https://arxiv.org/abs/2602.09959}, } - Dyson Equation for Correlated Linearizations and Test Error of Random Features RegressionHugo Latourelle-Vigeant and Elliot PaquetteRandom Matrices: Theory and Applications, 2026
This paper develops some theory of the Dyson equation for correlated linearizations and uses it to solve a problem on asymptotic deterministic equivalent for the test error in random features regression. The theory developed for the correlated Dyson equation includes existence-uniqueness, spectral support bounds and stability properties. This theory is new for constructing deterministic equivalents for pseudo-resolvents of a class of linearizations with correlated entries. In the application, this theory is used to give a deterministic equivalent of the test error in random features ridge regression, in a proportional scaling regime, wherein we have conditioned on both training and test datasets.
@article{latourelle2026dyson, title = {Dyson Equation for Correlated Linearizations and Test Error of Random Features Regression}, author = {Latourelle-Vigeant, Hugo and Paquette, Elliot}, year = {2026}, journal = {Random Matrices: Theory and Applications}, volume = {15}, number = {01}, pages = {2550026}, publisher = {World Scientific}, doi = {10.1142/S2010326325500261}, article = {https://www.worldscientific.com/doi/epdf/10.1142/S2010326325500261}, }
2024
- The matrix Dyson equation for machine learning: Correlated linearizations and the test error in random features regressionHugo Latourelle-Vigeant2024
Contemporary machine learning models, particularly deep learning models, are frequently trained on large datasets within high-dimensional feature spaces, presenting challenges for traditional analytical approaches. Notably, the effective generalization of highly overparameterized models contradicts conventional statistical wisdom. Furthermore, the presence of non-linear activations in artificial neural networks adds complexity to their analysis. To simplify theoretical analysis, it is often assumed that training data is sampled from an unstructured distribution. While such analyses offer insights into certain aspects of machine learning, they fall short in elucidating how neural networks extract information from the structure of the data, crucial for their success in real-world applications. Fortunately, random matrix theory has emerged as a valuable tool for theoretically understanding certain machine learning procedures. Various techniques have been employed to explore large random matrices through asymptotic deterministic equivalents. One such approach involves substituting the random resolvent associated with a large random matrix with the solution of a deterministic fixed-point equation known as the matrix Dyson equation. Another effective technique, known as the linearization trick, involves embedding a matrix expression into a larger random matrix, termed a linear matrix pencil, with a simplified correlation structure. In this thesis, we extend the matrix Dyson equation framework to derive an anisotropic global law for a broad class of pseudo-resolvents with general correlation structures. This extension enables the analysis of spectral properties of a wide range of random matrices using a simpler and deterministic solution to the matrix Dyson equation. Through the development of this theory, we address critical aspects such as existence-uniqueness, spectral support bounds, and stability properties. These considerations are essential for constructing deterministic equivalents for pseudo-resolvents of a class of correlated linear pencils. Leveraging this theoretical framework, we provide an asymptotically exact deterministic expression for the empirical test error of random features ridge regression. The random features model, characterized by its non-linear activation function and potential for overparameterization, emerges as a powerful model for studying phenomena observed in real-life machine learning models, such as multiple descent and implicit regularization. Our exact expression facilitates a precise characterization of the implicit regularization of the model and unveils connections between random features regression and closely related kernel methods. Since we make no particular assumptions about the distribution of the data and response variable, our work represents a significant step towards understanding how neural networks exploit specific data structures.
@mastersthesis{latourelle2024matrix, title = {The matrix Dyson equation for machine learning: Correlated linearizations and the test error in random features regression}, author = {Latourelle-Vigeant, Hugo}, year = {2024}, publisher = {McGill University}, eprint = {https://escholarship.mcgill.ca/concern/theses/m900p100w?locale=en}, article = {https://escholarship.mcgill.ca/concern/theses/m900p100w?locale=en}, }