Spectral Analysis of Deep Neural Network Weight Matrices Using Random Matrix Theory
DOI:
https://doi.org/10.29304/jqcsm.2026.18.32652Keywords:
Implicit Organization., Power-Law Distribution, Generalization Performance, Spectral Analysis, Deep Neural Networks, Random Array TheoryAbstract
This paper establishes a framework for systematic spectral analysis of weight matrices in deep neural networks, grounded in Random Matrix Theory. For various architectures, such as multilayer perceptrons, convolutional networks (ResNet-18, VGG-16), and transformers trained on standard benchmarks, the eigenvalue distributions of weight matrices are studied. Well-trained networks deviate significantly from the predictions made by Marchenko-Pasture and are found with eigenvalue spectra that are heavy-tailed and have power-law exponents α ∈ [0.7, 0.9]. We demonstrate that spectral properties and generalization performance are significantly correlated (p < 0.05). For instance, while training, there is a 30.7% improvement in the stable rank and a -0.87 Pearson correlation coefficient (p < 0.001) between the stable rank and the generalization gap. Looking at the layers separately, the first few keep a very random spectrum, but as we go deeper, we start to see a more structured spectrum that starts to diverge from the random matrix models. Spectral parameters of outlier eigenvalues increase by factors ranging from 2.3 to 12.9 during training. These results show how spectral metrics can be used in real life to check the health of a network and how link optimization dynamics, implicit organization, and the quality of learned representation are connected. The findings show that spectral metrics can be used to check the health of a network in real life. They also show a connection between link optimization dynamics, implicit organization, and the quality of learned representations.
Downloads
References
V.A. Marchenko and L.A. Pastur, "Distribution of eigenvalues for some sets of random matrices," Mathematics of the USSR-Sbornik, vol. 1, no. 4, pp. 457–483, 1967.
E.P. Wigner, "On the distribution of the roots of certain symmetric matrices," Annals of Mathematics, vol. 67, no. 2, pp. 325–327, 1958.
C.H. Martin and M.W. Mahoney, "Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning," Journal of Machine Learning Research, vol. 22, no. 165, pp. 1–73, 2021.
C.H. Martin and M.W. Mahoney, "Heavy-tailed self-regularization as a mechanism for deep neural network generalization," arXiv:1901.08276v5, 2021.
J. Pennington and Y. Bahri, "Geometry of neural network loss surfaces via random matrix theory," in Proc. 34th Int. Conf. Mach. Learn. (ICML), vol. 70, pp. 2798–2806, 2017.
M.S. Advani and A.M. Saxe, "High-dimensional dynamics of generalization error in neural networks," Neural Networks, vol. 132, pp. 428–446, 2020.
P.L. Bartlett, D.J. Foster, and M. Telgarsky, "Spectrally-normalized margin bounds for neural networks," Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro, "Exploring generalization in deep learning," Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
A.M. Saxe, J.L. McClelland, and S. Ganguli, "Exact solutions to the nonlinear dynamics of learning in deep linear networks," in Int. Conf. Learn. Represent. (ICLR), 2014.
A. Jacot, F. Gabriel, and C. Hongler, "Neural tangent kernel: Convergence and generalization in neural networks," Advances in Neural Information Processing Systems (NeurIPS), vol. 31, 2018.
B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro, "The role of over-parametrization in generalization of neural networks," in Int. Conf. Learn. Represent. (ICLR), 2019.
P.L. Bartlett and S. Mendelson, "Rademacher and Gaussian complexities: Risk bounds and structural results," Journal of Machine Learning Research, vol. 3, pp. 463–482, 2002.
K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 770–778, 2016.
K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," in Int. Conf. Learn. Represent. (ICLR), 2015.
A. Vaswani et al., "Attention is all you need," Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA: MIT Press, 2016.
G. Yang and S.S. Schoenholz, "Mean field residual networks: On the edge of chaos," Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
S. Hochreiter and J. Schmidhuber, "Flat minima," Neural Computation, vol. 9, no. 1, pp. 1–42, 1997.
S.L. Smith and Q.V. Le, "A Bayesian perspective on generalization and stochastic gradient descent," in Int. Conf. Learn. Represent. (ICLR), 2018.
N. Golowich, A. Rakhlin, and O. Shamir, "Size-independent sample complexity of neural networks," in Conf. Learn. Theory (COLT), pp. 297–299, 2018.
B. Neyshabur, R. Tomioka, and N. Srebro, "Norm-based capacity control in neural networks," in Conf. Learn. Theory (COLT), pp. 1376–1401, 2015.
A. Krizhevsky, I. Sutskever, and G.E. Hinton, "ImageNet classification with deep convolutional neural networks," Commun. ACM, vol. 60, no. 6, pp. 84–90, 2017.
M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, and J. Dickstein, "On the expressive power of deep neural networks," in Proc. 34th Int. Conf. Mach. Learn. (ICML), pp. 2847–2854, 2017.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Mohammed T.A.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.








