Kernel ridge regression has long been one of the workhorses of modern machine learning, powering everything from Gaussian-process style prediction to the theoretical study of wide neural networks through the neural tangent kernel. Yet a new study published in the journal Machine Learning argues that one of the field’s most trusted signals, the training error, is close to meaningless when it comes to predicting how well these models will perform on unseen data. The paper, authored by Yang Liu, Ernest Fokoue, Richard Lange, and Daniel Krutz of the Rochester Institute of Technology, builds a fresh account of generalization around a single concept: the alignment between the eigenvectors of the kernel matrix and the learning targets.
The core mathematical observation is deceptively simple. In kernel ridge regression, the prediction for the training data can be written as a weighted sum of the eigenvectors of the kernel matrix, where each weight is the projection of the target vector onto that eigenvector, scaled by the corresponding eigenvalue divided by the eigenvalue plus the regularization parameter. This spectral decomposition means that whatever the model learns is confined to the span of the kernel’s eigenvectors, and the contribution of each direction depends jointly on how strongly the target aligns with it and how large its eigenvalue is. A direction with a huge eigenvalue but poor alignment contributes almost nothing, while a weakly aligned trailing eigenvector can quietly dominate the error.
From this decomposition, the authors derive a striking result about reconstruction, the model’s ability to reproduce its own training labels. They prove that as long as the kernel matrix has sufficient rank, the reconstruction error can be made arbitrarily small simply by shrinking the regularization parameter, an algebraic fact that holds regardless of the data distribution. In high-dimensional settings, near-zero training loss is therefore nearly always achievable. The researchers demonstrate this empirically with synthetic targets built from eigenvectors of a linear kernel: once the input dimension exceeded roughly 500, reconstruction errors collapsed toward zero across the board. The practical implication is unsettling for practitioners who treat a low training loss as evidence of a healthy model, because the training error carries almost no predictive power for test performance.
The study goes further and shows that near-zero reconstruction error is not even a reliable sign of overfitting. In experiments on MNIST images of the digits four and nine, the team constructed learning targets aligned either with the top eigenvectors of the kernel matrix or with its trailing eigenvectors. When the target was aligned with the leading directions, both training and test predictions were nearly perfect. When it was aligned with the trailing directions, the model still reproduced the training data almost exactly, yet its test performance collapsed. The same pattern held on CIFAR10, FashionMNIST, and SVHN2, confirming that the gap between training and test behavior is governed by which part of the eigenspectrum the target inhabits, not by the training loss itself.
To explain these observations, the authors turn to matrix perturbation theory, treating the kernel matrix built from any finite sample as a perturbed version of a hypothetical optimal kernel matrix of the same size. Using classical tools such as the Davis-Kahan theorem, they derive an explicit finite-sample bound on the generalization error expressed in terms of three quantities: the alignment of the learning targets with the eigenvectors, the magnitude of the eigenvalues, and the gaps between consecutive eigenvalues. Strong generalization, the theorem says, requires increasing any of these. Targets aligned with the top eigenvectors benefit from all three at once, because leading eigenvalues are large, well separated, and estimated most accurately from finite data.
This perturbation perspective yields a reinterpretation of the familiar bias-variance tradeoff. In their framing, bias arises when the top eigenvectors cannot represent the learning targets, while variance arises when the targets must be represented by trailing eigenvectors, whose sample estimates are unstable and noisy. Overfitting in kernel ridge regression, on this view, is essentially a failure of the kernel to distinguish information from noise, either because the data are fundamentally too noisy or because the kernel choice is suboptimal. The finding echoes recent work on benign overfitting, but grounds it in concrete spectral quantities that can be computed from a single training set.
Perhaps the most provocative claim concerns regularization. The ridge parameter, long regarded as the main lever for controlling model complexity in the non-parametric setting, turns out to have minimal influence on generalization for most learning targets. Because leading eigenvalues are typically large relative to any reasonable regularization strength, the parameter barely changes their weights. It matters only when targets align with trailing eigenvectors, precisely the regime where generalization is already poor. The authors summarize this bluntly: regularization is more about salvaging bad performance than improving good performance, and its primary role is mitigating numerical instability rather than shrinking the hypothesis space. Their experiments confirm the pattern, showing that stronger regularization often increases generalization error except in the trailing-eigenvector regime.
The paper also delivers a practical payoff in the form of a simple empirical risk estimator, built from the projections of the target onto each eigenvector weighted by the inverse eigenvalue gaps. When benchmarked against the Kernel Alignment Risk Estimator of Jacot and colleagues and the eigenlearning estimator of Simon and colleagues on MNIST and FashionMNIST, the new estimator tracked the true testing error most closely in the leading eigendirections, roughly up to index 250 for MNIST and 300 for FashionMNIST. Existing estimators significantly underestimated the risk in trailing eigendirections, where the new bound, though loose, at least captured the upward trend. The authors note that sharper eigenvector perturbation bounds could improve the estimator further.
Methodologically, the work distinguishes itself from prior analyses by avoiding Gaussian assumptions on the sampling operator, an assumption that recent critiques have shown to be unrealistic in practice, and by focusing on the finite-sample regime rather than asymptotic limits. The comparison with a hypothetical optimal training set, chosen to minimize misalignment with the actual sample, allows the authors to state their bounds without distributional assumptions. They also highlight a connection to empirical observations in deep learning, noting that well-generalized neural networks tend to have weight matrices with isolated eigenvalues, consistent with the emphasis on eigenvalue gaps in their theory.
The authors are candid about limitations. Their analysis assumes noiseless labels, does not handle missing data, and produces bounds that may not be tight, since they are built from a chain of inequalities. Even so, the message is likely to resonate widely: what a kernel machine learns is determined not by how well it fits its training data, but by whether the structure of the data and the kernel conspire to place the target in the stable, well-estimated part of the spectrum. For a field increasingly concerned with understanding why overparameterized models generalize at all, that reframing turns the spotlight away from the loss curve and toward the geometry of the eigenspectrum.
Subject of Research: Eigen-alignment and generalization in kernel ridge regression
Article Title: On Kernel Eigen-Alignments of KRR: Reconstruction and Generalization
Article References: Liu, Y., Fokoue, E., Lange, R., & Krutz, D. (2026). On Kernel Eigen-Alignments of KRR: Reconstruction and Generalization. Machine Learning, 115(10), Article 238. https://doi.org/10.1007/s10994-026-07177-w
Image Credits: AI Generated
DOI: 10.1007/s10994-026-07177-w
Keywords: kernel ridge regression, eigenvectors, generalization, machine learning, kernel alignment, matrix perturbation theory, regularization, overfitting, learning theory, spectral bias, eigenvalues, risk estimation
Cite Scienmag News
Denise Maddox. (October 7, 2026). Eigenvector Alignment, Not Training Error, Predicts How Kernel Machines Generalize. Scienmag. https://scienmag.com/eigenvector-alignment-not-training-error-predicts-how-kernel-machines-generalize/
Denise Maddox. "Eigenvector Alignment, Not Training Error, Predicts How Kernel Machines Generalize." Scienmag, 7 October 2026, https://scienmag.com/eigenvector-alignment-not-training-error-predicts-how-kernel-machines-generalize/. Accessed 7 October 2026.
Denise Maddox. "Eigenvector Alignment, Not Training Error, Predicts How Kernel Machines Generalize." Scienmag. October 7, 2026. https://scienmag.com/eigenvector-alignment-not-training-error-predicts-how-kernel-machines-generalize/

