Adversarial training has become the workhorse defense of modern machine learning, yet its theoretical foundations have long rested on a mathematical assumption that real neural networks rarely satisfy. A new study published in the journal Machine Learning by Hailiang Ye and Ziqi An of China Jiliang University and Feilong Cao of Zhejiang Normal University dismantles that reliance, developing a generalization and optimization framework for adversarial training that works under far weaker smoothness conditions than the Lipschitz continuity assumed in most prior analyses. The work, published on 7 October 2026 as volume 115, article 240 of the journal, offers what the authors describe as a systematic step toward understanding one of deep learning’s most stubborn pathologies: robust overfitting.
To appreciate why this matters, it helps to recall what adversarial training actually does. Models trained on clean data can be catastrophically fooled by inputs perturbed by tiny, imperceptible amounts of noise, a phenomenon first explained systematically by Goodfellow, Shlens and Szegedy in their 2015 work on adversarial examples. Adversarial training, formalized in its most influential modern form by Madry and colleagues in 2018, counters this by solving a minimax game: the learner minimizes the loss while an inner adversary maximizes it over a bounded neighborhood of each training point. The resulting robust loss surface is notoriously irregular, and the theory describing how well such models generalize has leaned heavily on the assumption that this loss is Lipschitz continuous, meaning its value cannot change faster than a fixed multiple of the distance moved in the input.
Lipschitz continuity, however, is a strong demand. It implies, among other things, that the function’s rate of change is uniformly bounded everywhere, and it underpins many classical results connecting algorithmic stability to generalization. In practice, adversarial losses computed over deep networks frequently violate this assumption, exhibiting behavior that is smoother than arbitrary but not Lipschitz. The new paper addresses this gap by introducing Hölder continuity for the adversarial loss, a weaker condition in which the function’s change is bounded by a power of the distance rather than the distance itself, and by proposing a new concept the authors call approximate Hölder smoothness, which extends the classical analysis toolkit into this more general regime.
The mathematical machinery at the heart of the paper is a stability analysis of adversarial training. Algorithmic stability, a line of inquiry stretching back to Bousquet and Elisseeff’s 2002 work and famously connected to stochastic gradient descent by Hardt, Recht and Singer in 2016, measures how much a trained model changes when a single training example is swapped. Stable algorithms tend to generalize, because they cannot be memorizing individual points. The authors develop this stability analysis for adversarial training under their Hölder-type conditions, and this stability result then serves as the theoretical foundation for everything that follows: the generalization bounds, the optimization guarantees, and the excess risk analysis.
With that foundation in place, the paper establishes upper bounds on the generalization error of adversarial training under both convex and non-convex adversarial losses. This convex and non-convex split is significant. Convex losses, where the loss surface has a single well-behaved basin, admit cleaner analysis and correspond to simpler models, while non-convex losses describe the deep networks actually deployed in practice, where the loss landscape is riddled with saddle points and multiple minima. By covering both regimes under the same Hölder framework, the authors extend the reach of generalization theory to a broad class of losses that are non-Lipschitz but still Hölder smooth, a class they argue is common across deep learning.
The paper does not stop at generalization. It also bounds the optimization gap of stochastic gradient descent, the workhorse training algorithm, under the same conditions. The optimization gap measures how far the solution found by the algorithm sits from the true optimum of the adversarial objective, and bounding it is essential if generalization guarantees are to translate into statements about models that practitioners actually train. The authors then connect the two threads through an analysis of the induced excess risk, the quantity that captures how much worse the learned robust classifier performs than the best possible one. Together, these results offer theoretical insight into how optimization and generalization interact in adversarial training, an interaction that has remained murky even as empirical evidence of their entanglement has accumulated.
That entanglement is most visible in robust overfitting, the phenomenon that motivated the study. When models are trained adversarially for many epochs, their robust accuracy on test data does not simply plateau; it often peaks early and then declines, even as training robustness keeps improving. Rice, Wong and Kolter documented this striking behavior in 2020, and subsequent work has explored remedies ranging from data augmentation to input loss landscape regularization, including contributions published in this same journal by Li and Spratling in 2023. Yet a principled account of why robust overfitting happens, grounded in the smoothness properties of the loss, has been missing. The new framework contributes to such an account by characterizing generalization behavior under precisely the weaker smoothness conditions observed in practice, giving theorists a vocabulary in which the phenomenon can be analyzed rather than merely mitigated.
The authors do not leave the theory untested. They empirically validate their theoretical findings using a practical example, demonstrating that the framework accurately captures generalization behavior under adversarial training. While the abstract does not detail the full experimental suite, the paper’s notes reference the CIFAR image dataset hosted by the University of Toronto and the SVHN house-numbers dataset from Stanford, both standard benchmarks in robustness research, suggesting the validation was carried out on realistic image classification tasks rather than toy problems. This empirical grounding matters, because generalization bounds in deep learning have often been criticized as too loose to be informative; a framework whose predictions track observed behavior is a more useful instrument for the field.
The study also situates itself within a rapidly growing literature on the stability of adversarial learning. Recent years have seen a flurry of results: Xing, Song and Cheng analyzed the algorithmic stability of adversarial training in 2021; Xiao and colleagues established stability analyses and generalization bounds at NeurIPS in 2022 and followed with PAC-Bayesian spectrally-normalized bounds in 2023 and uniformly stable algorithms in 2024; Farnia and Ozdaglar studied the stability of gradient-based minimax learners; and Tian and Mao presented stability-based generalization bounds for adversarial training at ICLR in 2025. What distinguishes the new contribution is its explicit departure from Lipschitz continuity, aligning it with parallel efforts elsewhere in optimization, such as Gao and Deng’s 2024 work on stochastic weakly convex optimization beyond Lipschitz continuity and Deng and colleagues’ sharper asynchronous SGD bounds beyond Lipschitz and smoothness assumptions.
The broader significance of the work lies in the widening of the theoretical lens through which robust machine learning can be viewed. Hölder-type conditions have recently surfaced in adjacent domains, including a 2025 ICLR analysis of Hölder stability for multiset and graph neural networks by Davidson and Dym, and a survey of Lipschitz calculus perspectives on adversarial robustness by Zühlke and Kudenko in ACM Computing Surveys. By proving that generalization and optimization guarantees survive the relaxation from Lipschitz to Hölder smoothness, and by introducing approximate Hölder smoothness as a workable middle ground, the authors give researchers a tool that matches the messiness of real adversarial loss surfaces. The research was supported by the National Natural Science Foundation of China under grants 62536006 and 12671639 and by the Zhejiang Provincial Natural Science Foundation under grant LZ26F030003. For a field in which the gap between theoretical guarantees and practical robustness has often seemed unbridgeable, a theory built for the conditions actually observed in practice is a meaningful step forward, and one that may shape how the next generation of robustness analyses, and robustness remedies, are designed.
Subject of Research: Generalization and optimization theory of adversarial training under Hölder-type smoothness conditions
Article Title: Beyond Lipschitz: generalization and optimization analysis for adversarial training under Hölder-type conditions
Article References: Ye, H., An, Z., & Cao, F. (2026). Beyond Lipschitz: generalization and optimization analysis for adversarial training under Hölder-type conditions. Machine Learning, 115(10), Article 240. https://doi.org/10.1007/s10994-026-07181-0
Image Credits: AI Generated
DOI: 10.1007/s10994-026-07181-0
Keywords: adversarial training, Hölder continuity, Lipschitz continuity, generalization bounds, algorithmic stability, stochastic gradient descent, robust overfitting, optimization gap, excess risk, machine learning theory, robust generalization, deep learning
Cite Scienmag News
Blake Davidson. (October 8, 2026). New Theory Pushes Adversarial Training Analysis Beyond the Lipschitz Comfort Zone. Scienmag. https://scienmag.com/new-theory-pushes-adversarial-training-analysis-beyond-the-lipschitz-comfort-zone/
Blake Davidson. "New Theory Pushes Adversarial Training Analysis Beyond the Lipschitz Comfort Zone." Scienmag, 8 October 2026, https://scienmag.com/new-theory-pushes-adversarial-training-analysis-beyond-the-lipschitz-comfort-zone/. Accessed 8 October 2026.
Blake Davidson. "New Theory Pushes Adversarial Training Analysis Beyond the Lipschitz Comfort Zone." Scienmag. October 8, 2026. https://scienmag.com/new-theory-pushes-adversarial-training-analysis-beyond-the-lipschitz-comfort-zone/

