Deep neural networks have transformed fields from medical imaging to machine translation, yet their reliability in the real world remains fragile in two stubborn ways. They tend to settle into sharp, narrow regions of the loss landscape that generalize poorly, and they are notoriously easy to derail when training data contains mislabeled examples. A research team at the University of Malakand in Pakistan now reports a unified framework that tackles both problems at once. The method, called SLAW for Sharpness- and Loss-Adaptive Weighting, was published in the International Journal of Data Science and Analytics by Zaryab Rahman, Fakhrud Din and Shah Khalid, and its central promise is deceptively simple: make the training process itself aware of both the geometry of the loss surface and the statistical behavior of the samples flowing through each mini-batch.
The first half of the framework, Sharpness-Adaptive Label Smoothing, builds on a technique that practitioners have used for years. Label smoothing softens the hard one-hot targets of standard classification, replacing a certain answer with a distribution that leaves some probability mass on incorrect classes. This discourages the network from becoming overconfident and pushes it toward flatter minima, which earlier work by Hochreiter and Schmidhuber and later large-batch studies linked to better generalization. The catch is that the smoothing strength is usually a fixed hyperparameter, tuned by trial and error, even though the need for regularization changes dramatically over the course of training. SLAW makes that strength dynamic. It computes a first-order proxy of local sharpness, an inexpensive estimate of how steeply the loss rises around the current parameters, and uses it to modulate the smoothing applied at each step.
The intuition behind this adaptive scheme is geometric. Early in training, when gradients are large and the optimizer is traversing high-curvature regions of the loss landscape, the network benefits from strong regularization that discourages it from diving into narrow valleys. As training stabilizes and the parameters approach flatter basins, aggressive smoothing becomes counterproductive, blurring the decision boundaries the model needs to fit the data. SALS therefore applies heavier smoothing in steep, high-curvature areas and gradually relaxes it as the landscape flattens. Because the sharpness estimate relies only on first-order information rather than the full Hessian, which would be computationally prohibitive for modern architectures, the method keeps the overhead low enough for routine use. The authors report that this component improves not only generalization accuracy but also predictive calibration, addressing the well-documented tendency of modern networks to be confidently wrong.
The second component, Loss-Adaptive Reweighting, addresses the label-noise problem from a different angle. Deep networks have a well-known and troubling habit: given enough epochs, they memorize even blatantly wrong labels, because the capacity of these models is sufficient to fit arbitrary noise. Memorization typically shows up late in training, when the loss on corrupted samples that the network has already fit collapses toward zero while the loss on genuinely hard, clean examples remains elevated. LAW exploits this statistical signature. Within each mini-batch, it standardizes the per-sample loss values and flags samples whose losses are anomalous relative to their batch peers, then down-weights their contribution to the parameter update. Corrupted labels, which the model fits quickly and confidently, thus lose influence before they can drag the shared weights in the wrong direction.
What distinguishes LAW from the crowded field of noise-robust training methods is what it does not require. Many existing approaches, such as co-teaching, loss correction, and joint methods like DivideMix, depend on estimates of the noise rate or assumptions about the structure of the noise, for example that it is symmetric and flips labels uniformly across classes. In practice, real-world annotation errors are rarely so well-behaved, and misestimating the noise rate can be as damaging as ignoring it. LAW involves no prior knowledge of noise characteristics and no noise estimation rate at all. It simply reacts to the empirical distribution of losses in each batch, making it applicable to datasets whose corruption level is unknown or heterogeneous, which is the norm rather than the exception in web-scraped and crowd-annotated data.
The experimental evaluation covers three standard benchmarks: MNIST, CIFAR-10 and CIFAR-100. On clean data, SLAW matches conventional training algorithms, an important sanity check showing that the added machinery does not sacrifice baseline performance. The more striking results appear under synthetic label corruption. At moderate noise levels, with twenty percent of labels flipped symmetrically, SLAW nearly eliminates the catastrophic overfitting that typically afflicts late training, when standard models begin absorbing the corrupted labels and their test accuracy degrades. At extreme noise of forty percent, a regime in which the baseline models usually collapse entirely, SLAW continues to learn stably and retains meaningful generalization, a result the authors present as evidence that the framework can function where most pipelines fail outright.
Ablation studies disentangle the contributions of the two components and reveal a clean division of labor. SALS, the sharpness-adaptive smoothing, is the primary driver of improved generalization and calibration on clean and moderately noisy data, consistent with the landscape-flattening role of label smoothing identified in earlier analyses by Müller, Kornblith and Hinton. LAW, by contrast, is the main engine of robustness to label noise, doing the heavy lifting when corrupted samples threaten to poison the updates. The synergy matters because the two failure modes interact: noisy labels tend to push optimizers toward sharp minima, since memorizing outliers often requires fitting narrow spikes in the loss surface, so a method that addresses only one vulnerability leaves the other exposed.
The work sits at the intersection of several research threads that have matured over the past decade. Sharpness-aware minimization, introduced by Foret and colleagues, explicitly optimizes for flat minima by perturbing weights to worst-case nearby points before taking a gradient step, and adaptive variants such as ASAM refined the idea. Sample reweighting has its own lineage, from MentorNet’s learned curricula to meta-learned example weighting by Ren and co-workers and recent dynamic loss-based reweighting for large language model pretraining. SLAW’s contribution is less a single new primitive than a demonstration that these ideas can be fused cheaply and adaptively, with the sharpness signal controlling regularization strength and the loss statistics controlling sample influence, without the two mechanisms interfering with each other.
For practitioners, the appeal lies in the framework’s economy of assumptions. Real datasets are messy: medical registries contain transcription errors, web labels reflect the biases of their annotators, and crowdsourced annotations disagree in ways no symmetric noise model captures. A training method that neither assumes a known noise rate nor demands expensive second-order curvature computations fits the constraints of applied machine learning, where the authors argue such techniques are most needed. The code has been released on GitHub, lowering the barrier to independent replication, and the benchmarks, while standard, are the same ones on which the field’s noisy-label methods are conventionally compared.
The broader significance may lie in what SLAW implies about the relationship between optimization geometry and data quality. The results support a view in which these are not separate concerns to be handled by separate tools, but coupled aspects of a single training dynamic: the shape of the loss landscape determines how eagerly a network memorizes bad examples, and the composition of the data determines which minima the optimizer can reach. A framework that senses both, adjusting its regularization as the terrain flattens and its trust in each sample as the loss statistics shift, offers a template for the kind of self-correcting training that reliable, real-world deep learning systems will likely require as models are deployed on data that is, as the authors put it, imperfect, noisy, or of low quality.
Subject of Research: A sharpness- and loss-adaptive weighting framework for robust deep learning under label noise
Article Title: Slaw: sharpness- and loss-adaptive weighting for robust deep learning
Article References: Slaw: sharpness- and loss-adaptive weighting for robust deep learning. (n.d.). https://doi.org/10.1007/s41060-026-01272-w
Image Credits: AI Generated
DOI: 10.1007/s41060-026-01272-w
Keywords: deep learning, label noise, loss landscape, sharp minima, label smoothing, sample reweighting, generalization, calibration, robustness, sharpness-aware minimization, neural networks, CIFAR-10
Cite Scienmag News
Cassandra Pierce. (October 2, 2026). New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels. Scienmag. https://scienmag.com/new-training-method-steers-neural-networks-toward-flat-minima-and-away-from-bad-labels/
Cassandra Pierce. "New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels." Scienmag, 2 October 2026, https://scienmag.com/new-training-method-steers-neural-networks-toward-flat-minima-and-away-from-bad-labels/. Accessed 2 October 2026.
Cassandra Pierce. "New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels." Scienmag. October 2, 2026. https://scienmag.com/new-training-method-steers-neural-networks-toward-flat-minima-and-away-from-bad-labels/

