Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels

October 2, 2026
in Technology and Engineering
Cassandra Pierce
By Cassandra Pierce Scienmag Editorial Profile - Systems Neuroscience
Reading Time: 5 mins read
0
New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels

New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels

New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Deep neural networks have transformed fields from medical imaging to machine translation, yet their reliability in the real world remains fragile in two stubborn ways. They tend to settle into sharp, narrow regions of the loss landscape that generalize poorly, and they are notoriously easy to derail when training data contains mislabeled examples. A research team at the University of Malakand in Pakistan now reports a unified framework that tackles both problems at once. The method, called SLAW for Sharpness- and Loss-Adaptive Weighting, was published in the International Journal of Data Science and Analytics by Zaryab Rahman, Fakhrud Din and Shah Khalid, and its central promise is deceptively simple: make the training process itself aware of both the geometry of the loss surface and the statistical behavior of the samples flowing through each mini-batch.

The first half of the framework, Sharpness-Adaptive Label Smoothing, builds on a technique that practitioners have used for years. Label smoothing softens the hard one-hot targets of standard classification, replacing a certain answer with a distribution that leaves some probability mass on incorrect classes. This discourages the network from becoming overconfident and pushes it toward flatter minima, which earlier work by Hochreiter and Schmidhuber and later large-batch studies linked to better generalization. The catch is that the smoothing strength is usually a fixed hyperparameter, tuned by trial and error, even though the need for regularization changes dramatically over the course of training. SLAW makes that strength dynamic. It computes a first-order proxy of local sharpness, an inexpensive estimate of how steeply the loss rises around the current parameters, and uses it to modulate the smoothing applied at each step.

The intuition behind this adaptive scheme is geometric. Early in training, when gradients are large and the optimizer is traversing high-curvature regions of the loss landscape, the network benefits from strong regularization that discourages it from diving into narrow valleys. As training stabilizes and the parameters approach flatter basins, aggressive smoothing becomes counterproductive, blurring the decision boundaries the model needs to fit the data. SALS therefore applies heavier smoothing in steep, high-curvature areas and gradually relaxes it as the landscape flattens. Because the sharpness estimate relies only on first-order information rather than the full Hessian, which would be computationally prohibitive for modern architectures, the method keeps the overhead low enough for routine use. The authors report that this component improves not only generalization accuracy but also predictive calibration, addressing the well-documented tendency of modern networks to be confidently wrong.

The second component, Loss-Adaptive Reweighting, addresses the label-noise problem from a different angle. Deep networks have a well-known and troubling habit: given enough epochs, they memorize even blatantly wrong labels, because the capacity of these models is sufficient to fit arbitrary noise. Memorization typically shows up late in training, when the loss on corrupted samples that the network has already fit collapses toward zero while the loss on genuinely hard, clean examples remains elevated. LAW exploits this statistical signature. Within each mini-batch, it standardizes the per-sample loss values and flags samples whose losses are anomalous relative to their batch peers, then down-weights their contribution to the parameter update. Corrupted labels, which the model fits quickly and confidently, thus lose influence before they can drag the shared weights in the wrong direction.

What distinguishes LAW from the crowded field of noise-robust training methods is what it does not require. Many existing approaches, such as co-teaching, loss correction, and joint methods like DivideMix, depend on estimates of the noise rate or assumptions about the structure of the noise, for example that it is symmetric and flips labels uniformly across classes. In practice, real-world annotation errors are rarely so well-behaved, and misestimating the noise rate can be as damaging as ignoring it. LAW involves no prior knowledge of noise characteristics and no noise estimation rate at all. It simply reacts to the empirical distribution of losses in each batch, making it applicable to datasets whose corruption level is unknown or heterogeneous, which is the norm rather than the exception in web-scraped and crowd-annotated data.

The experimental evaluation covers three standard benchmarks: MNIST, CIFAR-10 and CIFAR-100. On clean data, SLAW matches conventional training algorithms, an important sanity check showing that the added machinery does not sacrifice baseline performance. The more striking results appear under synthetic label corruption. At moderate noise levels, with twenty percent of labels flipped symmetrically, SLAW nearly eliminates the catastrophic overfitting that typically afflicts late training, when standard models begin absorbing the corrupted labels and their test accuracy degrades. At extreme noise of forty percent, a regime in which the baseline models usually collapse entirely, SLAW continues to learn stably and retains meaningful generalization, a result the authors present as evidence that the framework can function where most pipelines fail outright.

Ablation studies disentangle the contributions of the two components and reveal a clean division of labor. SALS, the sharpness-adaptive smoothing, is the primary driver of improved generalization and calibration on clean and moderately noisy data, consistent with the landscape-flattening role of label smoothing identified in earlier analyses by Müller, Kornblith and Hinton. LAW, by contrast, is the main engine of robustness to label noise, doing the heavy lifting when corrupted samples threaten to poison the updates. The synergy matters because the two failure modes interact: noisy labels tend to push optimizers toward sharp minima, since memorizing outliers often requires fitting narrow spikes in the loss surface, so a method that addresses only one vulnerability leaves the other exposed.

The work sits at the intersection of several research threads that have matured over the past decade. Sharpness-aware minimization, introduced by Foret and colleagues, explicitly optimizes for flat minima by perturbing weights to worst-case nearby points before taking a gradient step, and adaptive variants such as ASAM refined the idea. Sample reweighting has its own lineage, from MentorNet’s learned curricula to meta-learned example weighting by Ren and co-workers and recent dynamic loss-based reweighting for large language model pretraining. SLAW’s contribution is less a single new primitive than a demonstration that these ideas can be fused cheaply and adaptively, with the sharpness signal controlling regularization strength and the loss statistics controlling sample influence, without the two mechanisms interfering with each other.

For practitioners, the appeal lies in the framework’s economy of assumptions. Real datasets are messy: medical registries contain transcription errors, web labels reflect the biases of their annotators, and crowdsourced annotations disagree in ways no symmetric noise model captures. A training method that neither assumes a known noise rate nor demands expensive second-order curvature computations fits the constraints of applied machine learning, where the authors argue such techniques are most needed. The code has been released on GitHub, lowering the barrier to independent replication, and the benchmarks, while standard, are the same ones on which the field’s noisy-label methods are conventionally compared.

The broader significance may lie in what SLAW implies about the relationship between optimization geometry and data quality. The results support a view in which these are not separate concerns to be handled by separate tools, but coupled aspects of a single training dynamic: the shape of the loss landscape determines how eagerly a network memorizes bad examples, and the composition of the data determines which minima the optimizer can reach. A framework that senses both, adjusting its regularization as the terrain flattens and its trust in each sample as the loss statistics shift, offers a template for the kind of self-correcting training that reliable, real-world deep learning systems will likely require as models are deployed on data that is, as the authors put it, imperfect, noisy, or of low quality.

Subject of Research: A sharpness- and loss-adaptive weighting framework for robust deep learning under label noise

Article Title: Slaw: sharpness- and loss-adaptive weighting for robust deep learning

Article References: Slaw: sharpness- and loss-adaptive weighting for robust deep learning. (n.d.). https://doi.org/10.1007/s41060-026-01272-w

Image Credits: AI Generated

DOI: 10.1007/s41060-026-01272-w

Keywords: deep learning, label noise, loss landscape, sharp minima, label smoothing, sample reweighting, generalization, calibration, robustness, sharpness-aware minimization, neural networks, CIFAR-10

Cite Scienmag News

Cassandra Pierce. (October 2, 2026). New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels. Scienmag. https://scienmag.com/new-training-method-steers-neural-networks-toward-flat-minima-and-away-from-bad-labels/

Cassandra Pierce. "New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels." Scienmag, 2 October 2026, https://scienmag.com/new-training-method-steers-neural-networks-toward-flat-minima-and-away-from-bad-labels/. Accessed 2 October 2026.

Cassandra Pierce. "New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels." Scienmag. October 2, 2026. https://scienmag.com/new-training-method-steers-neural-networks-toward-flat-minima-and-away-from-bad-labels/

Tags: adaptive loss functions in machine learningcalibrationCIFAR-10deep learningflat minima in deep learninggeneralizationgeneralization in deep neural networkslabel noiselabel smoothinglabel smoothing techniques in neural networksloss landscapeloss landscape geometrymethods for avoiding sharp minimamini-batch training strategiesneural network reliability in real-world applicationsneural network training optimizationneural networksovercoming overconfidence in classification modelsrobustnessrobustness to mislabeled datasample reweightingsharp minimaSharpness- and Loss-Adaptive Weighting (SLAW)sharpness-aware minimization
Share26Tweet16
Previous Post

Machine Learning Pinpoints Ten-Gene Signature Behind Diabetic Wounds That Refuse to Heal

Next Post

Study Finds Inefficient Natural Gas Combustion Drives NYC Methane Emissions and Financial Losses

Related Posts

Transformer-based model outperforms CNNs in tomato disease detection study
Technology and Engineering

Transformer-based model outperforms CNNs in tomato disease detection study

October 2, 2026
Researchers propose lightweight DDMnet for efficient instance segmentation
Technology and Engineering

Researchers propose lightweight DDMnet for efficient instance segmentation

October 2, 2026
New Compression Method Proposed for Fetal Heart Sound Signals
Technology and Engineering

New Compression Method Proposed for Fetal Heart Sound Signals

October 2, 2026
Neural Networks Crack the Secrets of Dual-Non-Newtonian Fluid Flow Under Magnetic and Bioconvective Forces
Technology and Engineering

Neural Networks Crack the Secrets of Dual-Non-Newtonian Fluid Flow Under Magnetic and Bioconvective Forces

October 2, 2026
UCSB Engineer Bolin Liao to Develop Platform for Imaging Quantum Interactions in 2D Materials
Technology and Engineering

UCSB Engineer Bolin Liao to Develop Platform for Imaging Quantum Interactions in 2D Materials

October 2, 2026
When Breaking the Speed Limit Saves Lives: The Ethical Dilemma of Self-Driving Cars
Technology and Engineering

When Breaking the Speed Limit Saves Lives: The Ethical Dilemma of Self-Driving Cars

October 2, 2026
Next Post
Study Finds Inefficient Natural Gas Combustion Drives NYC Methane Emissions and Financial Losses

Study Finds Inefficient Natural Gas Combustion Drives NYC Methane Emissions and Financial Losses

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Transformer-based model outperforms CNNs in tomato disease detection study
  • Audio Diaries Reveal What Shapes Surgical Residents’ Learning in the Clinic
  • Study Finds Inefficient Natural Gas Combustion Drives NYC Methane Emissions and Financial Losses
  • New Training Method Steers Neural Networks Toward Flat Minima and Away from Bad Labels

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading