Deep neural networks have become the workhorses of modern artificial intelligence, powering everything from image recognition on smartphones to diagnostic tools in hospitals. Yet their power comes at a price: these models are often bloated with millions of connections, many of which contribute little to the final answer. Every redundant weight still has to be stored, moved through memory, and multiplied during computation, which translates directly into wasted energy, slower responses, and a larger carbon footprint. A new study published in PLOS Complex Systems by Amirhossein Douzandeh Zenoozi, Laura Erhan, Antonio Liotta, and Lucia Cavallaro takes a hard, quantitative look at this problem and arrives at a conclusion that could reshape how engineers design efficient networks: the way you decide which connections to cut matters just as much as how many you cut.
The technique at the heart of the research is pruning, a strategy inspired by an old observation in neuroscience that biological brains prune synapses as they mature. In artificial networks, pruning means removing weights, or entire groups of weights, that have been identified as unimportant, leaving behind a sparser, leaner model. The idea has been around for decades, but its practical benefits depend heavily on implementation details. The research team compared two fundamentally different strategies. Global pruning treats the whole network as a single pool of weights and removes the least important ones wherever they happen to sit, which can leave some layers almost untouched while hollowing out others. Layer-wise pruning, by contrast, applies the sparsity target separately to each hidden layer, guaranteeing that every part of the network contributes to the reduction.
To test which strategy delivers better real-world efficiency, the researchers ran a systematic set of experiments across two multilayer perceptron models, one small-scale and one medium-scale, and then extended the analysis to VGG-16, a well-known convolutional neural network architecture, to check whether the findings generalize beyond simple fully connected networks. The models were evaluated on five widely used datasets spanning different domains and difficulty levels: MNIST and FashionMNIST for handwritten digits and clothing images, EMNIST for extended character recognition, CIFAR-10 for more complex natural images, and OctMNIST, a medical imaging dataset of optical coherence tomography scans. This breadth matters, because a pruning method that only works on toy problems would be of limited value for practitioners deploying models in medicine or embedded systems.
The sparsity levels examined were 50 percent and 80 percent, meaning half or four-fifths of the weights were removed, alongside the unpruned dense networks serving as the benchmark at 0 percent sparsity. The evaluation went well beyond the usual accuracy metric. The team measured inference time, the latency a user experiences when the trained model makes a prediction, and inference energy consumption, the electrical cost of every classification. They also tracked training energy use, estimated carbon dioxide emissions associated with training, and peak memory usage during training. Together these metrics paint a fuller picture of what it actually costs to build and run a neural network, costs that are often hidden when researchers report only top-line accuracy figures.
The results were strikingly consistent. Layer-wise pruning emerged as the clear winner, offering the best trade-offs by reliably reducing both inference time and inference energy while keeping accuracy essentially intact. One headline example illustrates the scale of the gains: when the small-scale model was trained on MNIST at 50 percent sparsity, inference energy dropped by 33 percent and inference time fell by 33 percent as well, while accuracy declined by a negligible 0.49 percent. In other words, the network delivered nearly identical answers at two-thirds of the computational cost. For battery-powered devices, edge sensors, or data centers processing millions of inferences per day, savings of that magnitude compound into substantial economic and environmental benefits.
Why does layer-wise pruning outperform the global approach? The answer lies in how each method distributes its cuts. Global pruning follows the magnitude of weights wherever it leads, and in practice this often means some layers are pruned far more aggressively than others. A layer that loses most of its connections can become a bottleneck or lose critical representational capacity, while layers left dense continue to impose their full memory and computation costs. Layer-wise pruning avoids this imbalance by enforcing the sparsity target uniformly across hidden layers, so the computational savings are realized everywhere in the network rather than concentrated in a few heavily thinned regions. The result is a model whose reduced cost is matched by preserved function, which explains the consistent advantage the researchers observed across datasets, model sizes, and architectures.
The extension to VGG-16 is particularly important for the credibility of these findings. Convolutional neural networks like VGG-16 have a very different structure from multilayer perceptrons, with convolutional filters, pooling stages, and far larger parameter counts, so a pruning strategy that only worked on simple perceptrons would raise doubts about its generality. By showing that the layer-wise approach retains its advantages on a representative CNN, the study strengthens the case that the principle applies across the architectures practitioners actually deploy. The same pattern held when the team turned to the sustainability metrics: training energy consumption, estimated CO2 emissions, and peak memory usage all favored the layer-wise strategy over global pruning, reinforcing the recommendation from every angle the researchers measured.
The environmental dimension of the study deserves particular emphasis. Training large models has become a recognized contributor to carbon emissions, and inference at scale, where a deployed model may answer billions of queries, can dwarf the training footprint over a model’s lifetime. By quantifying CO2 estimates alongside energy and memory metrics, the authors connect an abstract algorithmic choice to concrete sustainability outcomes. Their findings suggest that organizations seeking to reduce the carbon intensity of their AI systems do not necessarily need exotic hardware or entirely new model families; a disciplined pruning strategy applied layer by layer can deliver meaningful reductions with minimal loss in performance, making it one of the more accessible levers available for greener machine learning.
The practical implications extend across the deployment spectrum. On edge devices such as smartphones, wearables, and medical sensors, reduced inference time and energy translate into longer battery life and faster, more responsive applications. In data centers, where inference workloads run continuously, the same savings scale into lower electricity bills and reduced cooling demands. Even during development, the lower peak memory usage of pruned networks means researchers can train and iterate on models with more modest hardware, lowering the barrier to entry for smaller labs and companies. Because the study demonstrates these benefits across five datasets and multiple architectures, practitioners have reason to expect similar trade-offs in their own domains, from character recognition to medical image analysis.
Overall, the study makes a compelling case that layer-wise pruning is a practical, immediately usable approach for designing energy-efficient neural networks. Rather than treating accuracy and efficiency as opposing goals locked in a zero-sum contest, the research shows that a carefully structured sparsity strategy can deliver large reductions in inference time, inference energy, training energy, emissions estimates, and memory usage while sacrificing almost nothing in predictive performance. As artificial intelligence continues to spread into every corner of technology and society, work like this points toward a future in which the networks we rely on are not only intelligent but also lean, frugal, and environmentally responsible, proving that sometimes the smartest thing a neural network can do is know which of its connections it can live without.
Subject of Research: Energy-efficient pruning strategies for multilayer perceptron and convolutional neural networks
Article Title: Inference and training efficiency in pruned multilayer perceptron networks
Article References: Douzandeh Zenoozi, A., Erhan, L., Liotta, A., & Cavallaro, L. (2026). Inference and training efficiency in pruned multilayer perceptron networks. PLOS Complex Systems, 3(3), e0000095. https://doi.org/10.1371/journal.pcsy.0000095
Image Credits: AI Generated
DOI: 10.1371/journal.pcsy.0000095
Keywords: neural networks, pruning, sparsity, energy efficiency, inference, layer-wise pruning, global pruning, multilayer perceptron, VGG-16, CO2 emissions, edge computing, machine learning
Cite Scienmag News
Blake Davidson. (October 10, 2026). Trimming the Fat: Layer-Wise Pruning Slashes Neural Network Energy Use. Scienmag. https://scienmag.com/trimming-the-fat-layer-wise-pruning-slashes-neural-network-energy-use/
Blake Davidson. "Trimming the Fat: Layer-Wise Pruning Slashes Neural Network Energy Use." Scienmag, 10 October 2026, https://scienmag.com/trimming-the-fat-layer-wise-pruning-slashes-neural-network-energy-use/. Accessed 10 October 2026.
Blake Davidson. "Trimming the Fat: Layer-Wise Pruning Slashes Neural Network Energy Use." Scienmag. October 10, 2026. https://scienmag.com/trimming-the-fat-layer-wise-pruning-slashes-neural-network-energy-use/

