Deep learning has been built, almost without interruption, on a single computational habit: the backward pass. Since the popularization of back-propagation in the 1980s, neural networks have learned by pushing error signals backwards through their layers, a process that demands that every intermediate activation be stored during the forward pass so that gradients can be computed on the way back. That memory bill has become one of the defining constraints of modern artificial intelligence, shaping what can be trained and on what hardware. Now, a study published in Artificial Intelligence Review by Yomna Abdulgawad Elnady, Amany Mahmoud Sarhan and Mohammad Ali Eita of Tanta University in Egypt examines what happens when the backward pass is abandoned altogether, and shows that with the right automated tuning, the alternative can come surprisingly close to matching the incumbent.
The alternative in question is the Forward-Forward (FF) algorithm, introduced by Geoffrey Hinton, one of the founding figures of deep learning. Where back-propagation performs one forward pass to generate predictions and one backward pass to distribute blame, FF performs two forward passes. The first runs on positive, real data drawn from the training set; the second runs on negative data, which may be generated by the network itself. Each pass produces a measure of goodness at every layer, and layers are trained to assign high goodness to positive data and low goodness to negative data. Because no gradients need to flow backwards, activations do not need to be cached for a reverse sweep, and the memory footprint of training shrinks dramatically.
That memory efficiency is precisely what makes FF interesting for a particular and growing class of problems: training neural networks on low-resource hardware. Edge devices, embedded controllers and medical monitoring systems often cannot afford the memory overhead of back-propagation, which scales with the number of stored activations. FF’s two-pass structure, by contrast, opens the door to on-device learning without the same storage demands. The catch, as the Tanta University team documents across six diverse datasets, is that FF in its basic form lags behind back-propagation in predictive accuracy. The algorithm trades memory for performance, and until now that trade has been steep enough to keep it largely a research curiosity rather than a practical training method.
The new study attacks that gap not by redesigning the algorithm itself but by treating its configuration as an optimization problem in its own right. Neural network training is governed by a constellation of hyper-parameters, settings such as learning rates, layer configurations and goodness thresholds that are fixed before training begins and that can make the difference between a model that converges gracefully and one that stalls. In conventional practice these values are chosen by hand, guided by intuition and trial and error. The researchers instead deployed automated hyper-parameter optimization, using metaheuristic search algorithms, computational procedures inspired by natural processes that explore large configuration spaces far more systematically than manual tuning ever could.
The scope of the evaluation is notable. The team first assessed FF’s baseline performance across six diverse datasets to establish how the algorithm behaves in general. They then narrowed to four datasets for the optimization phase, applying two well-established metaheuristic algorithms alongside four state-of-the-art ones, a total of six search strategies competing to find the hyper-parameter settings that would squeeze the most performance out of FF. This comparative design matters, because the choice of optimizer for hyper-parameter search is itself a hyper-parameter of sorts, and the study’s structure allows the relative strength of different search strategies to be assessed on the same footing.
The results are striking. On Pneumonia-MedMnist, a medical imaging benchmark, validation accuracy rose from 93 percent to 97 percent after optimization. On Fashion-MNIST, a standard clothing-image classification set, accuracy climbed from 87 percent to 91 percent. The most dramatic gains appeared on the organ-segmentation datasets from the MedMnist collection: OrganC-MedMnist jumped from 82 percent to 94 percent, a twelve-point improvement, and OrganA-MedMnist rose from 88 percent to 95 percent. Crucially, these gains came while maintaining or even reducing training time, meaning the optimization did not simply buy accuracy by burning more compute. The study also reports that metaheuristic optimization accelerated convergence, allowing the networks to reach good solutions faster than they otherwise would.
Those numbers deserve unpacking, because they speak to a broader tension in machine learning research. The gap between memory efficiency and accuracy has long been framed as a fundamental trade-off: algorithms that avoid back-propagation’s memory bill pay for it in performance. What this study demonstrates is that at least part of that gap is not fundamental at all, but an artifact of under-tuned configuration. FF’s hyper-parameter space is evidently sensitive enough that default or manually chosen settings leave substantial performance on the table. When a systematic search is applied, much of that lost performance is recovered, in one case recovering nearly the entire distance to the sort of accuracy figures typically associated with conventionally trained networks on these benchmarks.
The technical explanation for why hyper-parameters matter so much to FF lies in its layered goodness objective. Because each layer is trained locally to discriminate positive from negative data, the overall behavior of the network emerges from the interaction of many local objectives rather than from a single global loss gradient. Small changes in threshold settings or layer-wise learning dynamics can therefore ripple through the network in ways that are hard to anticipate manually. Metaheuristic algorithms, which iteratively propose, evaluate and refine candidate solutions using mechanisms borrowed from processes such as evolution or swarm behavior, are well suited to navigating such rugged configuration landscapes, where the relationship between settings and outcomes is nonlinear and riddled with local optima.
The implications extend beyond a single algorithm. Interest in back-propagation-free training has grown as the field confronts the energy and memory costs of ever-larger models, and as researchers look to the brain, which clearly learns without any known reverse-phase error propagation, for inspiration. Hinton’s Forward-Forward proposal was framed partly in these terms, as a step toward learning procedures that might be more biologically plausible and more hardware-friendly. The Egyptian team’s contribution is to show that the practical viability of such procedures depends heavily on the surrounding optimization machinery, and that investing in automated tuning can transform a promising but underperforming idea into a competitive one.
For practitioners working at the edge, the message is concrete. Medical imaging tasks of the kind tested here, including pneumonia detection and organ identification, are exactly the workloads targeted at resource-constrained clinical hardware, and the demonstrated accuracy levels after optimization, reaching 94 to 97 percent validation accuracy on several benchmarks, suggest that memory-efficient training is approaching the point where it can be taken seriously for such applications. The study, published open access on 30 September 2026, received open access funding from Egypt’s Science, Technology and Innovation Funding Authority in cooperation with the Egyptian Knowledge Bank. Whether Forward-Forward can scale to the large architectures that dominate contemporary artificial intelligence remains an open question, but this work establishes that the algorithm’s accuracy deficit is neither fixed nor mysterious, and that the tools for closing it were already sitting in the metaheuristic toolbox.
Subject of Research: Optimizing the Forward-Forward neural network training algorithm with metaheuristic hyper-parameter search
Article Title: Towards optimized forward–forward training for neural networks
Article References: Elnady, Y. A., Sarhan, A. M., & Eita, M. A. (2026). Towards optimized forward–forward training for neural networks. Artificial Intelligence Review, 59(11), Article 232. https://doi.org/10.1007/s10462-026-11668-6
Image Credits: AI Generated
DOI: 10.1007/s10462-026-11668-6
Keywords: Forward-Forward algorithm, back-propagation, neural networks, metaheuristics, hyper-parameter optimization, Geoffrey Hinton, memory efficiency, machine learning, MedMnist, Fashion-MNIST, Tanta University, Artificial Intelligence Review
Cite Scienmag News
Blake Davidson. (September 30, 2026). Hinton’s Forward-Forward Algorithm Gets a Hyper-Parameter Tune-Up That Closes the Accuracy Gap. Scienmag. https://scienmag.com/hintons-forward-forward-algorithm-gets-a-hyper-parameter-tune-up-that-closes-the-accuracy-gap/
Blake Davidson. "Hinton’s Forward-Forward Algorithm Gets a Hyper-Parameter Tune-Up That Closes the Accuracy Gap." Scienmag, 30 September 2026, https://scienmag.com/hintons-forward-forward-algorithm-gets-a-hyper-parameter-tune-up-that-closes-the-accuracy-gap/. Accessed 30 September 2026.
Blake Davidson. "Hinton’s Forward-Forward Algorithm Gets a Hyper-Parameter Tune-Up That Closes the Accuracy Gap." Scienmag. September 30, 2026. https://scienmag.com/hintons-forward-forward-algorithm-gets-a-hyper-parameter-tune-up-that-closes-the-accuracy-gap/

