Cybersecurity researchers have unveiled a machine learning framework that can identify malicious software with flawless accuracy while relying on just four memory features instead of the dozens typically required. The study, published in Discover Artificial Intelligence, combines dynamic malware analysis with a structurally enhanced version of the Harris Hawks Optimizer, a nature-inspired algorithm that mimics the cooperative hunting tactics of one of the animal kingdom’s most agile predators. Tested on the widely used CIC-MalMem-2022 dataset, the approach pushed the Extra Trees classifier to a perfect 100 percent accuracy, precision, and F1-score, while cutting convergence times across every classifier it touched.
The scale of the malware problem provides the backdrop for this work. According to the Symantec Internet Security Threat Report cited by the authors, more than 430 million new malware samples have been discovered, with an average of 1.2 million malware attacks launched every day. Traditional detection methods, which often depend on matching known signatures, struggle against new, unknown, or deliberately obfuscated attacks. As attackers grow more sophisticated, defenders need techniques that can spot malicious behavior even when the code has never been seen before or is designed to lie dormant until it is safe to strike.
Dynamic Malware Analysis, or DMA, offers one promising answer. Rather than examining a program’s static code, DMA runs the suspicious program inside a controlled environment and watches what it actually does: how it alters memory spaces, injects processes, makes unusual system or API calls, modifies the registry, and interacts with the network. Because it observes behavior rather than appearance, DMA can catch malware that evades signature-based tools, including obfuscated variants and fileless threats that hide in memory. The system memory data collected during execution is then handed to a machine learning classifier, which decides whether the activity looks malicious or benign.
But DMA carries a heavy computational burden. Memory dumps are enormous, and the features extracted from them frequently include redundant or irrelevant information that inflates processing time, consumes resources, and can even degrade classification accuracy. This is where feature selection comes in. By identifying the most informative subset of features, feature selection reduces dimensionality, lowers computational cost, and reduces the risk of overfitting, in which a model memorizes noise in the training data instead of learning the underlying patterns that generalize to new samples. Metaheuristic optimization algorithms have proven particularly effective at navigating the vast combinatorial space of possible feature subsets.
The Harris Hawks Optimizer, first introduced as a metaheuristic for solving optimization problems, models candidate solutions as hawks searching for prey, where the prey represents the optimal solution. The algorithm alternates between exploration, in which hawks perch randomly or position themselves relative to the best solution found so far, and exploitation, in which they converge on promising regions using strategies such as soft siege, hard siege, and progressive rapid dives modeled by Levy flight mathematics. A key transition parameter, the prey’s escaping energy, decreases over iterations and governs when the search shifts from broad exploration to focused refinement. While HHO has already been applied successfully to high-dimensional DMA data, the authors note it can suffer from premature convergence, settling on suboptimal feature subsets without adequately exploring the search space.
The team’s enhancement addresses this weakness with two structural modifications. First, they replaced the standard random initialization with a hybrid scheme: half of the initial hawk population is generated using Latin Hypercube Sampling, a statistical technique that divides the solution space into segments and samples from each, ensuring a more uniform spread of starting points and faster convergence toward the global optimum. Second, they integrated a Gaussian crossover mechanism borrowed from evolutionary computation. After each iteration, pairs of hawks are blended using Gaussian distributions to produce offspring that inherit traits from both parents, and the two fittest hawks are also crossed to generate elite offspring. The population is then trimmed back to its original size by keeping only the best individuals, balancing diversity with convergence.
The experimental pipeline began with the CIC-MalMem-2022 dataset, created at the Canadian Institute for Cybersecurity. It contains 58,596 memory dump records, evenly split between benign and malicious samples, described by 55 features extracted with the VolMemLyzer-V2 analyzer from VirtualBox memory snapshots. These features capture injected code, file access, registry activity, process behavior, and redirected API calls. The researchers label-encoded the class column, applied Min-Max normalization to forty features whose raw values ranged up to 1.05 million, and then ran their improved HHO with a population of 26 hawks over 260 iterations, using a fitness function weighted toward accuracy with a small penalty for the number of selected features.
The results were striking. The improved optimizer distilled 55 features down to just four: the number of threads, the number of timers, the number of services detected in memory, and the number of callbacks, a 92.7 percent reduction in dimensionality. Among these, the service count feature showed the highest permutation importance at 0.35. The best feature subset was found roughly 40 iterations into the search. When the reduced feature set was fed to six supervised classifiers, Extra Trees achieved perfect 100 percent scores, Random Forest reached 99.98 percent, and Decision Trees, XGBoost, Gradient Boosted Trees, and AdaBoost followed closely behind. Compared with the original HHO, accuracy improved across the board, for example rising from 99.96 to 100 percent for Extra Trees and from 99.93 to 99.98 percent for Random Forest, while convergence times fell for every classifier, with Decision Trees dropping from 0.0312 to 0.0156 seconds.
The improvements were not statistical flukes. A paired Wilcoxon signed-rank test over 20 independent runs comparing the improved and original optimizers on Extra Trees accuracy yielded a test statistic of zero and a p-value of 1.9 times ten to the negative sixth, confirming significance at the 0.05 level, with a large effect size of 0.87. Against other published models built on the same CIC-MalMem-2022 dataset, which reported accuracies between 97.67 and 99.97 percent, the proposed framework achieved the highest reported binary classification accuracy while using a dramatically smaller feature subset. The authors also analyzed computational complexity, noting that the crossover operations add minimal overhead relative to the dominant cost of classifier training, and that the four-feature subset makes deployment far lighter than the original 55-feature space.
The study is not without acknowledged limitations. All experiments were conducted on a single dataset, evaluation used one stratified train-test split rather than cross-validation, and variations in sandbox configurations and memory acquisition environments were not explored. The authors say future work will add k-fold cross-validation, external dataset validation, and comparisons with other enhanced HHO variants. Still, the core message is compelling: by letting a flock of virtual hawks hunt through a high-dimensional feature space with smarter initialization and genetic recombination, the researchers showed that near-perfect malware detection does not require mountains of data, only the right handful of behavioral clues.
Subject of Research: Feature selection with an enhanced Harris Hawks Optimizer for dynamic malware detection using memory analysis
Article Title: Malware detection using dynamic analysis with a modified Harris Hawk optimizer
Article References: Anabousi, H. A., Abualhaj, M. M., Abu-Shareha, A., Yousif, M., Daoud, M. S., Faheem, M. R., & Soltani, H. (2026). Malware detection using dynamic analysis with a modified Harris Hawk optimizer. Discover Artificial Intelligence, 6(1), Article 1309. https://doi.org/10.1007/s44163-026-01359-0
Image Credits: AI Generated
DOI: 10.1007/s44163-026-01359-0
Keywords: malware detection, dynamic malware analysis, Harris Hawks Optimizer, feature selection, machine learning, metaheuristics, CIC-MalMem-2022, cybersecurity, Extra Trees, memory forensics, optimization, classification
Cite Scienmag News
Denise Maddox. (September 30, 2026). Hawk-Inspired Algorithm Slashes Malware Data to Four Features for Perfect Detection. Scienmag. https://scienmag.com/hawk-inspired-algorithm-slashes-malware-data-to-four-features-for-perfect-detection/
Denise Maddox. "Hawk-Inspired Algorithm Slashes Malware Data to Four Features for Perfect Detection." Scienmag, 30 September 2026, https://scienmag.com/hawk-inspired-algorithm-slashes-malware-data-to-four-features-for-perfect-detection/. Accessed 30 September 2026.
Denise Maddox. "Hawk-Inspired Algorithm Slashes Malware Data to Four Features for Perfect Detection." Scienmag. September 30, 2026. https://scienmag.com/hawk-inspired-algorithm-slashes-malware-data-to-four-features-for-perfect-detection/

