Colorectal cancer remains one of the deadliest and most common malignancies in the world, with roughly 1.9 million new cases recorded globally in 2020. When caught early, the outlook is dramatically better: data from the US National Cancer Institute put the five-year survival rate for early-stage disease at around 90 percent. The problem is that diagnosis still depends heavily on histopathology, in which a pathologist examines hematoxylin and eosin stained tissue under a microscope. The method is considered the gold standard, but it is slow, expensive to scale, and vulnerable to the natural variability of human interpretation, particularly when different tissue types look deceptively similar under the lens.
A team of researchers in India has now built an artificial intelligence system designed to ease that burden, and their results suggest that combining two very different kinds of neural networks, then letting nature-inspired algorithms tune the combination, can push diagnostic accuracy to a striking level. Writing in the journal Discover Artificial Intelligence, the group reports that their hybrid model classified colorectal histopathology images with 96.27 percent accuracy, up from 92.13 percent for the same architecture before optimization. The work was led by Hemanth K S, Samiksha Sandeep Zokande and Prathana Sharma of CHRIST University in Bangalore, together with N Kartik of the Manipal Academy of Higher Education.
The central technical insight behind the study is that convolutional neural networks and vision transformers have complementary blind spots. Convolutional models such as EfficientNetB0 excel at detecting local texture and structure, because their filters scan small neighborhoods of pixels and build up hierarchical features layer by layer. But convolution is inherently local: even as the receptive field grows with depth, the network struggles to connect distant regions of an image. In histopathology that is a serious limitation, because diagnostically meaningful patterns, such as the relationship between glandular formations, stromal tissue and cell nuclei, can span an entire tissue patch.
Vision transformers solve that problem with self-attention, a mechanism borrowed from natural language processing that lets a model relate any patch of an image to any other patch, regardless of distance. The researchers chose BEiT, a transformer pretrained through masked image modeling, in which patches of an image are hidden and the network must reconstruct them from context, a strategy adapted from BERT in language. That pretraining gives BEiT rich visual representations even when labeled medical data is scarce. Yet transformers carry their own costs: they are computationally demanding and notoriously sensitive to hyperparameter settings, which means their performance can collapse if the learning rate, dropout or weight decay are chosen poorly.
The team’s solution was to cascade the two architectures rather than fuse them in parallel. BEiT first processes the image and produces a 768-dimensional embedding of its global context, which is then projected into the input space of EfficientNetB0 for convolutional refinement and final classification. The authors tested a parallel fusion of the two feature streams early on and found it inflated the parameter count by roughly 60 percent without improving validation accuracy, so they settled on the sequential design, which also made it easier to freeze and unfreeze parts of the pipeline during tuning.
That tuning is where the biology comes in. Instead of hand-picking hyperparameters or exhaustively searching a grid, the researchers deployed two bio-inspired metaheuristics. The Whale Optimization Algorithm mimics the bubble-net feeding behavior of humpback whales, which spiral upward beneath prey while releasing bubbles; mathematically, candidate solutions either encircle the best solution found so far or spiral toward it along a shrinking helix, balancing exploration and exploitation. Particle Swarm Optimization, meanwhile, models a flock of particles that each remember their own best position and are pulled toward the swarm’s global best, blending inertia, personal experience and social learning.
Both algorithms searched the same four-dimensional space of learning rate, weight decay, dropout rate and fully connected layer width, using 15 candidates over up to 30 iterations, with the BEiT backbone frozen during candidate evaluation to keep computation manageable. Once the best configuration was found, the full model was retrained from scratch with the transformer unfrozen. Notably, the two optimizers converged on different settings, with WOA favoring smaller learning rates and dropout values while PSO preferred wider classification layers, evidence that they explored genuinely different regions of the search space.
The experiments used a publicly available benchmark of 5,000 H&E stained tissue patches, evenly divided among eight classes: tumor epithelium, simple stroma, complex stroma, lymphocytes, debris, mucosa, adipose tissue and background. Images were resized to 224 by 224 pixels, contrast-enhanced with histogram equalization, denoised with Gaussian blur and normalized to ImageNet statistics. The team deliberately skipped data augmentation after finding it ballooned the dataset to roughly 25,000 images and pushed training time from about 8 hours to more than 14, an unacceptable cost when the optimizer must evaluate hundreds of candidate configurations.
The payoff was clear across every metric. PSO-optimized training reached 96.27 percent accuracy and cut the standard deviation across repeated runs by 45.56 percent, while WOA achieved 96.00 percent with a 23 percent reduction, both comfortably ahead of the unoptimized baseline. Confusion matrices showed substantial gains on the hardest classes: the complex stroma class improved its F1 score from 0.900 to 0.962 under PSO, and stroma rose from 0.892 to 0.956. Area under the ROC curve approached unity for most classes, with tumor and adipose tissue reaching 0.997 or higher, and mucosa hitting a perfect 1.000. Critically, false negatives, the most dangerous error in cancer diagnostics, dropped sharply for nearly every tissue type.
The authors are candid about the road ahead. The benchmark, though balanced and well curated, cannot capture the staining variability and scanner differences of real multi-center clinical data, and the black-box nature of deep networks remains an obstacle to clinical adoption, one that explainability tools such as Grad-CAM and SHAP may eventually address. Future work could extend the framework to multi-modal inputs combining genomic or proteomic data, or test alternative metaheuristics such as the Grey Wolf Optimizer. Still, the study makes a compelling case that when transformers see the whole picture, convolutions see the fine detail, and whale songs and bird flocks handle the tuning, automated pathology moves a meaningful step closer to the clinic.
Subject of Research: A hybrid BEiT and EfficientNetB0 deep learning model optimized with bio-inspired algorithms for classifying colorectal cancer histopathology images
Article Title: Optimized hybrid BEiT and EfficientNetB0 deep learning model with whale and particle swarm algorithms for colorectal cancer histopathological image classification
Article References: K S, H., Kartik, N., Zokande, S. S., & Sharma, P. (2026). Optimized hybrid BEiT and EfficientNetB0 deep learning model with whale and particle swarm algorithms for colorectal cancer histopathological image classification. Discover Artificial Intelligence, 6(1), Article 1295. https://doi.org/10.1007/s44163-026-02213-z
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02213-z
Keywords: colorectal cancer, histopathology, BEiT, EfficientNetB0, vision transformer, Whale Optimization Algorithm, Particle Swarm Optimization, deep learning, medical imaging, hyperparameter optimization, image classification, computational pathology
Cite Scienmag News
Nathaniel Bowman. (October 4, 2026). Whales and Swarms Help AI Read Cancer Slides With Record Accuracy. Scienmag. https://scienmag.com/whales-and-swarms-help-ai-read-cancer-slides-with-record-accuracy/
Nathaniel Bowman. "Whales and Swarms Help AI Read Cancer Slides With Record Accuracy." Scienmag, 4 October 2026, https://scienmag.com/whales-and-swarms-help-ai-read-cancer-slides-with-record-accuracy/. Accessed 4 October 2026.
Nathaniel Bowman. "Whales and Swarms Help AI Read Cancer Slides With Record Accuracy." Scienmag. October 4, 2026. https://scienmag.com/whales-and-swarms-help-ai-read-cancer-slides-with-record-accuracy/

