Tomato growers may soon have a diagnostician in their pocket that rivals the experts, after researchers unveiled an artificial intelligence system capable of identifying tomato leaf diseases with near-perfect accuracy while explaining its reasoning in human-readable heatmaps. In a study published in the journal Plant Methods, a team led by Ran Wang and Xiao Yu of Shandong University of Technology in China describes DAPR-AM-Net, an end-to-end smart farming framework that combines dual-attention progressive refinement with a novel adaptive data augmentation scheme to classify and forecast tomato leaf diseases under real-world field conditions.
Tomatoes are among the world’s most economically significant vegetable crops, but they are also notoriously vulnerable to disease. The Food and Agriculture Organization estimates that plant diseases destroy between 20 and 40 percent of global crop yields every year, and tomato pathogens—from early and late blight to bacterial spot, leaf mold, and tomato yellow leaf curl virus—spread rapidly and produce symptoms that are maddeningly similar to one another and to pest damage. Traditional scouting by trained personnel is labor-intensive and poorly suited to large-scale monitoring, which has driven a decade-long push toward computer vision systems that can diagnose diseases automatically from photographs.
Deep learning models have made impressive progress on this problem, but the field has been haunted by three persistent failures. Field photographs contain cluttered backgrounds—soil, hands, tools, other plants—that lure models into focusing on the wrong visual cues. Many disease categories look almost identical in their early stages, producing high intra-class similarity that confounds fine-grained classification. And agricultural datasets are severely imbalanced: common diseases have tens of thousands of images while rare conditions have only a handful, causing models to systematically underperform on the very classes where a missed diagnosis costs the most. Interpretability has also lagged behind accuracy, with most explainability techniques bolted on after training rather than woven into the model itself.
DAPR-AM-Net attacks all of these problems simultaneously through four tightly coupled innovations. The first is a Dual Attention Fusion Mechanism, or DAFM, built on top of an EfficientNet-B0 backbone. DAFM chains together two complementary attention architectures: a squeeze-and-excitation block that recalibrates the importance of individual feature channels, followed by a convolutional block attention module that sharpens focus both on which channels matter and where in the image the informative regions lie. In practice, this means the network learns to amplify the texture, color, and structural signatures of lesions—concentric rings of late blight, browned margins, chlorotic halos—while actively suppressing soil and background noise.
The second innovation addresses the augmentation problem. Standard MixUp training blends two training images into a synthetic hybrid with a randomly drawn mixing coefficient, which improves generalization but can blur precisely the lesion semantics the model needs to learn. The researchers’ Adaptive MixUp with Attention-Aware Sampling, or AMAAS, replaces the blind random blend with a guided one. The system consults attention maps from the previous training epoch, computes an importance score for each image reflecting how salient its diseased regions are, and rescales the mixing coefficient accordingly. Images with clear, informative lesions receive greater weight in the blend, while background-dominated or noisy samples are down-weighted. AMAAS additionally folds in a class-frequency compensation factor that boosts the representation of rare categories inside synthetic training samples, so that long-tailed diseases are not merely seen more often but seen more informatively.
The third component, Progressive Feature Refinement with Dual Attention (PFR-DA), reframes feature extraction as a multi-stage refinement rather than a single pass. Features from different network depths interact through gated cross-level fusion: high-level semantic information is progressively injected downward into low-level texture representations, and lightweight auxiliary classification heads attached to intermediate layers impose supervision at every stage. This design preserves fine-grained detail in early layers while ensuring the network’s final judgments remain consistent with its earlier visual evidence. The fourth element, an Imbalance-Aware Multi-Objective Optimization strategy called IAMOO, tackles class skew at three levels at once—through a weighted random sampler that oversamples rare classes, through the class-aware mixing already embedded in AMAAS, and through a composite training objective that balances overall accuracy against minority-class recall when selecting the final model.
The team evaluated the system on two datasets spanning opposite ends of the realism spectrum. The first is Plant-Village, a widely used public benchmark of 54,305 images across 38 disease classes and 14 crop species, captured under controlled conditions with plain backgrounds. The second is Tomato-DD, a self-constructed dataset of 48,584 images covering 11 tomato disease categories, compiled from independently collected field photographs and public sources and spanning multiple lighting conditions, occlusion levels, and background complexities. The dataset was rigorously re-annotated and deliberately includes ambiguous lesion boundaries, co-occurring symptoms, and environmental interference—the messy realities of an actual field.
The results were striking. On the Tomato-DD test set, DAPR-AM-Net achieved 99.73 percent accuracy, 99.73 percent precision, 99.74 percent recall, and a 99.73 percent F1-score, outperforming a battery of strengthened baselines including DenseNet-169, EfficientNet-B3, VGG-19, Xception, and a custom CNN, which reached 97.04, 99.16, 97.19, 96.19, and 85.98 percent respectively. On the full Plant-Village dataset, the model reached 99.85 percent accuracy with a 99.81 percent F1-score. Remarkably, it does so with a compact architecture of only 4.72 million parameters—VGG-19, by comparison, carries roughly 140 million—and sustains an end-to-end inference speed of about 302 frames per second on a standard GPU, including preprocessing and post-processing. That combination of speed and small footprint makes the model a realistic candidate for drones, mobile devices, and resource-constrained agricultural IoT nodes.
Ablation studies confirmed that every module earns its place. A baseline EfficientNet model scored 98.87 percent accuracy; adding each component in isolation pushed results upward, and the full four-module configuration delivered the best performance. The gains were most pronounced precisely where the design predicted they would be. For Powdery Mildew, a class with few training samples, the model achieved an F1-score of 99.73 percent, a 2.43-point improvement over the baseline; for Spider Mites it reached a perfect 100 percent. For the notoriously confusable pair of early blight and late blight, it posted F1-scores of 99.35 and 99.71 percent. Across all classes, the gap between precision and recall stayed below 0.65 percentage points, indicating consistent performance rather than strength concentrated in a few easy categories. Fivefold cross-validation yielded 99.64 percent on every metric with a standard deviation of just 0.12, signaling strong generalization.
Crucially, the system does not just answer—it shows its work. Grad-CAM visualizations integrated with the model’s own attention maps generate heatmaps that highlight the exact regions driving each diagnosis. For late blight, the channel attention first amplifies chromatic and textural descriptors of the lesion, and the spatial attention then locks onto browned leaf margins and concentric ring patterns, suppressing healthy green tissue. The model also proved sensitive to early-stage symptoms, detecting small initial lesions of leaf mold that could easily escape visual notice. Quantitative tests of explanation quality bore this out: the mean Insertion AUC of the saliency maps was approximately 0.85, and cosine robustness under Gaussian noise reached 0.985, meaning the explanations remain essentially stable even when inputs are perturbed. The model itself was similarly resilient, retaining 99.61 percent accuracy under motion blur and 99.62 percent under brightness changes, with only a modest dip to 99.27 percent under Gaussian noise.
To close the loop between laboratory and field, the researchers built a complete smart agriculture platform around DAPR-AM-Net. The three-tier web system lets growers capture or upload leaf images and receive, within seconds, the disease prediction, a confidence score, a Grad-CAM heatmap, and tailored pesticide recommendations. It links diagnosis to practical action: a spraying module fuses disease outputs with real-time meteorological data from Open-Meteo to compute a spray suitability index for the coming week, triggering recommendations when conditions are favorable and drift warnings when they are not. Additional modules provide three-day irrigation forecasts driven by weather and soil-moisture simulation, and early warnings for heavy rainfall, drought, and high winds. Performance testing on 40 CPU-only laptops showed an average latency of 0.87 seconds per image including visualization, and on mobile phones 1.12 seconds, with over 90 percent of requests completed within two seconds—fast enough for genuine field use. When confidence falls below 75 percent, the system prompts users for additional multi-angle images rather than guessing, and the researchers explicitly position the platform as a decision-support tool that keeps humans in charge of final calls, particularly in uncertain or high-risk scenarios.
The authors acknowledge limitations. RGB imagery constrains performance under extreme illumination or for spectrally subtle pathologies, and overconfident errors remain possible for extremely rare or unseen diseases—hence the emphasis on confidence thresholds and human review. They also note that the study has not yet included quantitative field trials measuring agronomic outcomes such as pesticide reduction or disease incidence, which they flag as a priority for future work, alongside multispectral data integration, pixel-level lesion segmentation, and active learning to adapt the model to regional and temporal shifts. Still, by unifying attention modeling, attention-guided augmentation, progressive refinement, and imbalance-aware optimization into a single deployable system, DAPR-AM-Net offers a blueprint for agricultural AI that is accurate, fast, and—perhaps most importantly—willing to explain itself.
Cite Scienmag News
Alan Morgan. (September 8, 2026). DAPR-AM-Net: explainable AI detects and forecasts tomato leaf diseases. Scienmag. https://scienmag.com/dapr-am-net-explainable-ai-detects-and-forecasts-tomato-leaf-diseases/
Alan Morgan. "DAPR-AM-Net: explainable AI detects and forecasts tomato leaf diseases." Scienmag, 8 September 2026, https://scienmag.com/dapr-am-net-explainable-ai-detects-and-forecasts-tomato-leaf-diseases/. Accessed 8 September 2026.
Alan Morgan. "DAPR-AM-Net: explainable AI detects and forecasts tomato leaf diseases." Scienmag. September 8, 2026. https://scienmag.com/dapr-am-net-explainable-ai-detects-and-forecasts-tomato-leaf-diseases/








