Grapevines are among the most economically valuable crops on Earth, and they are under constant assault from foliar diseases such as Black Rot, ESCA, and Leaf Blight. For generations, detecting these threats has depended on the trained eye of agricultural experts, a process that is slow, expensive, and vulnerable to human subjectivity. Now, a pair of researchers from the Data Science Laboratory at the Industrial University of Ho Chi Minh City in Vietnam has unveiled an artificial intelligence system that not only diagnoses grapevine diseases with remarkable accuracy but also explains its reasoning in plain language. The study, published in Discover Artificial Intelligence, describes a Global–Local Hybrid Framework built on a vision-language model that mimics how a human pathologist actually works: first surveying the whole leaf, then zooming in on the lesion with a magnifying glass.
The team, Lang Hoang Son and Bui Thanh Hung, chose to move beyond the convolutional neural networks (CNNs) that have dominated plant disease classification for years. CNNs such as ResNet and EfficientNet can achieve near-perfect scores on curated laboratory datasets, but they suffer from two fundamental weaknesses. They extract raw pixel-level features without high-level semantic reasoning, and they operate as inscrutable black boxes, offering farmers a class label with no justification. Worse, researchers have documented the so-called Clever Hans phenomenon, in which CNNs quietly learn to recognize background artifacts—soil, lighting, uniform laboratory backdrops—rather than the biological morphology of the disease itself. When such models confront real-world field images, their performance can collapse.
Vision-language models (VLMs) such as LLaVA, which combine a CLIP image encoder with a large language model backbone, promised a way out by generating human-readable diagnostic rationales. But they brought their own problem: the attention sink phenomenon. When fed complex real-world photographs, these models squander computational attention on noisy background elements—rocks, shadows, variable illumination—instead of focusing on the microscopic necrotic lesions that actually matter. The Vietnamese team’s solution was architectural: give the model two views of every leaf at once, one global and one local, so that attention is anchored to genuine pathology.
The local view is produced by an unsupervised Saliency Auto-Crop algorithm that requires no bounding box annotations, a major practical advantage because manual annotation is prohibitively expensive at agricultural scale. The algorithm first converts the input image into the CIE-LAB color space, decoupling luminance from chrominance. This matters because the a* channel is especially sensitive to the green-red spectrum, allowing necrotic brown lesions to stand out against green leaf tissue regardless of lighting conditions. A weighted saliency score is computed across the image, then Otsu’s adaptive thresholding and mathematical morphological operations filter out background noise. Finally, image moments are used to calculate the centroid of the isolated lesion, and a 128-by-128-pixel patch is cropped around it. The researchers empirically tested the patch size: crops of 64 pixels risk truncating diagnostic transition zones such as chlorotic halos, while 224-pixel crops capture so much healthy tissue that the patch becomes redundant with the global view.
These two images—a 256-pixel global view and the 128-pixel lesion patch—are encoded by a pre-trained CLIP ViT-L/14 encoder into 576 visual tokens each, projected into the language model’s embedding space, and concatenated with a carefully engineered text prompt that explicitly tells the model which image is which. The prompt instructs the system to identify exactly one of four categories: Black Rot, ESCA, Healthy, or Leaf Blight. Fine-tuning the 7-billion-parameter LLaVA-1.5 model conventionally would demand enormous computational resources, so the team applied 4-bit QLoRA, a quantized low-rank adaptation technique. The pre-trained weights are frozen and quantized into a 4-bit NormalFloat format, and lightweight trainable adapter matrices are injected into the decoder layers. Only 7.68 percent of the total parameters were updated, allowing the entire training run to complete on free cloud GPUs—an NVIDIA Tesla T4 pair with 32 gigabytes of memory—while the model acquired specialized plant pathology knowledge.
The results on the PlantVillage grapevine dataset, comprising 9,027 images split into training, validation, and blind test sets, were striking. The hybrid framework achieved an overall accuracy of 98.7 percent on the blind test, with Leaf Blight recognized perfectly at 100 percent, Healthy leaves at 99.8 percent, ESCA at 98.1 percent, and Black Rot at 97.0 percent—the last slightly hampered by the visual resemblance between early black rot lesions and leaf blight. An ablation study confirmed that both image streams are indispensable: the global image alone reached 97.8 percent, the local patch alone just 90.3 percent, and the un-tuned zero-shot model a dismal 18.6 percent. The QLoRA rank mattered too, with accuracy climbing from 62.1 percent at rank 8 to 98.0 percent at rank 32 before the team settled on rank 64 for maximum stability.
But the most consequential experiment was the one designed to expose the Clever Hans effect. The team evaluated all models in a strictly zero-shot manner on PlantDoc, a dataset of grapevine images captured in uncontrolled field conditions with complex backgrounds and volatile lighting. Advanced CNNs that had scored near 100 percent in the laboratory plummeted: EfficientNet-B4 fell from 99.8 percent to 54.9 percent, a catastrophic domain gap of 44.9 percentage points. The hybrid VLM framework, by contrast, retained 80.5 percent accuracy overall, with healthy leaf recognition peaking at 98.6 percent, though Black Rot degraded to 60.9 percent. The researchers attribute this resilience to the saliency-cropped local patches acting as semantic anchors that prevent the model from being derailed by environmental noise. The framework also scored 98.55 percent on the external AgroMind benchmark and a perfect 100 percent on AgroCoT, both laboratory-style datasets with subtle domain shifts.
Interpretability was engineered into the system rather than bolted on afterward. Instead of relying on post-hoc explanation tools like Grad-CAM, which merely visualize local gradients, the team extracted cross-attention matrices directly from the 31st and final decoder layer of the language model—the layer of highest semantic abstraction, immediately preceding the output head. A forward hook captures the mean attention weights that text tokens assign to visual tokens, which are then upsampled by bilinear interpolation and overlaid on the original image as heatmaps. Layer-wise analysis showed that shallow layers produce sparse, texture-focused activations, while deeper layers concentrate attention on pathological regions. Quantitatively, the model allocated roughly 73 percent of its attention to the global image and 27 percent to the local patch, a ratio that held steady whether the leaf was diseased or healthy—a Welch’s t-test found no significant difference between the two conditions, with a negligible effect size—indicating a stable, balanced diagnostic strategy rather than arbitrary attention spikes.
The authors are candid about the limitations. Inference is slow: the fine-tuned LLaVA requires 2,231 milliseconds per image, or 0.448 frames per second, and consumes 149 joules per image on a Tesla T4, making real-time deployment on drones or mobile devices impractical for now. Black Rot’s scattered, multi-spot lesions sometimes defeat the single-patch cropping strategy, motivating future multi-scale, multi-crop attention mechanisms. And the residual 17-point domain gap on field imagery shows that laboratory overfitting has not been fully eliminated. The team’s roadmap includes knowledge distillation to compress the model’s reasoning into lighter architectures, deeper quantization and pruning for on-device diagnostics, and evolving the static CIE-LAB extraction module into an autonomous visual agent that can propose and refine its own lesion bounding boxes through iterative reasoning.
Even with those caveats, the study marks a meaningful shift in agricultural AI. Where previous systems chased marginal accuracy gains on sterile benchmarks, this framework trades a sliver of in-domain performance for transparency, natural-language reasoning, and genuine cross-domain robustness. For vineyard managers, the difference is tangible: instead of an opaque percentage, they receive a diagnosis, a rationale, and a heatmap showing exactly which lesion triggered the decision. As vision-language models continue to mature, the grapevine may become the proving ground for a new generation of AI systems that agriculture can actually trust.
Subject of Research: Interpretable vision-language AI for grapevine disease diagnosis
Article Title: A global-local hybrid vision-language framework for interpretable grapevine disease diagnosis
Article References: A global-local hybrid vision-language framework for interpretable grapevine disease diagnosis. (n.d.). https://doi.org/10.1007/s44163-026-02463-x
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02463-x
Keywords: vision-language models, grapevine disease, plant pathology, explainable AI, LLaVA, QLoRA, saliency auto-crop, precision agriculture, convolutional neural networks, domain shift, cross-attention, PlantVillage
Cite Scienmag News
Alan Morgan. (October 8, 2026). AI Learns to Diagnose Grapevine Diseases Like a Plant Pathologist. Scienmag. https://scienmag.com/ai-learns-to-diagnose-grapevine-diseases-like-a-plant-pathologist/
Alan Morgan. "AI Learns to Diagnose Grapevine Diseases Like a Plant Pathologist." Scienmag, 8 October 2026, https://scienmag.com/ai-learns-to-diagnose-grapevine-diseases-like-a-plant-pathologist/. Accessed 8 October 2026.
Alan Morgan. "AI Learns to Diagnose Grapevine Diseases Like a Plant Pathologist." Scienmag. October 8, 2026. https://scienmag.com/ai-learns-to-diagnose-grapevine-diseases-like-a-plant-pathologist/

