Every autumn, across the rolling farmland of northern Europe, fleets of massive harvesters rumble through sugar beet fields, tearing the pale, tapering roots from the soil at a rate of dozens of tonnes per hour. It is a brutal, high-speed operation, and the crop pays a price. Beets get their tops sliced too deeply during defoliation, snap apart during lifting, or are bruised and scuffed by the rotating cleaning turbines that strip away clods of mud. Each wound bleeds sugar, and sugar is money. Now a team of researchers in Greece has built an artificial intelligence system that watches the crop flow inside the machine and flags damaged beets within milliseconds, a step toward harvesters that could adapt themselves on the fly to protect every root they pull from the ground.
The new framework, described in the journal Smart Agricultural Technology, tackles a problem that has long frustrated agricultural engineers. Sugar beet is one of Europe’s most valuable arable crops, with France and Germany leading production according to the United Nations Food and Agriculture Organization, yet a meaningful share of every harvest is degraded by mechanical damage between the soil and the truck. Because the damage happens inside a machine moving at full field speed, farmers have had no practical way to see it, measure it, or respond to it. The researchers, led by Konstantinos Gkountakos of the Centre for Research and Technology Hellas, set out to give harvesters something close to a nervous system: cameras, neural networks, and the ability to distinguish a healthy beet from a broken one in real time.
The journey from field to tank inside a modern harvester is short but violent. First, defoliators gently remove the leafy tops; cut too high and leaves remain to regrow, cut too low and the head of the beet, which is rich in sugar, is shaved away. Next, lifting shares pry the roots from the ground, ideally without cracking them. Then come the cleaning turbines, cylindrical drums that spin the beets against one another to knock off soil, stones, and debris. Residual soil inflates transport costs and accelerates erosion, but overly aggressive cleaning shatters and abrades the roots. Finally, a metallic sieve belt carries the cleaned beets into the tank. At the end of this conveyor, just before loading, the researchers saw an opportunity: a natural inspection point where every beet passes the camera, and where a damaged crop could still trigger adjustments to the lifting and cleaning mechanisms upstream.
To train and test their system, the team created something the field has lacked: a genuinely realistic dataset, which they named VerBeet. Using an industrial Basler camera with 8.3 megapixel resolution mounted in a white housing above the sieve belt of a Vervaet Q-621 harvester, they captured 694 frames over five harvesting days between early December 2025 and early January 2026, spaced 60 to 75 seconds apart to avoid overlapping content and to sample entire fields. Because daylight shifted with the harvester’s position relative to the sun, producing frames that ranged from dark to glare-bright, they installed three 40-watt LED lights on magnetic mounts, providing 13,200 lumens of consistent illumination. Two human annotators then spent roughly 17 hours drawing bounding boxes around every beet, cross-validating each other’s work, ultimately labeling 5,158 individual beets across three categories: undamaged, breakage, and abrasion.
The class distribution itself tells a story about mechanized harvesting. As expected, undamaged beets dominate with 3,685 instances, but 1,293 beets showed breakage, structural damage to the root from lifting or cleaning, and 180 displayed surface abrasion. The imbalance between the two damage types proved instructive: abrasions, the researchers found, are often hidden by clinging soil or simply invisible from a single camera angle, whereas a broken beet announces itself regardless of perspective. This matters because the ultimate goal is not just detection but attribution, knowing whether damage originated in the topping, lifting, or cleaning stages, so the harvester can adjust the right component rather than guessing.
The framework itself operates in two stages. A state-of-the-art object detector first localizes every beet in each frame; the researchers benchmarked lightweight members of the YOLO family, including YOLOv11n, YOLOv12n, and the newest YOLO26, against the transformer-based RT-DETR. Each cropped beet is then passed to a convolutional neural network classifier, drawn from EfficientNet, DenseNet, and MobileNet architectures, that decides whether the root is whole, broken, or abraded. Speed was non-negotiable: YOLOv11n processed more than 160 frames per second on the VerBeet test set, while MobileNetV2 classified beets at around 130 instances per second on modest consumer-grade hardware, confirming the pipeline could run on the edge devices that a real harvester would carry.
Two clever training tricks pushed performance further. The first is a size-aware weighted loss. Because the camera always views the belt from the same distance, a beet’s apparent size in the image reveals how much of it is visible; occluded or partially buried beets appear smaller and are harder to classify. The loss function weights each training sample by the inverse of its normalized area, so smaller, harder instances contribute more to learning, with weights clipped between 0.1 and 1.0 to keep training stable. The second trick is generative: the team used FLUX.2-4B Klein, a lightweight open-source diffusion model, to synthesize augmented training images in which each beet is placed on a randomized background. This counteracts overfitting to the unchanging metallic belt behind the beets and, crucially, multiplies scarce examples of the rare abrasion class. Every synthetic image was manually inspected, and hallucinated samples where the model invented or erased beets were discarded.
The results reveal both the promise and the stubborn difficulty of real-world machine vision. On the controlled, laboratory-style Semantic Sugar Beet dataset, models reached F1-scores above 87 percent, but on VerBeet’s dusty, occluded, unevenly lit frames, the best binary damage detection, EfficientNetB3 with size-aware loss and diffusion augmentation, achieved an F1-score of 71.48 percent. Cross-dataset experiments drove the point home: models trained on the controlled dataset collapsed when tested in the field, with recall as low as 12.7 percent, while models trained on VerBeet generalized far better, reaching 74.6 percent mAP on the cleaner data. In the most demanding test, a field-disjoint end-to-end evaluation where the test set always contained beets from a harvesting day entirely unseen during training, the full pipeline of YOLOv11n and MobileNetV2-Size achieved a mean F1-score of 67.09 percent across five folds, distinguishing beets, breakage, and abrasion in realistic conditions.
The researchers are candid about limitations. The dataset came from a single harvester model, so cameras mounted at different angles or heights on other machines could shift the statistics the networks rely on. Abrasion remains the weakest link, with multiclass F1-scores hovering near 51 percent, limited by the scarcity of genuine examples and the visual similarity between abraded and broken surfaces. Future work, the authors suggest, will extend VerBeet to more harvesting conditions, beet varieties, and even other root crops such as potatoes, and will explore adapting modern vision foundation models to withstand occlusion, dust, and erratic light under real-time constraints on semi-autonomous harvesting platforms.
Still, the trajectory is clear. Agriculture is racing toward machines that perceive, decide, and react, and the sugar beet harvester, long a blunt instrument of force and throughput, is acquiring finesse. A future machine that detects rising breakage rates and automatically softens its lifting shares, or notices abraded beets and eases the turbine speed, would translate directly into tonnes of sugar saved and soil left in the field where it belongs. The VerBeet dataset, with its 5,158 honestly annotated beets captured in the mud and glare of actual harvest days, may well become the benchmark on which that generation of gentle machines is trained.
Subject of Research: Real-time computer vision detection of sugar beet damage during mechanical harvesting
Article Title: Sugar beet localization and damage detection during harvesting
Article References: Gkountakos, K., Pasios, S., Ioannidis, K., Demestichas, K., Vrochidis, S., & Kompatsiaris, I. (2026). Sugar beet localization and damage detection during harvesting. Smart Agricultural Technology, 15, Article 102560. https://doi.org/10.1016/j.atech.2026.102560
Image Credits: AI Generated
DOI: 10.1016/j.atech.2026.102560
Keywords: sugar beet, computer vision, deep learning, YOLO, damage detection, precision agriculture, harvesting, VerBeet dataset, diffusion models, object detection, agricultural technology, neural networks
Cite Scienmag News
Alan Morgan. (September 23, 2026). AI Spots Bruised Sugar Beets in Real Time to Cut Harvest Losses. Scienmag. https://scienmag.com/ai-spots-bruised-sugar-beets-in-real-time-to-cut-harvest-losses/
Alan Morgan. "AI Spots Bruised Sugar Beets in Real Time to Cut Harvest Losses." Scienmag, 23 September 2026, https://scienmag.com/ai-spots-bruised-sugar-beets-in-real-time-to-cut-harvest-losses/. Accessed 23 September 2026.
Alan Morgan. "AI Spots Bruised Sugar Beets in Real Time to Cut Harvest Losses." Scienmag. September 23, 2026. https://scienmag.com/ai-spots-bruised-sugar-beets-in-real-time-to-cut-harvest-losses/

