<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>challenges in autonomous harvesting &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/challenges-in-autonomous-harvesting/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 14:50:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>challenges in autonomous harvesting &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Enhanced AI Vision System Teaches Greenhouse Robots to Actually Pick the Fruit</title>
		<link>https://scienmag.com/enhanced-ai-vision-system-teaches-greenhouse-robots-to-actually-pick-the-fruit/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 14:50:34 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[addressing labor shortages with robots]]></category>
		<category><![CDATA[agricultural robotics]]></category>
		<category><![CDATA[AI vision systems in greenhouses]]></category>
		<category><![CDATA[AI-powered agricultural robots]]></category>
		<category><![CDATA[challenges in autonomous harvesting]]></category>
		<category><![CDATA[collaborative robots]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[depth validation]]></category>
		<category><![CDATA[E-YOLOv11-RGBD framework]]></category>
		<category><![CDATA[fruit detection]]></category>
		<category><![CDATA[fruit detection in complex environments]]></category>
		<category><![CDATA[fruit picking automation]]></category>
		<category><![CDATA[greenhouse automation]]></category>
		<category><![CDATA[greenhouse robot fruit picking]]></category>
		<category><![CDATA[improving harvesting success rates]]></category>
		<category><![CDATA[machine learning in agriculture]]></category>
		<category><![CDATA[perception-to-action]]></category>
		<category><![CDATA[RGB-D sensing]]></category>
		<category><![CDATA[robotic arm manipulation for fruit harvesting]]></category>
		<category><![CDATA[robotic harvesting]]></category>
		<category><![CDATA[UR3e manipulator]]></category>
		<category><![CDATA[vision system for robotic harvesting]]></category>
		<category><![CDATA[YOLOv11]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=262446</guid>

					<description><![CDATA[A new RGB-D perception-to-action framework called E-YOLOv11-RGBD raised robotic greenhouse harvesting success from 51.7 to 92.5 percent by converting AI fruit detections into depth-validated, reachable and graspable robot actions.]]></description>
										<content:encoded><![CDATA[<p>Agricultural robots have long been able to see fruit with impressive accuracy, yet many still struggle to actually pick it. A new study published in Smart Agricultural Technology tackles precisely that gap, presenting a framework called E-YOLOv11-RGBD that links what a camera sees to what a robotic arm can physically do. Developed by Filipe Pereira, Paulo Vieira, Luís Freitas, José M. Machado, António Ramos Silva and António M. Lopes at the University of Minho&#8217;s Mechatronics Laboratory, the system was tested on tomato, strawberry and pepper harvesting in a controlled greenhouse-like setting, and it more than doubled the overall harvesting success rate compared with a conventional baseline.</p>
<p>The motivation is straightforward. Labour shortages are squeezing agricultural sectors worldwide, and harvesting remains one of the most physically demanding and repetitive farm tasks. Robotic harvesting has attracted enormous attention, but reliable selective harvesting outside demonstrations remains elusive. Fruits hide behind leaves, illumination shifts through the day, and the same crop appears in different sizes, colours and orientations. Greenhouses reduce some of this chaos but introduce their own constraints: restricted spaces, metallic structures, dense vegetation and, in some cases, the need to coexist safely with human workers.</p>
<p>The researchers identified a recurring blind spot in the literature. Most studies report image-level metrics such as precision, recall and mean average precision, but far fewer analyse what happens after detection. A high-confidence bounding box does not guarantee reliable depth information. A geometrically valid target may still lie outside the manipulator&#8217;s workspace or exceed the gripper&#8217;s aperture. Because perception studies often evaluate these stages independently, it has been difficult to know how detection errors propagate into physical harvesting failures. The team&#8217;s central insight is that the real problem is not just detecting crops, but converting a visual detection into a robot-executable action while explicitly accounting for sensing, kinematic and grasping constraints.</p>
<p>The hardware platform combines commercially available components: an Intel RealSense D435i RGB-D camera for perception, a UR3e collaborative manipulator for harvesting motion, an OnRobot 2FG7 electric gripper for grasping, and an Omron LD-60 mobile robot as an integration layer for future greenhouse mobility. The manipulator was mounted on the mobile base using a modular aluminium structure, verified through finite-element analysis that showed a minimum safety factor of 15. Importantly, the harvesting experiments reported here were performed with the stationary manipulation subsystem; autonomous navigation was not evaluated.</p>
<p>At the heart of the system is an enhanced version of the YOLOv11-M object detector, adapted through what the authors call a failure-driven approach. Rather than claiming novelty for any single component, they combined mechanisms chosen to fix specific observed errors. A new P2 small-target detection branch preserves fine-scale information for small or distant fruit, particularly strawberries. Bidirectional P2–P5 feature fusion strengthens multiscale information exchange. A C2PSA-iEMA attention module sharpens discrimination between crops and visually similar backgrounds, while Wise-PIoU bounding-box regression improves localization quality under variable scale, partial occlusion and irregular fruit geometry. Hard-negative learning exposed the detector to leaves, stems, shadows and support wires that had triggered false positives, and scale-aware training increased exposure to fruits at different apparent sizes.</p>
<p>The results at detector level were clear. On a fixed test set of 600 labelled crop instances drawn from a merged 3,233-image multicrop dataset, E-YOLOv11 raised precision from 95.0 to 97.4 percent, recall from 93.0 to 95.8 percent, and mAP at the strict 0.5–0.95 threshold from 90.0 to 93.1 percent, at a modest inference cost of 18 milliseconds per image. The most striking error-mode change involved background confusion: the baseline model sometimes mistook background regions for tomatoes, which in a real robot would trigger pointless and time-wasting movements. The enhanced configuration substantially reduced these false positives.</p>
<p>But the true contribution lies in the perception-to-action pipeline that follows detection. Every accepted detection passes through a sequence of validation gates. The depth frame is aligned to the RGB image, and depth is sampled as a local average around the bounding-box centre to suppress sensor noise. Detections with missing or inconsistent depth are rejected outright. Valid candidates are projected into three-dimensional camera coordinates using the camera&#8217;s intrinsic parameters, and an approximate crop width is estimated from the bounding box and measured depth. Coordinates are then transformed into the robot&#8217;s reference frame, corrected for the tool centre point offset, and checked for inverse-kinematics feasibility on the UR3e controller, which tests 121 candidate end-effector orientations. Finally, graspability is verified against the gripper&#8217;s 73-millimetre maximum aperture. Only candidates passing every gate generate a robot command, transmitted through the real-time RTDE interface.</p>
<p>The physical validation was rigorous in its accounting. RGB-D localization was assessed independently at 42 known workspace positions with three repetitions each, yielding a mean three-dimensional Euclidean error of 10.20 ± 3.74 millimetres. The harvesting comparison involved 120 attempts per configuration, 40 per crop class, for 240 physical attempts overall. The baseline YOLOv11-M configuration succeeded in 62 of 120 attempts, a 51.7 percent success rate. The enhanced E-YOLOv11-RGBD framework achieved 111 of 120, or 92.5 percent, a difference of 40.8 percentage points. The most dramatic improvement came with strawberries, which jumped from zero successful harvests out of 40 in the baseline to 35 out of 40, largely because the enhanced detector resolved distance-related ambiguity in which strawberries were misclassified as tomatoes. Pepper harvesting rose from 55 to 90 percent, while tomato remained perfect in both configurations. Mean cycle time also fell slightly, from 18.35 to 17.20 seconds, because fewer invalid targets wasted robot motion.</p>
<p>The authors are notably candid about the limits of their claims. The detector comparison rests on a single training run per configuration, so no between-seed variability is available, and the merged public dataset used a fixed image-level split whose provenance-disjointness was not formally established. The harvesting-success difference reflects the integrated configuration, detector plus validation refinements, and cannot be attributed to the detector alone because no factorial experiment separated them. The trials were conducted under stable artificial illumination at roughly 45 centimetres camera-to-scene distance, and transfer to commercial greenhouses with changing light, plant motion and dense foliage has not been demonstrated. The OnRobot 2FG7 gripper also proved restrictive for larger or irregularly shaped peppers, and the UR3e workspace bounds the reachable harvesting region.</p>
<p>Even with those caveats, the study offers a compelling proof of concept for a shift in how agricultural robotics should be evaluated. The authors argue that depth reliability, metric localization, kinematic feasibility and end-effector compatibility must be assessed together with image-level detection, since perception accuracy alone is a poor predictor of harvesting success. Future work will pursue larger greenhouse trials, formal hand-eye calibration, adaptive crop-specific grasping, feature-level fusion of RGB and depth data inside the detector, deployment on low-power embedded hardware with quantized models, and coordinated navigation between the mobile base and the manipulator. If those steps succeed, the gap between seeing fruit and picking it may finally close, bringing robotic harvesters closer to the greenhouses where labour is scarcest.</p>
<p><strong>Subject of Research:</strong> An RGB-D perception-to-action framework for multicrop greenhouse robotic harvesting using an enhanced YOLOv11 detector</p>
<p><strong>Article Title:</strong> E-YOLOv11-RGBD perception-to-action framework for multicrop greenhouse robotic harvesting</p>
<p><strong>Article References:</strong> Pereira, F., Vieira, P., Freitas, L., Machado, J. M., Silva, A. R., &amp; Lopes, A. M. (2026). E-YOLOv11-RGBD perception-to-action framework for multicrop greenhouse robotic harvesting. <em>Smart Agricultural Technology, 15</em>, Article 102609. <a href="https://doi.org/10.1016/j.atech.2026.102609" rel="noopener noreferrer">https://doi.org/10.1016/j.atech.2026.102609</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> robotic harvesting, greenhouse automation, YOLOv11, RGB-D sensing, computer vision, collaborative robots, agricultural robotics, deep learning, fruit detection, UR3e manipulator, depth validation, perception-to-action</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">262446</post-id>	</item>
	</channel>
</rss>
