Sichuan pepper, the spice behind the numbing tingle of Chinese cuisine, has long resisted mechanization. Its clusters grow in random orientations on thorny branches, its skin is studded with fragile oil glands that release the prized aroma only when intact, and picking it still consumes roughly one-third of total production costs. Now a research team reporting in Artificial Intelligence in Agriculture has built a picking robot that combines multi-task computer vision, principal component analysis-based pose estimation, and a flexible swallowing end-effector to harvest the delicate spice without rupturing its oil glands.
The problem with existing tools is fundamentally mechanical. Hand-held shear-type pickers are accurate but slow and prone to slicing neighboring leaf buds, while electric comb-type and saw-blade devices rely on rigid, fast-moving components that smash into dense clusters. Those collisions squeeze and rupture the tiny oil glands, which are between 0.5 and 1 millimeter in diameter, causing aroma volatilization and quality degradation. Rigid mechanisms elsewhere in agriculture tell a similar story: canopy-contact vibration harvesters for plums operate nearly 40 times faster than manual labor but damage the skin of 10 to 18 percent of fruit. For a small, thorny, densely clustered berry like Sichuan pepper, the team concluded that neither rigid gripping nor vibration would work.
Their solution begins with perception. The researchers collected 941 images of Sichuan pepper plants in Hanyuan County, Sichuan Province, a region known as the Hometown of Sichuan Pepper, capturing them with a smartphone at 3024 by 4032 resolution and annotating them for both object detection and semantic segmentation. They then trained YOLOP2, a multi-task network in which detection and segmentation share a single E-ELAN encoder, a lightweight feature-fusion neck with spatial pyramid pooling, and decoupled decoder heads for each task. On a test set, the model achieved a mean average precision at an IoU threshold of 0.5 of 0.9101 and a mean intersection-over-union of 0.8883, outperforming YOLOv8x, Faster R-CNN, DeepLabv3+, and U-Net, while running joint inference at roughly 34 frames per second on a single GPU, fast enough for real-time field operation.
A striking finding emerged from how the data were labeled. The team split their detection dataset into large-cluster and small-cluster annotations, using an objective criterion of 40 millimeters maximum cluster diameter measured with vernier calipers in the field. Large-cluster annotation yielded an mAP of 0.9101, while small-cluster annotation dropped to 0.6972, a statistically significant gap of 0.2129 confirmed by t-test. The reason lies in the non-maximum suppression stage: because pepper clusters grow densely and overlap, fine-grained small-cluster labels generate many highly overlapping predicted boxes that suppress one another, producing missed detections. Segmentation performance, by contrast, was essentially unaffected by annotation granularity, suggesting that detection and segmentation respond differently to how humans carve up continuous fruit masses into discrete targets.
With masks in hand, the pipeline converts 2D perception into 3D geometry. Segmentation masks for clusters and branches constrain which pixels are back-projected through the pinhole camera model into 3D point clouds, separating foreground from background at the data-generation stage. A depth main-peak filtering method then builds a histogram of point depths, identifies the most frequent depth as the cluster’s primary spatial position, and expands an adaptive window around that peak until the growth rate of retained points falls below a 5 percent threshold. Ablation tests showed this adaptive window achieved a 72.30 percent point retention rate, sitting sensibly between fixed windows that were too narrow, which truncated valid cluster edges, and too wide, which admitted background noise.
The heart of the pose estimation is principal component analysis applied to the cleaned point cloud. The cluster’s point coordinates are zero-centered, a covariance matrix is computed, and eigenvalue decomposition yields three orthogonal eigenvectors corresponding to the major, middle, and minor axes of a fitted ellipsoid, along with the geometric center that anchors the model. Error statistics across 189 tests showed the middle axis was fitted most reliably, with a mean absolute error of 7.9 millimeters and a bias of just 1.1 millimeters, while the minor axis was the dominant error source, underestimated with a mean absolute percentage error of 35.67 percent. Importantly, this error pattern held stable across left, parallel, and right camera viewing angles, meaning the systematic bias is predictable rather than random, a property the team exploits rather than fights.
Because knowing a cluster’s orientation alone cannot prevent a robotic arm from colliding with branches, the researchers also modeled the branch itself. They extracted a skeleton of center points from the branch point cloud and fitted a cubic B-spline curve, then found the point on that curve closest to each cluster center. From the branch tangent and the radial vector from branch to cluster, they constructed a local orthogonal coordinate frame and classified each cluster into one of five poses: upward, downward, leftward, rightward, or forward relative to its branch. Geometric analysis showed that feeding the end-effector along the radial direction, with its opening plane perpendicular to the pedicel, lets the tool align precisely with the pedicel base while avoiding both lateral interference and frontal collisions with the cluster body.
The feed depth is likewise adaptive. Rather than relying on a fixed 2D projection center, the system selects the ellipsoid principal axis most closely aligned with the feed direction and uses that axis’s semi-axis length as the insertion depth. In a 20-sample comparison, the fixed-depth method produced a mean absolute error of 5.80 millimeters against the true shearing-in distance, with 45 percent of samples suffering either insufficient feed, which left the blade short of the pedicel, or excessive feed, which collided with branches. The adaptive method cut the error to 2.40 millimeters. Meanwhile, the end-effector itself pairs a double-cycloid shearing mechanism with a flexible TPU spiral conveying drum driven by a single 7.8-watt brushless motor through a double-ratchet transmission that decouples shearing from conveying. Force calculations confirmed the blade delivers about 14.53 newtons, roughly 160 percent above the 5.58-newton maximum required to cut the toughest pedicels, and simulation identified 83 revolutions per minute as the optimal drum speed, where peak skin force of 50.3 newtons stays just under the 51.4-newton rupture threshold measured with a texture analyzer.
Field trials conducted from September 25 to 29, 2025, under illumination ranging from 12,600 to 76,200 lux, put the whole system to the test on a JAKA Zu5 six-degree-of-freedom collaborative arm. A complete picking cycle, from recognition to collection, took 8.26 seconds, with the vision step consuming a mere 0.06 seconds. Across 30 trials the robot achieved a mean picking net rate of 50.2 percent and a picking success rate of 56.7 percent, with successful trials averaging 81.4 percent net rate against 8.6 percent for failures, a pronounced all-or-nothing pattern in which correct feed positioning nearly guarantees a clean harvest. Analysis of the 13 failures attributed 46.1 percent to overlapping branches and 30.7 percent to thicker pedicels, with leaf obstruction and scattered cluster growth accounting for the rest. Critically, an acid-value test-strip colorimetric assay showed no color change in peppers picked by the robot, while controls with artificially ruptured oil glands turned clearly yellow, demonstrating that the swallowing mechanism genuinely avoids mechanical oil gland damage.
The team is candid about limitations. Wind-induced branch sway causes point cloud loss, single-view observation leaves spatial blind zones, and the current shearing mechanism struggles with thick pedicels and dense obstructions. Future work will fuse multi-frame temporal point clouds using iterative closest point registration, add active viewpoint planning and artificial potential field obstacle avoidance, and install a rigid-flexible guiding hood to passively push aside interfering foliage. Even so, the study marks a meaningful advance in an underexplored field: unlike single-fruit robots for tomatoes or apples, Sichuan pepper demands pose estimation and damage-free handling of thorny, interlocking clusters. By pairing a perception system that quantifies its own biases with hardware that tolerates them, the researchers have sketched a credible path toward automating one of the world’s most labor-intensive spice harvests.
Subject of Research: Robotic multi-pose harvesting of Sichuan pepper using multi-task visual perception and a flexible swallowing end-effector
Article Title: Research and trials on multi-pose picking of Sichuan pepper based on multi-task perception
Article References: Chen, C., Wang, Z., Cheng, T., Song, Z., Lu, J., Yang, F., & Wang, Z. (2026). Research and trials on multi-pose picking of Sichuan pepper based on multi-task perception. Artificial Intelligence in Agriculture. https://doi.org/10.1016/j.aiia.2026.09.005
Image Credits: AI Generated
DOI: 10.1016/j.aiia.2026.09.005
Keywords: Sichuan pepper, agricultural robotics, pose estimation, multi-task learning, YOLOP2, flexible end-effector, point cloud processing, PCA, precision agriculture, harvest automation, machine vision, damage-free picking
Cite Scienmag News
Denise Maddox. (October 2, 2026). Robot Learns to Read Pepper Clusters and Swallow Them Whole. Scienmag. https://scienmag.com/robot-learns-to-read-pepper-clusters-and-swallow-them-whole/
Denise Maddox. "Robot Learns to Read Pepper Clusters and Swallow Them Whole." Scienmag, 2 October 2026, https://scienmag.com/robot-learns-to-read-pepper-clusters-and-swallow-them-whole/. Accessed 2 October 2026.
Denise Maddox. "Robot Learns to Read Pepper Clusters and Swallow Them Whole." Scienmag. October 2, 2026. https://scienmag.com/robot-learns-to-read-pepper-clusters-and-swallow-them-whole/

