<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>unstructured scenes &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/unstructured-scenes/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 08:06:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>unstructured scenes &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Smarter Robot Hands: New Vision System Grabs Objects With 95% Accuracy in Cluttered Scenes</title>
		<link>https://scienmag.com/smarter-robot-hands-new-vision-system-grabs-objects-with-95-accuracy-in-cluttered-scenes/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 08:06:00 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced robotic grasping technology]]></category>
		<category><![CDATA[cluttered object grasping]]></category>
		<category><![CDATA[collaborative robots]]></category>
		<category><![CDATA[collaborative robots in unstructured scenes]]></category>
		<category><![CDATA[GR-ConvNet]]></category>
		<category><![CDATA[grasp pose estimation]]></category>
		<category><![CDATA[Hefei University of Technology]]></category>
		<category><![CDATA[industrial automation robotics]]></category>
		<category><![CDATA[intelligent manufacturing]]></category>
		<category><![CDATA[lightweight segmentation model]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[MobileSAM]]></category>
		<category><![CDATA[multi-object manipulation robotics]]></category>
		<category><![CDATA[object detection]]></category>
		<category><![CDATA[object detection success rate]]></category>
		<category><![CDATA[RGB-D camera object detection]]></category>
		<category><![CDATA[RGB-D vision]]></category>
		<category><![CDATA[robot grasp success rate]]></category>
		<category><![CDATA[robot vision system]]></category>
		<category><![CDATA[robotic grasping]]></category>
		<category><![CDATA[unstructured scene object recognition]]></category>
		<category><![CDATA[unstructured scenes]]></category>
		<category><![CDATA[YOLOv8n]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252717</guid>

					<description><![CDATA[Researchers at Hefei University of Technology report a vision-based grasping pipeline combining improved object detection, MobileSAM segmentation, and grasp pose estimation that achieved a 98 percent detection rate and 95 percent grasp success rate for collaborative robots in cluttered, unstructured scenes.]]></description>
										<content:encoded><![CDATA[<p>Robots that can reach into a jumbled bin of parts and pick out exactly the right object remain one of the stubborn challenges of industrial automation. In a new preprint posted on 7 September 2026 in Mechanical Sciences Discussions, a team of engineers at Hefei University of Technology in China reports a vision-based grasping system that pushes collaborative robots closer to that goal. Led by Junxiao Liu and corresponding author Jun Qian, the researchers combined an improved object detector, a lightweight segmentation model, and a grasp-pose estimation network into a single pipeline that lets a robot find a target object in a cluttered, unstructured scene and compute a stable way to pick it up. On a physical robot platform using RGB-D camera input, the method achieved a 98 percent target detection success rate and an overall grasp success rate of 95 percent across 100 trials.</p>
<p>The core problem the team set out to solve is deceptively simple to describe and notoriously hard to engineer around. In structured factories, robots operate in carefully staged environments where objects arrive in known positions and orientations. In unstructured scenes, by contrast, objects can lie at any angle, overlap one another, and sit against backgrounds that confuse machine vision systems. When a robot is asked to grasp one specific object among many, unclear target regions and interference from backgrounds and neighboring objects can badly corrupt the selection of a grasp point. A grasp that looks geometrically sound in the image may actually land on the wrong object or on a patch of background, causing the gripper to close on empty air or knock over adjacent items.</p>
<p>The first stage of the pipeline tackles detection. The researchers started with YOLOv8n, a compact member of the widely used You Only Look Once family of real-time object detectors, and modified two of its internal components. They replaced the standard convolution modules with RFAConv, or receptive-field attention convolution, a design that gives the network a more flexible way to weight the spatial extent of the features it extracts. This change is aimed at sharpening the detection of target boundaries and local details, which matters enormously when objects touch or partially occlude one another. They also swapped the conventional upsampling module for DySample, a dynamic upsampling technique that adapts how low-resolution feature maps are enlarged back to image scale, helping the network recover fine spatial structure that standard upsampling tends to blur away.</p>
<p>Detecting the object, however, is only half the problem. The robot also needs to know precisely where and how to close its gripper. For this, the team turned to GR-ConvNet, a generative residual convolutional network designed to predict grasp poses from images. A grasp pose in planar grasping is typically described by the position of the grasp center, the angle of the gripper approach, and the width of the gripper opening. The researchers enhanced GR-ConvNet with a residual attention structure, related to the CBAM attention mechanism, which lets the network emphasize the most informative feature channels and spatial locations when producing its grasp quality map. They also added a region-guidance mechanism that steers the grasp prediction toward the relevant part of the scene, improving the stability of grasp predictions when backgrounds are visually complex.</p>
<p>The most distinctive element of the method is how it constrains the grasp search to the target object itself. After the improved YOLOv8n produces a bounding box around the target, that box is fed as a prompt to MobileSAM, a lightweight variant of the Segment Anything Model. MobileSAM converts the coarse bounding box into a precise pixel-level mask of the target object. This mask is then used to filter the grasp quality map output by the improved GR-ConvNet, restricting the search for grasp points to pixels that actually belong to the target. The effect is a coarse-to-fine strategy: detection narrows the field, segmentation refines it, and grasp estimation operates only within the verified target region. This dramatically reduces the chance that the robot will be distracted by non-target objects or background clutter when choosing where to grasp.</p>
<p>The entire system was validated on a collaborative robot visual grasping experimental platform, with a UR5 robot arm at its center, using RGB-D images as input. RGB-D cameras capture both color and depth information, allowing the system to translate image-plane grasp poses into three-dimensional robot motions. In the reported experiments, the complete pipeline achieved a 98 percent success rate in detecting the correct target and a 95 percent overall grasp success rate, with 95 successful grasps out of 100 trials. The authors report that the method reduces interference from non-target regions during grasp point selection and improves the stability of target-specific grasping in unstructured multi-object scenes, supporting its potential application in intelligent manufacturing settings where robots must handle diverse, unpredictably arranged parts.</p>
<p>The work is not without its critics, and the open peer discussion attached to the preprint illustrates how modern robotic research is scrutinized. An anonymous referee, commenting on 12 September 2026, acknowledged the practical engineering relevance of the study and the breadth of its experiments, but raised pointed questions about novelty and rigor. Because RFAConv, DySample, CBAM, and MobileSAM are all existing components, the referee argued that the contribution could be seen as a combination of established modules rather than a fundamentally new method, and asked the authors to clarify where the main novelty lies relative to other detection-assisted and segmentation-assisted grasping approaches. The referee also requested fuller experimental documentation, including dataset sizes, training splits, optimizer settings, learning rates, batch sizes, and hardware details needed for reproducibility.</p>
<p>The referee&#8217;s technical concerns go to the heart of how such systems should be evaluated. One issue concerns the parameters of the Gaussian region-guidance mechanism and the coarse-to-fine prediction branches in the improved GR-ConvNet, whose values and selection criteria were not fully reported, making it hard to judge whether performance is sensitive to their tuning. Another concerns the evidence for the segmentation step&#8217;s benefit: while the paper shows that the MobileSAM mask yields higher overlap with the true target region and less background redundancy than the raw bounding box, the referee noted that this does not directly prove improved grasping performance, and suggested a controlled comparison between bounding-box-constrained and mask-constrained grasping under identical conditions. The referee further recommended ablation experiments comparing the original YOLOv8n plus GR-ConvNet baseline against the improved components, along with explicit criteria for what counts as a successful grasp.</p>
<p>These debates matter because the stakes for visual grasping research are rising quickly. Collaborative robots, or cobots, are designed to work alongside humans without safety cages, and their economic promise depends on flexibility: the same arm that packs boxes today should sort irregular parts tomorrow without expensive reprogramming or fixture redesign. A vision system that reliably isolates a requested object from clutter and computes a stable grasp is a key enabling technology for that flexibility, with applications ranging from bin picking in warehouses to parts handling in small-batch manufacturing and even assistive robotics. The Hefei team&#8217;s coarse-to-fine architecture, in which detection prompts segmentation and segmentation constrains grasp estimation, reflects a broader trend of chaining foundation-model components like SAM with task-specific networks to get the best of both general visual knowledge and domain-specific precision.</p>
<p>As a preprint under review at Mechanical Sciences, the paper remains a work in progress, and the authors&#8217; responses to the referee&#8217;s requests for ablation studies, parameter disclosure, and baseline comparisons will determine how strong the final contribution proves to be. What is already clear is the shape of the engineering solution: rather than building one monolithic network to solve detection, segmentation, and grasping simultaneously, the team composed specialized modules, each improved for its particular weakness, and used the output of each stage to discipline the next. With a 95 percent grasp success rate in genuinely unstructured multi-object scenes, the method offers a compelling data point that this modular strategy can deliver reliable, target-specific manipulation, even as the peer-review process works to establish exactly which pieces of the pipeline deserve the credit.</p>
<p><strong>Subject of Research:</strong> Vision-based robotic grasping in unstructured multi-object scenes using object detection and grasp pose estimation</p>
<p><strong>Article Title:</strong> A visual grasping method for collaborative robots in unstructured scenes based on object detection and grasp pose estimation</p>
<p><strong>Article References:</strong> Liu, J., Qian, J., Tan, Y., &amp; Zhou, R. (2026). A visual grasping method for collaborative robots in unstructured scenes based on object detection and grasp pose estimation. <a href="https://doi.org/10.5194/ms-2026-165" rel="noopener noreferrer">https://doi.org/10.5194/ms-2026-165</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/ms-2026-165" rel="noopener noreferrer">10.5194/ms-2026-165</a></p>
<p><strong>Keywords:</strong> robotic grasping, collaborative robots, object detection, YOLOv8n, GR-ConvNet, MobileSAM, grasp pose estimation, RGB-D vision, unstructured scenes, intelligent manufacturing, machine learning, Hefei University of Technology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252717</post-id>	</item>
	</channel>
</rss>
