<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>integrating depth-based imagery with neural networks &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/integrating-depth-based-imagery-with-neural-networks/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 04:39:09 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>integrating depth-based imagery with neural networks &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to See Hands Like Humans: New End-to-End Model Reads 3D Hand Poses and Actions from Egocentric Video</title>
		<link>https://scienmag.com/ai-learns-to-see-hands-like-humans-new-end-to-end-model-reads-3d-hand-poses-and-actions-from-egocentric-video/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 04:39:09 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3D hand pose estimation]]></category>
		<category><![CDATA[3D hand pose estimation from egocentric video]]></category>
		<category><![CDATA[3D hand shape and position detection]]></category>
		<category><![CDATA[applications in robotics and assistive technology]]></category>
		<category><![CDATA[autonomous interpretation of hand movements]]></category>
		<category><![CDATA[challenges in hand gesture recognition from]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[egocentric vision]]></category>
		<category><![CDATA[egocentric vision in computer vision]]></category>
		<category><![CDATA[end-to-end deep learning for hand activity recognition]]></category>
		<category><![CDATA[FPHAB]]></category>
		<category><![CDATA[graph convolutional network]]></category>
		<category><![CDATA[hand activity recognition]]></category>
		<category><![CDATA[HOI4D dataset]]></category>
		<category><![CDATA[human-object interaction]]></category>
		<category><![CDATA[integrating depth-based imagery with neural networks]]></category>
		<category><![CDATA[neural networks for hand gesture analysis]]></category>
		<category><![CDATA[PA-ResGCN]]></category>
		<category><![CDATA[PA-ResGCN for hand action classification]]></category>
		<category><![CDATA[part-aware residual graph convolutional networks]]></category>
		<category><![CDATA[robotics]]></category>
		<category><![CDATA[TriHorn-NET]]></category>
		<category><![CDATA[TriHorn-NET for hand pose estimation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233542</guid>

					<description><![CDATA[Researchers have built an end-to-end deep learning framework that combines 3D hand pose estimation with hand action recognition from egocentric video, achieving near-real-time speeds and precision of nearly 100 percent on the HOI4D dataset.]]></description>
										<content:encoded><![CDATA[<p>Teaching a robot to pour a cup of coffee, or helping a blind person grasp an unfamiliar object, sounds like a simple task for a human being. For a machine, however, it requires solving two of the most stubborn problems in computer vision at the same time: figuring out the precise three-dimensional shape and position of a hand from a camera view, and then interpreting what that hand is actually doing with the objects around it. A new study published in Multimedia Tools and Applications tackles both challenges in a single pipeline, offering one of the most complete demonstrations yet that deep learning can carry out this entire process automatically, from raw egocentric video to recognized hand activity.</p>
<p>The research, led by Thi-Loan Nguyen of the Institute of Information Technology at Hanoi Pedagogical University 2 and Thai Nguyen University of Information and Communication Technology, together with colleagues at Tan Trao University and Hung Vuong University, describes an end-to-end deep learning framework that couples two complementary neural networks. The first, called TriHorn-NET, estimates the 3D pose of the hand from depth-based egocentric imagery. The second, PA-ResGCN, a part-aware residual graph convolutional network, takes the estimated skeleton and classifies the hand action being performed. By chaining these models together, the system can look through a head-mounted or robot-mounted camera and answer two questions in one pass: where exactly are the fingers, and what is the hand doing?</p>
<p>The choice of architecture was not arbitrary. The team systematically compared a range of state-of-the-art convolutional neural networks for each half of the problem. For 3D hand pose estimation, they evaluated SimpleHand, HaMuCo, and TriHorn-NET. For hand action recognition, they benchmarked ISTA-Net, PA-ResGCN, DD-Net, and MS-G3D. They also considered transformer-based end-to-end approaches for the combined task. This head-to-head evaluation, conducted on standard egocentric benchmarks, allowed the researchers to select the components that delivered the best balance of accuracy and speed rather than relying on a single fashionable architecture.</p>
<p>The numbers tell an interesting story about the trade-offs involved. On the HOI4D dataset, a large 4D egocentric collection of category-level human-object interactions, HaMuCo achieved the lowest pose estimation error, with a Procrustes-Aligned Mean Per Joint Position Error of 11.2 millimeters and a Procrustes-Aligned Mean Per Vertex Position Error of 7.2 millimeters. TriHorn-NET followed closely with a PA-MPJPE of 14.2 millimeters, a gap small enough that the researchers judged its other advantages worthwhile. Those advantages became clear in the downstream task: when TriHorn-NET&#8217;s estimated hand poses were fed into PA-ResGCN trained under a 3:7 data-splitting configuration, the action recognition stage reached a precision of 99.77 percent on HOI4D, the best result reported in the study.</p>
<p>Speed matters just as much as accuracy for real-world deployment, and here the framework also performed strongly. The fastest configuration, combining TriHorn-NET with DD-Net on HOI4D, processed video at 29.74 frames per second, while the full proposed model ran at 29.51 frames per second on a graphics processing unit. That figure sits just below the 30 frames per second threshold generally considered the floor for smooth real-time perception, meaning the pipeline is effectively capable of keeping pace with live video. For a robot arm that must react to a human hand in motion, or an assistive wearable that must warn a user before a grasp fails, that near-real-time throughput is what separates a laboratory curiosity from a practical tool.</p>
<p>The experiments were designed to stress the system under different data regimes. The team fine-tuned and tested their model on HOI4D under two end-to-end splitting configurations, 7:3 and 3:7, referring to the ratio of training to testing data, and additionally evaluated the framework on the First-Person Hand Action Benchmark, FPHAB, under the 7:3 configuration. FPHAB is a widely used benchmark containing RGB-D videos of first-person hand actions with 3D hand pose annotations, making it a natural proving ground for egocentric perception. Reporting results across both datasets and multiple splits gives a fuller picture of how the framework generalizes, and the paper provides detailed error measurements, confusion matrices for the action recognition stage, and overall computation times for the proposed model and every comparison method.</p>
<p>What makes this work technically significant is the way it handles the classic chicken-and-egg problem of egocentric hand analysis. Action recognition models typically perform best when given accurate skeleton data, but in a real deployment there are no ground-truth 3D annotations available; the skeleton must itself be estimated from the video. By evaluating the full chain end to end, with the action recognizer consuming estimated rather than annotated poses, the researchers measured performance under realistic conditions. The near-perfect precision of PA-ResGCN on HOI4D suggests that graph convolutional networks, which treat the hand skeleton as a graph of joints and bones and learn how information should flow between neighboring parts, are remarkably robust to the small errors introduced by upstream pose estimation.</p>
<p>The motivation behind the work reaches well beyond the benchmark numbers. The authors frame the problem around two concrete use cases: robot arms that must perform complex manipulations the way human hands do, and visually impaired people who need assistance grasping complicated objects in daily life. Egocentric vision, in which the camera shares the viewpoint of the person or robot, is the natural sensing modality for both. A humanoid robot&#8217;s gripper camera, a smart glass lens, or a wearable assistive device all see the world from a first-person perspective, where hands constantly enter and leave the frame, objects occlude the fingers, and lighting is unpredictable. Any system that works in this setting must be robust to exactly these challenges, which is why egocentric datasets like HOI4D and FPHAB have become central to the field.</p>
<p>The study also situates itself within a rapidly accelerating research landscape. Recent years have seen an explosion of interest in egocentric hand understanding, driven by new datasets such as H2O, Assembly101, HOT3D, and ThermoHands, and by increasingly powerful architectures ranging from PointNet-style point set networks to voxel-to-voxel prediction models, mesh recovery transformers, and graph-based methods like Hope-Net and HandFoldingNet. Earlier approaches to gesture and activity recognition relied on classical techniques such as support vector machines and random forests applied to wearable motion sensors, but the deep learning era has shifted the field toward models that learn spatial and temporal structure directly from images and skeletons. The Vietnamese team&#8217;s contribution is to show that a carefully selected combination of existing convolutional and graph-based components, fine-tuned end to end, can match or exceed more complex alternatives while running at usable speeds.</p>
<p>The authors conclude that their results demonstrate the end-to-end deep framework based on convolutional neural network models can be applied for estimation and recognition in building practical applications. In other words, the pieces needed for machines that understand human hands, where they are in three-dimensional space and what they are doing, are no longer scattered across separate research threads. They can be assembled into a single working system that runs fast enough to matter. As robots move out of factories and assistive technologies reach more users, that kind of integrated, real-time hand understanding may prove to be one of the quiet foundations of the next generation of human-centered machines. The research was funded by Tan Trao University in Tuyen Quang province, Vietnam, and the full results, including all comparison tables and evaluation details, are available in the journal article.</p>
<p><strong>Subject of Research:</strong> End-to-end deep learning for 3D hand pose estimation and hand activity recognition from egocentric vision datasets</p>
<p><strong>Article Title:</strong> Automatically end-to-end hand activity recognition based on 3D hand pose estimation from egocentric vision dataset</p>
<p><strong>Article References:</strong> Nguyen, T.-L., Phan, V.-N., Nguyen, V.-T., &amp; Le, V.-H. (2026). Automatically end-to-end hand activity recognition based on 3D hand pose estimation from egocentric vision dataset. <em>Multimedia Tools and Applications, 85</em>(9), Article 749. <a href="https://doi.org/10.1007/s11042-026-21804-7" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21804-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21804-7" rel="noopener noreferrer">10.1007/s11042-026-21804-7</a></p>
<p><strong>Keywords:</strong> 3D hand pose estimation, hand activity recognition, egocentric vision, deep learning, graph convolutional network, TriHorn-NET, PA-ResGCN, HOI4D dataset, FPHAB, human-object interaction, robotics, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233542</post-id>	</item>
	</channel>
</rss>
