<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Depth estimation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/depth-estimation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 23:15:08 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Depth estimation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI</title>
		<link>https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 23:15:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-driven robotic sensing]]></category>
		<category><![CDATA[artificial intelligence for tactile sensing]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[CLIP]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision in robotics]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Depth estimation]]></category>
		<category><![CDATA[feature pyramid network]]></category>
		<category><![CDATA[multimodal AI]]></category>
		<category><![CDATA[multimodal AI in robotics]]></category>
		<category><![CDATA[multimodal perception in robots]]></category>
		<category><![CDATA[multimodal sensor fusion]]></category>
		<category><![CDATA[neural networks for robotic manipulation]]></category>
		<category><![CDATA[pressure estimation]]></category>
		<category><![CDATA[robotic force estimation]]></category>
		<category><![CDATA[robotic grasping force estimation]]></category>
		<category><![CDATA[robotic manipulation]]></category>
		<category><![CDATA[robotic manipulation without tactile sensors]]></category>
		<category><![CDATA[robotics]]></category>
		<category><![CDATA[sensorless force detection]]></category>
		<category><![CDATA[soft grippers]]></category>
		<category><![CDATA[tactile sensing]]></category>
		<category><![CDATA[visual-based pressure sensing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=250325</guid>

					<description><![CDATA[A new multimodal AI framework fuses RGB images, segmentation masks, depth maps, and text prompts to estimate robotic gripper pressure from vision alone, beating state-of-the-art methods on tendon-actuated grippers.]]></description>
										<content:encoded><![CDATA[<p>Robotic hands have long faced a fundamental dilemma: to grip an egg without crushing it or a wrench without dropping it, a robot needs to know exactly how hard it is pressing, yet the sensors that provide that answer are often bulky, expensive, and fragile. Now, a team of researchers has unveiled a new artificial intelligence framework that lets robotic grippers estimate the pressure they exert on objects using nothing more than cameras and a remarkably rich fusion of visual, geometric, and even linguistic information. The work, published in the journal Results in Engineering, demonstrates that when a neural network is allowed to look at a scene through several complementary lenses at once, it can outperform the best purely visual methods at reading the invisible language of force.</p>
<p>The research, led by Yawen Liu and Wei Tang with colleagues Zhen Zhang and Xinrong Chen, addresses a persistent bottleneck in robotic manipulation. Conventional approaches to force sensing rely on physical instruments such as force and torque sensors or tactile arrays mounted directly on the gripper. Rigid-body models calculate grasping force through explicit mechanical equations, while soft-body models attempt to capture the nonlinear deformations of compliant structures. Although advances in materials science, including conductive hydrogels that endow soft robots with human-like sensory capabilities, have pushed the field forward, physical sensors bring inherent drawbacks: they add system complexity and cost, and their measurement range is limited, making it difficult to adapt to the diverse and dynamic demands of different grasping tasks.</p>
<p>Vision-based pressure estimation has emerged as an elegant alternative. By pointing cameras at the interaction between a gripper and an object, computer vision techniques can infer the magnitude and distribution of applied forces without any physical contact with the measured surface. This eliminates the need for structural modifications to the gripper and offers flexible, scalable perception. Yet most existing methods, including the current state-of-the-art system known as HiVPE, rely on single RGB images. That restriction means the network sees only surface appearance, missing explicit three-dimensional geometric deformation and high-level semantic context, both of which become crucial when dealing with severe occlusions or the complex hyperelastic behavior of soft materials.</p>
<p>The new framework breaks free of that single-modality constraint by weaving together four distinct streams of information. The first is the ordinary RGB image capturing the interaction between the gripper and a high-resolution pressure sensing array. The second is a binary segmentation mask that precisely delineates the gripper&#8217;s contact regions, generated by a pre-trained deep segmentation model that isolates the hand from the background. The third is a depth map produced from the RGB image by a monocular depth estimation network, encoding the geometric and spatial relationships of the contact. The fourth is perhaps the most surprising: a short textual prompt, either the word Straight or Roll, describing the gripper&#8217;s action, which is converted into a semantic embedding by the CLIP text encoder, the same contrastive language-image model that underpins many modern multimodal AI systems.</p>
<p>Architecturally, the system proceeds in three stages: feature extraction, feature fusion, and decoding. During extraction, a dual-branch hierarchical encoding scheme employs two parallel SE-ResNeXt50 networks. The first processes the raw RGB image to produce five feature maps at progressively abstract scales, spanning downsampling factors from 2 to 32 and channel dimensions from 64 to 2048, capturing everything from low-level textures to high-level semantics. The second encoder jointly processes the concatenated segmentation and depth inputs, yielding a parallel set of features rich in structural and geometric cues. In this way, one branch learns what the scene looks like while the other learns how it is shaped in space.</p>
<p>The fusion stage is where the framework&#8217;s distinctive machinery comes into play. Multi-head self-attention, with four attention heads of 64 dimensions each operating on 256-dimensional features, first refines the highest-resolution feature maps within each modality, strengthening intra-modal contextual dependencies through residual connections and feed-forward networks. A bidirectional cross-attention operator then lets the two streams talk to each other: the appearance features absorb structural context from the segmentation and depth branch, while the structural features incorporate appearance cues from the RGB branch. Finally, the 512-dimensional textual embedding is projected through a multilayer perceptron into a 2880-dimensional feature space, reshaped to align spatially with the visual features and concatenated to form a unified multimodal representation. A Feature Pyramid Network decoder then performs top-down aggregation, fusing high-level semantic cues with low-level spatial details before a segmentation head produces the final dense pressure prediction.</p>
<p>Training and evaluation were carried out on a publicly available benchmark containing roughly 650,000 synchronized RGB-pressure frames from two very different soft grippers, one tendon-actuated and one pneumatic. Ground-truth pressure annotations came from a high-resolution Sensel Morph pressure sensing array, spatially aligned with the camera images through calibration and homography transformation. The dataset covers four interaction types: contact, slide, close, and no-contact, the last serving as adversarial samples. The authors discretized continuous pressure values into nine logarithmically spaced categories, from zero background to a top bin of 64 kilopascals and above, and optimized the network with a pixel-wise weighted cross-entropy loss that penalizes underrepresented contact classes to counteract the imbalance between vast non-contact regions and sparse pressure-bearing areas. All results were averaged over six independent runs to verify statistical stability.</p>
<p>The results are striking for the tendon-actuated gripper, where the multimodal method surpassed the previous state of the art on every metric. It achieved a temporal accuracy of 96.57 percent, a contact area intersection-over-union of 77.20 percent, a volumetric IoU of 62.87 percent, and a mean absolute error of just 4.33 pascals, compared with 5.0 pascals for HiVPE and 5.3 for the earlier VPEC-Net. For the pneumatic gripper, the method attained the highest temporal accuracy and volumetric IoU, though it fell marginally short of HiVPE on contact IoU and mean absolute error, a shortfall the authors attribute to multimodal fusion occasionally introducing noise when soft surfaces deform so severely that depth and segmentation maps become blurred. Ablation experiments confirmed that each component earns its place: segmentation masks alone improved performance, adding attention strengthened contact-region predictions, depth maps reduced error further, and textual prompts delivered the final gains in spatial-pressure consistency.</p>
<p>The implications extend well beyond a leaderboard. Because the approach requires no embedded sensors, it could make force-aware grasping accessible to cheap, commodity robotic hardware, from laboratory manipulators to warehouse pickers handling fragile goods. The demonstration that a single descriptive word about the gripper&#8217;s action can measurably sharpen pressure predictions hints at a future where robots are instructed in natural language and perceive the physical consequences of their movements through vision alone. The authors are candid about the current limitations: the framework has so far been validated on static, offline benchmark data rather than deployed on physical robots in dynamic manipulation, where real-time sensor noise, unpredictable physical interactions, and inference latency constraints all come into play. They also note that soft grippers, with their complex nonlinear hyperelastic deformation, remain the hardest case, and they point toward more advanced depth estimation, adaptive fusion architectures that dynamically re-weight unreliable modalities, and explicit deformation modeling as the next frontiers. If those challenges are met, the line between seeing and touching may grow ever thinner, bringing robots one step closer to hands that truly understand what they hold.</p>
<p><strong>Subject of Research:</strong> Vision-based multimodal pressure estimation for robotic gripper manipulators</p>
<p><strong>Article Title:</strong> Vision-Based multimodal pressure estimation for robotic gripper manipulators</p>
<p><strong>Article References:</strong> Liu, Y., Tang, W., Zhang, Z., &amp; Chen, X. (2026). Vision-Based multimodal pressure estimation for robotic gripper manipulators. <em>Results in Engineering, 32</em>, Article 113295. <a href="https://doi.org/10.1016/j.rineng.2026.113295" rel="noopener noreferrer">https://doi.org/10.1016/j.rineng.2026.113295</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.rineng.2026.113295" rel="noopener noreferrer">10.1016/j.rineng.2026.113295</a></p>
<p><strong>Keywords:</strong> robotics, pressure estimation, multimodal AI, soft grippers, computer vision, deep learning, attention mechanism, CLIP, tactile sensing, depth estimation, feature pyramid network, robotic manipulation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">250325</post-id>	</item>
		<item>
		<title>New AI Learns to Track Surgical Robots in Real Time, Even When Blood and Smoke Block the View</title>
		<link>https://scienmag.com/new-ai-learns-to-track-surgical-robots-in-real-time-even-when-blood-and-smoke-block-the-view/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 12:45:49 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[6DoF pose estimation]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[da Vinci]]></category>
		<category><![CDATA[Depth estimation]]></category>
		<category><![CDATA[even in challenging conditions like blood and smoke blockage]]></category>
		<category><![CDATA[geometric consistency]]></category>
		<category><![CDATA[markerless tracking]]></category>
		<category><![CDATA[multi-task learning]]></category>
		<category><![CDATA[occlusion robustness]]></category>
		<category><![CDATA[PICO]]></category>
		<category><![CDATA[real-time inference]]></category>
		<category><![CDATA[Surgical robotics]]></category>
		<category><![CDATA[surgical robots in real time]]></category>
		<category><![CDATA[SurgRIPE benchmark]]></category>
		<category><![CDATA[using advanced vision-based techniques]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212406</guid>

					<description><![CDATA[Researchers at the University of Leeds have developed PICO, a single-stage AI system that estimates the full 3D position and orientation of surgical tools in real time from one camera image, achieving near-benchmark accuracy even under occlusion.]]></description>
										<content:encoded><![CDATA[<p>Every time a surgical robot reaches into a patient&#8217;s body, its control system needs to know exactly where the instrument is: three coordinates of position and three angles of orientation, together known as a six degree-of-freedom, or 6DoF, pose. For decades, the da Vinci Surgical System and its research derivatives have relied on forward kinematics, chains of mathematical equations that translate joint-angle readings into tool-tip coordinates. The approach is elegant on paper, but the hardware betrays it. Cable slack, friction between cables and pulleys, and the accumulation of small errors across multiple joints can leave discrepancies of up to 1.02 millimetres, most visibly at the end effector, the very part of the instrument that touches tissue. In an operating theatre, a millimetre is not a rounding error; it is the difference between a clean incision and a severed nerve.</p>
<p>A team of researchers at the University of Leeds, led by Lucy Fothergill and Duygu Sarikaya of the School of Computer Science, together with Pietro Valdastri and Dominic Jones from the School of Electronic and Electrical Engineering, has now unveiled a vision-based alternative designed to close that gap. Their system, called PICO for Projection-Informed Consistency Optimisation, estimates the full 6DoF pose of a surgical tool directly from a single monocular RGB image, with no markers, no trackers and no external hardware attached to the instrument. Published in the International Journal of Computer Assisted Radiology and Surgery, the work demonstrates that a single, end-to-end trainable neural network can rival far more cumbersome multi-stage pipelines while running at true real-time speed, and it holds up even when blood, smoke or other instruments obscure the camera&#8217;s view.</p>
<p>The case for markerless vision is straightforward once the constraints of the operating room are understood. External markers and trackers demand sterilisation procedures that interrupt surgical workflow, and they require an unobstructed line of sight between camera and marker, something a surgical field full of tissue, fluid and smoke simply cannot guarantee. Yet most existing markerless approaches dodge the hardest part of the problem. Rather than predicting the 6DoF pose outright, they first extract intermediate representations, such as 2D keypoints, segmentation masks or dense 2D-3D correspondences, and only then compute the pose using a Perspective-n-Point solver, template matching, template-based rendering or iterative refinement. Each extra stage is another place where noise accumulates, another source of computational delay, and another dependency that can fail. If the initial estimate is poor, the refinement step may never fully correct it, and iterative render-and-compare loops are notoriously slow at inference time, degrading performance in the fast-changing environment of a live operation.</p>
<p>PICO&#8217;s architecture attacks the problem from a different direction. Given a cropped monocular RGB image, a shared encoder based on ResNet-50, a widely used deep residual network pre-trained on ImageNet, extracts visual features. Those features feed two parallel paths. One path runs through a U-Net style decoder, the workhorse of biomedical image segmentation, into two task-specific heads: a segmentation head that produces a binary mask of the tool, and a depth head that outputs a pseudo-depth map of the scene. The other path bypasses the decoder, passing the encoder features through fully connected layers into regression heads that directly output the tool&#8217;s rotation and translation. Rotation is regressed as a full 3&#215;3 matrix, which is projected onto the nearest valid rotation in the mathematical group SO(3) using Singular Value Decomposition, a parameterisation the authors chose to avoid the singularities and ambiguities that plague Euler angles and quaternions. Translation is split into the (x, y) pixel coordinates of the tool joint in the image frame and a separate depth value z relative to the camera, each with its own activation function tuned to the range of plausible values.</p>
<p>The genuinely novel ingredient is the pair of geometric consistency losses, or proxy tasks, that bind these outputs together. The projection loss takes the network&#8217;s predicted pose, applies it to a randomly sampled set of 3D points on the tool&#8217;s mesh model, and projects those points into the 2D image plane using the camera&#8217;s intrinsics. The contour of the resulting concave hull yields a projected binary mask, which acts as a pseudo-ground-truth against which the predicted segmentation mask is scored with a Dice loss. If the network&#8217;s pose and its segmentation disagree, the loss rises, forcing the model to keep its 2D appearance and its 3D geometry in register. The point-to-point loss works entirely in 3D: the predicted and ground-truth pose transformations are applied to the same sampled model points, and the root mean squared error between corresponding points is minimised. Together with a geodesic loss that measures the true angular distance between rotation matrices, and separate root-mean-squared-error terms for the (x, y) and z translation components, the full multi-task objective supervises the network simultaneously in 2D image space and 3D model space.</p>
<p>The auxiliary tasks proved to have complementary but delicate roles. In ablation studies across four test datasets, adding depth supervision alone yielded the lowest translation errors on three of the four sets, for example cutting error from 13.28 to 8.36 millimetres on the occluded large needle driver set, because it primarily constrains the depth component of the pose. Segmentation supervision, by contrast, sharpened rotation estimates by encoding the tool&#8217;s projected shape, improving rotation error from 18.89 to 11.11 degrees on the large needle driver and from 12.95 to 9.67 degrees on the Maryland bipolar forceps. Curiously, combining the two naively did not stack the gains and sometimes hurt performance, but the projection and point-to-point losses, which couple both signals through a single predicted transformation, resolved the trade-off and produced the best overall results.</p>
<p>Benchmarked on the SurgRIPE dataset, introduced at the MICCAI 2022 SurgRIPE challenge and still the only public benchmark with ground-truth 6DoF pose annotations for surgical instruments, PICO ranked second in rotational accuracy across all four test sets, covering two tools, the large needle driver and the Maryland bipolar forceps, each in occluded and unoccluded conditions. It recorded rotation errors of 5.78 degrees on the large needle driver and 21.02 degrees on the occluded forceps, trailing only the top-performing multi-stage entry from ImFusion, which combined SurfEmb surface embeddings with an iterative render-and-compare refinement. PICO&#8217;s translational performance remained competitive, particularly under occlusion, where its 8.87-millimetre error on the occluded needle driver far outperformed the 28.09 millimetres of PVNet, a widely cited two-stage method. Against the only other single-stage method in the benchmark, PICO cut rotation errors dramatically, from 27.21 to 5.78 degrees on the large needle driver and from 34.13 to 21.02 degrees on the occluded forceps.</p>
<p>Perhaps the most impressive figure is the runtime. On a consumer-grade NVIDIA Tesla T4 GPU, PICO completes inference at 33.3 frames per second, roughly 30 milliseconds per image, clearing the customary 30 FPS threshold for real-time performance. PVNet, by comparison, manages about 25 FPS even on a GTX 1080ti. At inference time the U-Net decoder is simply discarded, stripping away computational overhead, while a fine-tuned YOLOv5 detector locates the tool in the image, achieving intersection-over-union scores as high as 0.89. The combination of single-step prediction and geometric supervision means no iterative refinement, no correspondence solving and no render-and-compare loop stand between the camera image and the pose estimate.</p>
<p>The authors were unusually candid about the method&#8217;s remaining weakness. PICO scored lowest of all benchmarked methods on the ADD metric, which measures the fraction of samples whose average distance between ground-truth and predicted point clouds falls below 10 percent of the instrument&#8217;s diameter. An error decomposition revealed why: depth error along the camera axis, with mean values as high as 11.82 millimetres on the occluded forceps set, tracks ADD distance almost perfectly, with Spearman correlations of 0.955 to 0.982, dwarfing the correlations of image-plane error. In other words, the pseudo-depth maps generated by sampling mesh points and assigning them to the nearest pixel, then filling holes with neighbourhood averages, remain too coarse to supervise depth precisely. The team also acknowledges that resizing non-square crops to a fixed 224-pixel square can introduce mild aspect-ratio distortion, and that the scarcity of public surgical pose datasets restricts comparisons to benchmark-reported figures rather than fully reproducible implementations.</p>
<p>Even so, the trajectory of the work is hard to ignore. Accurate, markerless, real-time 6DoF tool pose estimation is a prerequisite for surgical autonomy, robotic proprioception and safe tissue interaction, and PICO demonstrates that an end-to-end network guided by geometric consistency can match multi-stage pipelines without their latency or fragility. The Leeds group&#8217;s next steps, refining depth modelling and extending generalisation to unseen instruments and environments through domain adaptation and data augmentation, will determine how quickly this kind of software settles into the operating theatre. But the central message already stands: a neural network that forces its own 2D and 3D views of the world to agree can track a robot&#8217;s instruments with the speed and reliability that surgical autonomy demands.</p>
<p><strong>Subject of Research:</strong> Real-time markerless 6DoF pose estimation of surgical instruments from monocular images using multi-task learning and geometric consistency losses</p>
<p><strong>Article Title:</strong> PICO: Projection-Informed Consistency Optimisation for 6DoF surgical tool pose estimation</p>
<p><strong>Article References:</strong> Fothergill, L., Valdastri, P., Jones, D., &amp; Sarikaya, D. (2026). PICO: Projection-Informed Consistency Optimisation for 6DoF surgical tool pose estimation. <em>International Journal of Computer Assisted Radiology and Surgery</em>. <a href="https://doi.org/10.1007/s11548-026-03802-0" rel="noopener noreferrer">https://doi.org/10.1007/s11548-026-03802-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11548-026-03802-0" rel="noopener noreferrer">10.1007/s11548-026-03802-0</a></p>
<p><strong>Keywords:</strong> surgical robotics, 6DoF pose estimation, PICO, multi-task learning, geometric consistency, SurgRIPE benchmark, markerless tracking, depth estimation, computer vision, da Vinci, real-time inference, occlusion robustness</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212406</post-id>	</item>
		<item>
		<title>AI Clears the Fog of Endoscopy: New Network Erases Glare and Flicker in Real Time</title>
		<link>https://scienmag.com/ai-clears-the-fog-of-endoscopy-new-network-erases-glare-and-flicker-in-real-time/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:02:11 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based specular reflection removal]]></category>
		<category><![CDATA[artificial intelligence in endoscopic imaging]]></category>
		<category><![CDATA[computational tools for endoscopy noise reduction]]></category>
		<category><![CDATA[deep learning for endoscopic video stabilization]]></category>
		<category><![CDATA[Depth estimation]]></category>
		<category><![CDATA[endoscopy video artifacts]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[flicker reduction in endoscopy]]></category>
		<category><![CDATA[GastroHUN]]></category>
		<category><![CDATA[Gastrointestinal endoscopy]]></category>
		<category><![CDATA[gastrointestinal lesion visibility improvement]]></category>
		<category><![CDATA[HyperKvasir]]></category>
		<category><![CDATA[Lumina-Net endoscopy image processing]]></category>
		<category><![CDATA[Medical image computing]]></category>
		<category><![CDATA[medical image inpainting techniques]]></category>
		<category><![CDATA[minimally invasive gastrointestinal diagnostics]]></category>
		<category><![CDATA[real-time endoscopic video enhancement]]></category>
		<category><![CDATA[Real-time processing]]></category>
		<category><![CDATA[real-time video denoising in medical procedures]]></category>
		<category><![CDATA[Robotic-assisted intervention]]></category>
		<category><![CDATA[Specular reflection removal]]></category>
		<category><![CDATA[Temporal learning]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[Video inpainting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204348</guid>

					<description><![CDATA[A new mask-guided video inpainting framework called Lumina-Net removes specular reflections and temporal flicker from endoscopic footage in near real time, improving both clinical visualization and downstream depth estimation.]]></description>
										<content:encoded><![CDATA[<p>Gastrointestinal endoscopy has transformed modern medicine, giving physicians a direct, minimally invasive window into the digestive tract for both diagnosis and therapy. Yet anyone who has watched raw endoscopic footage knows its most stubborn enemy: light itself. Because the procedure takes place in a wet, curved, enclosed cavity illuminated by a powerful point source, the mucosal surface frequently behaves like a mirror. The result is specular reflection, bright saturated patches that wash out tissue detail, together with temporal flicker that makes successive frames appear to pulse. For clinicians inspecting the lining of the stomach or colon, these artifacts can obscure early lesions. For the growing ecosystem of computational tools built on endoscopic video, from depth estimation to robotic navigation, they are a serious source of noise. A new artificial intelligence framework called Lumina-Net, described in Medical &amp; Biological Engineering &amp; Computing, now promises to strip away these distortions while keeping the video temporally stable and running fast enough for real-world use.</p>
<p>The study, led by Tianjun Yang, Xingfeng Xu and Siyang Zuo of Tianjin University together with gastroenterologist Xin Chen of Tianjin Medical University General Hospital, approaches the problem as an exercise in video inpainting: the task of filling in corrupted regions of an image sequence with plausible content drawn from the surrounding frames. Rather than treating each frame in isolation, as many earlier reflection-removal systems did, Lumina-Net exploits the fact that endoscopic video is inherently temporal. As the endoscope moves, the same patch of mucosa is viewed from slightly different angles at different moments, so information hidden behind a glare in one frame is often cleanly visible in its neighbors. The heart of the framework is a spatiotemporal Transformer that uses overlapping tokens, meaning small blocks of visual features that share context across both space and time. By allowing these tokens to aggregate complementary information from consecutive frames, the network can reconstruct the true tissue appearance beneath a reflection instead of simply painting over it with a generic texture.</p>
<p>Two lightweight modules in the decoder distinguish Lumina-Net from its predecessors. The first, Variance-Guided Feature Modulation, or VGFM, tackles a subtle but important problem: specular highlights in endoscopic imagery come in mixed scales, from tiny pinpoints of glare to large saturated blooms that cover a substantial fraction of the field of view. VGFM recalibrates network features using channel statistics, essentially measuring the variance of responses along each feature channel and using that measurement to decide how strongly to amplify or suppress them. This statistical steering allows a single network to handle both fine-grained and coarse-scale artifacts without resorting to separate models or heavy per-scale processing, keeping the added computational cost to a minimum.</p>
<p>The second module, Multi-Scale Energy Free-Space Attention, abbreviated MS-EFSA, is notable for carrying no learned parameters at all. Instead of adding weights that must be trained, it derives spatial attention weights directly from the energy of the features themselves. In practical terms, regions of the image that carry strong, reliable information about tissue structure receive more attention, while ambiguous or corrupted regions are down-weighted. The authors designed this mechanism with one anatomical priority in mind: preserving mucosal structures, the fine vascular and fold patterns of the gastrointestinal lining that clinicians rely on for diagnosis. Attention schemes that merely chase photometric consistency can smooth away exactly these clinically meaningful details. By anchoring attention to feature energy across multiple scales, MS-EFSA encourages the inpainted result to remain faithful to the underlying anatomy rather than producing a visually plausible but structurally hollow reconstruction.</p>
<p>The technical claims were tested on two public benchmarks: HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy, and GastroHUN, a dataset covering a complete systematic screening protocol for the stomach. Quantitatively, Lumina-Net achieved a peak signal-to-noise ratio of 30.20 decibels on these datasets and reduced mean squared error by approximately 5.3 percent compared with the strongest baseline method. While such numbers may sound incremental, in the tightly contested field of image restoration a 5 percent error reduction over a state-of-the-art competitor is meaningful, and PSNR above 30 decibels in a challenging medical domain indicates a reconstruction quality that is difficult to achieve. More striking is the speed: the network runs at 27 frames per second on a single graphics processing unit, placing it at the threshold of real-time video processing and making deployment alongside a live endoscopy workflow a realistic prospect rather than a laboratory aspiration.</p>
<p>Numbers alone, however, do not decide whether a restoration method is fit for clinical use, and the team supplemented the quantitative evaluation with a blinded study in which clinical experts compared the visual quality of outputs from Lumina-Net and competing approaches without knowing which method produced which result. The experts consistently preferred the proposed method, a finding the authors supported with established statistical procedures for ranked comparisons, including Wilcoxon-style rank testing and Kendall&#8217;s coefficient of concordance, which measures agreement among multiple raters. This convergence of expert judgment with objective metrics strengthens the case that the improvements are perceptually relevant, not merely artifacts of a particular error function.</p>
<p>Perhaps the most consequential demonstration concerns what happens downstream of the cleaned video. Modern endoscopy research increasingly depends on monocular depth estimation, the task of inferring three-dimensional scene structure from a single camera, which underpins applications such as autonomous scope navigation and robotic-assisted intervention. Glare and flicker are poison for these algorithms, because depth networks learn from photometric consistency between frames and reflections violate the assumptions that make that consistency meaningful. When the researchers applied depth estimation models to sequences processed by Lumina-Net, the resulting depth predictions improved, providing concrete evidence that reflection removal is not just cosmetic but functionally enables the robotic and navigational systems now under development in surgical laboratories worldwide.</p>
<p>The work fits into a research lineage stretching back nearly two decades, from early hand-crafted methods for detecting and inpainting specular highlights, through generative adversarial networks trained to synthesize glare-free tissue, to more recent temporal learning approaches and depth-aware endoscopic video inpainting presented at venues such as MICCAI. What Lumina-Net adds to this progression is a combination of architectural economy and clinical grounding. Its Transformer core borrows from general-purpose video inpainting frameworks such as FuseFormer and joint spatial-temporal transformation networks, but the VGFM and MS-EFSA modules are engineered specifically for the statistics of endoscopic imagery. The collaboration between engineering and clinical departments, funded by the National Natural Science Foundation of China under grant number 62133010, reflects a broader trend in which computer vision researchers and practicing gastroenterologists co-design tools around the actual failure modes of the imaging chain rather than abstract benchmarks.</p>
<p>The practical implications extend well beyond cleaner videos for human viewing. Reliable, glare-free endoscopic video is a prerequisite for the next generation of computer-assisted interventions: self-navigating capsule endoscopes that must map the stomach, surgical robots that need accurate tissue geometry, and diagnostic support systems that flag subtle early-stage lesions before they become advanced cancers. Because Lumina-Net operates at near-real-time speed and its complete code and pretrained weights are slated for release on GitHub upon acceptance, the barrier to integrating it into these pipelines is low. The study relied exclusively on publicly available, anonymized datasets, requiring no new ethics approval, and the authors report no conflicts of interest.</p>
<p>Caveats remain, as they always do with deep learning in medicine. The method was validated on two datasets and its generalization to unusual patient populations, atypical lighting hardware, or pathological tissue with markedly different reflectance properties will require further study. And like all generative restoration systems, inpainting networks must be used with care in diagnostic contexts, since any reconstructed pixel is by definition inferred rather than observed. Still, the combination of statistical rigor, expert validation and demonstrated downstream utility marks Lumina-Net as a serious step toward endoscopic video that is as clean as the underlying anatomy deserves. If the glare can be removed as reliably as this work suggests, both the eyes of the endoscopist and the algorithms of the robotic future may finally see the digestive tract clearly.</p>
<p><strong>Subject of Research:</strong> Deep learning-based specular reflection removal in gastrointestinal endoscopy video using temporal learning and feature fusion.</p>
<p><strong>Article Title:</strong> Lumina-Net: temporal learning with feature fusion for endoscopic artifact removal</p>
<p><strong>Article References:</strong> Lumina-Net: temporal learning with feature fusion for endoscopic artifact removal. (n.d.). <a href="https://doi.org/10.1007/s11517-026-03674-1" rel="noopener noreferrer">https://doi.org/10.1007/s11517-026-03674-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11517-026-03674-1" rel="noopener noreferrer">10.1007/s11517-026-03674-1</a></p>
<p><strong>Keywords:</strong> Gastrointestinal endoscopy, Specular reflection removal, Video inpainting, Temporal learning, Transformer, Feature fusion, Depth estimation, HyperKvasir, GastroHUN, Robotic-assisted intervention, Medical image computing, Real-time processing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204348</post-id>	</item>
	</channel>
</rss>
