<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Real-time processing &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/real-time-processing/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 07:59:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Real-time processing &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Drones Get Smarter: Graph Attention Networks Boost Real-Time Aerial Object Detection</title>
		<link>https://scienmag.com/drones-get-smarter-graph-attention-networks-boost-real-time-aerial-object-detection/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 07:59:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in autonomous drone navigation]]></category>
		<category><![CDATA[Aerial object detection using graph attention networks]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dense small object detection in computer vision]]></category>
		<category><![CDATA[drone imagery]]></category>
		<category><![CDATA[drone-based mapping and coordinate projection]]></category>
		<category><![CDATA[enhanced YOLOv7 for small object detection]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[graph attention networks]]></category>
		<category><![CDATA[handling occlusions in aerial object detection]]></category>
		<category><![CDATA[indoor and outdoor drone scene understanding]]></category>
		<category><![CDATA[integrating graph reasoning with deep learning for drones]]></category>
		<category><![CDATA[lightweight spatial projection in drone vision]]></category>
		<category><![CDATA[object detection]]></category>
		<category><![CDATA[real-time drone imagery analysis]]></category>
		<category><![CDATA[Real-time processing]]></category>
		<category><![CDATA[single GPU real-time aerial imagery processing]]></category>
		<category><![CDATA[small object detection]]></category>
		<category><![CDATA[spatial mapping]]></category>
		<category><![CDATA[UAV]]></category>
		<category><![CDATA[urban scene object recognition from aerial views]]></category>
		<category><![CDATA[VisDrone2019]]></category>
		<category><![CDATA[YOLOv7]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=237296</guid>

					<description><![CDATA[Researchers have built a drone perception framework that pairs an enhanced YOLOv7 detector with graph attention network reasoning and spatial projection, lifting small-object detection accuracy on VisDrone2019 while sustaining 65 frames per second.]]></description>
										<content:encoded><![CDATA[<p>Drones hovering over crowded city streets face one of the hardest problems in computer vision: spotting tiny, tightly packed objects—cars, pedestrians, bicycles—while flying at speed and mapping what they see onto a real-world coordinate system. A new study published in Discover Artificial Intelligence by Yang Gao, Congwei Liu, and Xiangyu Han of Handan University tackles both challenges at once, presenting a unified framework that couples an enhanced YOLOv7 detector with graph attention network reasoning and a lightweight spatial projection branch. The result is a system that not only detects more of the small targets that conventional detectors routinely miss, but also projects those detections onto a global map with measurably lower error, all while sustaining real-time performance on a single consumer GPU.</p>
<p>The core problem the researchers set out to solve is well known to anyone working with aerial imagery. Single-stage detectors from the YOLO family are prized for their inference speed, which makes them attractive for onboard drone processing, but their repeated downsampling operations tend to suppress the weak spatial responses produced by dense, small objects. In a typical urban scene captured from altitude, a vehicle or pedestrian may occupy only a handful of pixels and may be partially occluded by neighboring objects. When a detector evaluates each candidate in isolation, it cannot exploit the spatial arrangement or semantic compatibility of surrounding objects to resolve ambiguous features. The authors identify this absence of candidate-level context exchange within an efficient detector as the specific research gap their framework addresses, rather than simply weak multi-scale representation.</p>
<p>Their solution unfolds in three coordinated stages. First, the YOLOv7 backbone is fortified with a set of feature-enhancement modules chosen specifically to preserve small-object cues. SPD-Conv is inserted at the early downsampling stage, where it rearranges local spatial information into the channel dimension instead of discarding it through strided convolution, keeping fine-grained detail alive before compression begins. Coordinate Attention, applied after the E-ELAN feature extraction block, jointly encodes channel dependency and directional position information by pooling along the horizontal and vertical axes, allowing the network to focus precisely on target areas while suppressing background noise. SimAM, a parameter-free attention mechanism, recalibrates salient neurons after the major backbone stages by estimating neuron importance through an energy function, sharpening the contrast between small targets and complex backgrounds without adding trainable parameters.</p>
<p>The second enhancement concerns how features at different scales are combined. A Bidirectional Feature Pyramid Network, or BiFPN, replaces the conventional feature pyramid in the detector&#8217;s neck, fusing the P3, P4, and P5 feature maps with learnable weights along both top-down and bottom-up pathways. This bidirectional flow lets shallow layers rich in localization detail exchange information with deeper layers carrying semantic context, a combination that proves especially valuable when targets span a wide range of apparent sizes. Together, these four modules—SPD-Conv, Coordinate Attention, SimAM, and BiFPN—form the enhanced detection front end that feeds the framework&#8217;s more distinctive component: graph-based relational reasoning.</p>
<p>That component treats retained candidate detections as nodes in a sparse graph. After confidence filtering and non-maximum suppression, each surviving candidate region contributes a node containing its fused feature vector, bounding-box center, and category-confidence context. Edges connect nodes whose bounding-box centers lie within a distance threshold, with weights derived from the cosine similarity between feature vectors and gated by an indicator function enforcing local connectivity. A graph attention network then propagates information across this structure, using learnable attention coefficients to decide how much each neighbor should influence each node. Multi-head attention stabilizes training by concatenating outputs from several attention heads, each aggregating complementary contextual cues from different representation subspaces. After several layers of propagation, the graph-enriched features are projected back to the detector&#8217;s feature dimension and fused with the original candidate features through a residual connection followed by layer normalization, a design intended to mitigate over-smoothing.</p>
<p>Crucially, the graph is kept computationally tractable. A fully connected graph over hundreds of candidates in a dense urban frame would incur quadratic cost, so the authors sparsify it using a joint top-k and distance-threshold strategy: neighbors must satisfy the spatial threshold and are then ranked by joint spatial-feature relevance, with only the top eight retained per node. This reduces construction cost from O(N²) to approximately O(Nk), keeping computation proportional to the retained candidate set. Training is driven by a composite loss combining CIoU bounding-box regression, Focal Loss for classification with parameters fixed at 0.25 and 2.0, and a graph-consistency regularization term that encourages connected nodes to develop similar feature representations. The graph term is deliberately down-weighted at 0.1, serving as an auxiliary regularizer, and the coefficients were fixed after preliminary validation runs monitoring loss magnitudes, validation mAP, and convergence stability.</p>
<p>The third stage closes the loop between perception and geography. Rather than the traditional detect-first, locate-later workflow that accumulates delays, the framework synchronizes video frames with GPS and IMU metadata and runs two coordinated branches in parallel. The detection branch outputs object classes and refined image coordinates, while the mapping branch estimates camera pose from attitude and position data. Using a pinhole camera model, detected pixel coordinates are projected into world coordinates, with the missing depth information recovered from the drone&#8217;s altitude sensor and gimbal pitch angle under a locally flat ground assumption. Projected points are then smoothed across frames to generate a dynamic semantic map. Notably, the spatial reprojection error is used only for calibration and validation, not backpropagated through the detector, so the system should be understood as an integrated inference pipeline rather than a fully end-to-end optimized detection-and-mapping model.</p>
<p>The experimental results, obtained on the VisDrone2019 and UAVDT benchmarks with images resized to 640 by 640, are striking. The proposed YOLOv7-GAT framework achieves 42.6 percent mAP@0.5 and 25.8 percent mAP@0.5:0.95 on VisDrone2019, gains of 4.8 and 4.2 percentage points over the YOLOv7 baseline, while sustaining 65 frames per second on an NVIDIA RTX 3090 at batch size one. Confusion-matrix analysis reveals where the improvements concentrate: recall for pedestrians rises from 0.65 to 0.87, for people from 0.65 to 0.89, and for bicycles from 0.76 to 0.92. Inter-class confusion drops sharply as well, with pedestrian-to-people misclassification falling from 13 to 6 percent and tricycle-to-awning-tricycle errors falling from 14 to 6 percent. An ablation study confirms that each module contributes positively, with BiFPN delivering the largest fusion gain and the GAT module supplying additional contextual reasoning for ambiguous targets after candidate formation.</p>
<p>The mapping branch shows consistent benefits too. Across three tested scenario-altitude combinations on campus road, crossroad, and park settings, the corrected projection branch reduces root mean square error relative to traditional geometric projection, with most localization deviations falling below three pixels and axis-aligned errors remaining within roughly half a meter. Visualizations of the learned graph connectivity show dense urban scenes producing strong merged connectivity fields while sparse highway scenes form isolated local interaction regions, supporting the module&#8217;s adaptive behavior. The authors are candid about limitations: the flat-ground assumption degrades over rapidly changing terrain, low illumination reduces detection confidence, the fixed graph thresholds may limit adaptability in extreme scenes, and the 65 FPS figure applies only to the reported GPU configuration, with no edge-device benchmark yet performed.</p>
<p>Even with those caveats, the study offers a compelling demonstration that context is a resource detectors can learn to spend wisely. By letting each candidate consult its neighbors before committing to a prediction, and by folding mapping into the same pipeline rather than bolting it on afterward, the framework points toward drone systems that understand not just what they see but where it is—accurately, quickly, and within a single coherent architecture. The authors&#8217; stated next steps, including adaptive graph construction, benchmarking against recent DETR-style detectors, multi-sensor fusion with LiDAR, RTK-GPS, or digital elevation data, and lightweight deployment through graph pruning and quantization, suggest this integrated detection-and-mapping paradigm is only beginning to take flight.</p>
<p><strong>Subject of Research:</strong> A UAV object detection and spatial mapping framework integrating enhanced YOLOv7 with graph attention networks</p>
<p><strong>Article Title:</strong> A UAV object detection and spatial mapping framework integrating enhanced YOLOv7 with graph attention networks</p>
<p><strong>Article References:</strong> Gao, Y., Liu, C., &amp; Han, X. (2026). A UAV object detection and spatial mapping framework integrating enhanced YOLOv7 with graph attention networks. <em>Discover Artificial Intelligence, 6</em>(1), Article 1339. <a href="https://doi.org/10.1007/s44163-026-02004-6" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02004-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02004-6" rel="noopener noreferrer">10.1007/s44163-026-02004-6</a></p>
<p><strong>Keywords:</strong> UAV, object detection, YOLOv7, graph attention networks, spatial mapping, computer vision, VisDrone2019, small-object detection, real-time processing, deep learning, feature fusion, drone imagery</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">237296</post-id>	</item>
		<item>
		<title>AI Clears the Fog of Endoscopy: New Network Erases Glare and Flicker in Real Time</title>
		<link>https://scienmag.com/ai-clears-the-fog-of-endoscopy-new-network-erases-glare-and-flicker-in-real-time/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:02:11 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based specular reflection removal]]></category>
		<category><![CDATA[artificial intelligence in endoscopic imaging]]></category>
		<category><![CDATA[computational tools for endoscopy noise reduction]]></category>
		<category><![CDATA[deep learning for endoscopic video stabilization]]></category>
		<category><![CDATA[Depth estimation]]></category>
		<category><![CDATA[endoscopy video artifacts]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[flicker reduction in endoscopy]]></category>
		<category><![CDATA[GastroHUN]]></category>
		<category><![CDATA[Gastrointestinal endoscopy]]></category>
		<category><![CDATA[gastrointestinal lesion visibility improvement]]></category>
		<category><![CDATA[HyperKvasir]]></category>
		<category><![CDATA[Lumina-Net endoscopy image processing]]></category>
		<category><![CDATA[Medical image computing]]></category>
		<category><![CDATA[medical image inpainting techniques]]></category>
		<category><![CDATA[minimally invasive gastrointestinal diagnostics]]></category>
		<category><![CDATA[real-time endoscopic video enhancement]]></category>
		<category><![CDATA[Real-time processing]]></category>
		<category><![CDATA[real-time video denoising in medical procedures]]></category>
		<category><![CDATA[Robotic-assisted intervention]]></category>
		<category><![CDATA[Specular reflection removal]]></category>
		<category><![CDATA[Temporal learning]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[Video inpainting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204348</guid>

					<description><![CDATA[A new mask-guided video inpainting framework called Lumina-Net removes specular reflections and temporal flicker from endoscopic footage in near real time, improving both clinical visualization and downstream depth estimation.]]></description>
										<content:encoded><![CDATA[<p>Gastrointestinal endoscopy has transformed modern medicine, giving physicians a direct, minimally invasive window into the digestive tract for both diagnosis and therapy. Yet anyone who has watched raw endoscopic footage knows its most stubborn enemy: light itself. Because the procedure takes place in a wet, curved, enclosed cavity illuminated by a powerful point source, the mucosal surface frequently behaves like a mirror. The result is specular reflection, bright saturated patches that wash out tissue detail, together with temporal flicker that makes successive frames appear to pulse. For clinicians inspecting the lining of the stomach or colon, these artifacts can obscure early lesions. For the growing ecosystem of computational tools built on endoscopic video, from depth estimation to robotic navigation, they are a serious source of noise. A new artificial intelligence framework called Lumina-Net, described in Medical &amp; Biological Engineering &amp; Computing, now promises to strip away these distortions while keeping the video temporally stable and running fast enough for real-world use.</p>
<p>The study, led by Tianjun Yang, Xingfeng Xu and Siyang Zuo of Tianjin University together with gastroenterologist Xin Chen of Tianjin Medical University General Hospital, approaches the problem as an exercise in video inpainting: the task of filling in corrupted regions of an image sequence with plausible content drawn from the surrounding frames. Rather than treating each frame in isolation, as many earlier reflection-removal systems did, Lumina-Net exploits the fact that endoscopic video is inherently temporal. As the endoscope moves, the same patch of mucosa is viewed from slightly different angles at different moments, so information hidden behind a glare in one frame is often cleanly visible in its neighbors. The heart of the framework is a spatiotemporal Transformer that uses overlapping tokens, meaning small blocks of visual features that share context across both space and time. By allowing these tokens to aggregate complementary information from consecutive frames, the network can reconstruct the true tissue appearance beneath a reflection instead of simply painting over it with a generic texture.</p>
<p>Two lightweight modules in the decoder distinguish Lumina-Net from its predecessors. The first, Variance-Guided Feature Modulation, or VGFM, tackles a subtle but important problem: specular highlights in endoscopic imagery come in mixed scales, from tiny pinpoints of glare to large saturated blooms that cover a substantial fraction of the field of view. VGFM recalibrates network features using channel statistics, essentially measuring the variance of responses along each feature channel and using that measurement to decide how strongly to amplify or suppress them. This statistical steering allows a single network to handle both fine-grained and coarse-scale artifacts without resorting to separate models or heavy per-scale processing, keeping the added computational cost to a minimum.</p>
<p>The second module, Multi-Scale Energy Free-Space Attention, abbreviated MS-EFSA, is notable for carrying no learned parameters at all. Instead of adding weights that must be trained, it derives spatial attention weights directly from the energy of the features themselves. In practical terms, regions of the image that carry strong, reliable information about tissue structure receive more attention, while ambiguous or corrupted regions are down-weighted. The authors designed this mechanism with one anatomical priority in mind: preserving mucosal structures, the fine vascular and fold patterns of the gastrointestinal lining that clinicians rely on for diagnosis. Attention schemes that merely chase photometric consistency can smooth away exactly these clinically meaningful details. By anchoring attention to feature energy across multiple scales, MS-EFSA encourages the inpainted result to remain faithful to the underlying anatomy rather than producing a visually plausible but structurally hollow reconstruction.</p>
<p>The technical claims were tested on two public benchmarks: HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy, and GastroHUN, a dataset covering a complete systematic screening protocol for the stomach. Quantitatively, Lumina-Net achieved a peak signal-to-noise ratio of 30.20 decibels on these datasets and reduced mean squared error by approximately 5.3 percent compared with the strongest baseline method. While such numbers may sound incremental, in the tightly contested field of image restoration a 5 percent error reduction over a state-of-the-art competitor is meaningful, and PSNR above 30 decibels in a challenging medical domain indicates a reconstruction quality that is difficult to achieve. More striking is the speed: the network runs at 27 frames per second on a single graphics processing unit, placing it at the threshold of real-time video processing and making deployment alongside a live endoscopy workflow a realistic prospect rather than a laboratory aspiration.</p>
<p>Numbers alone, however, do not decide whether a restoration method is fit for clinical use, and the team supplemented the quantitative evaluation with a blinded study in which clinical experts compared the visual quality of outputs from Lumina-Net and competing approaches without knowing which method produced which result. The experts consistently preferred the proposed method, a finding the authors supported with established statistical procedures for ranked comparisons, including Wilcoxon-style rank testing and Kendall&#8217;s coefficient of concordance, which measures agreement among multiple raters. This convergence of expert judgment with objective metrics strengthens the case that the improvements are perceptually relevant, not merely artifacts of a particular error function.</p>
<p>Perhaps the most consequential demonstration concerns what happens downstream of the cleaned video. Modern endoscopy research increasingly depends on monocular depth estimation, the task of inferring three-dimensional scene structure from a single camera, which underpins applications such as autonomous scope navigation and robotic-assisted intervention. Glare and flicker are poison for these algorithms, because depth networks learn from photometric consistency between frames and reflections violate the assumptions that make that consistency meaningful. When the researchers applied depth estimation models to sequences processed by Lumina-Net, the resulting depth predictions improved, providing concrete evidence that reflection removal is not just cosmetic but functionally enables the robotic and navigational systems now under development in surgical laboratories worldwide.</p>
<p>The work fits into a research lineage stretching back nearly two decades, from early hand-crafted methods for detecting and inpainting specular highlights, through generative adversarial networks trained to synthesize glare-free tissue, to more recent temporal learning approaches and depth-aware endoscopic video inpainting presented at venues such as MICCAI. What Lumina-Net adds to this progression is a combination of architectural economy and clinical grounding. Its Transformer core borrows from general-purpose video inpainting frameworks such as FuseFormer and joint spatial-temporal transformation networks, but the VGFM and MS-EFSA modules are engineered specifically for the statistics of endoscopic imagery. The collaboration between engineering and clinical departments, funded by the National Natural Science Foundation of China under grant number 62133010, reflects a broader trend in which computer vision researchers and practicing gastroenterologists co-design tools around the actual failure modes of the imaging chain rather than abstract benchmarks.</p>
<p>The practical implications extend well beyond cleaner videos for human viewing. Reliable, glare-free endoscopic video is a prerequisite for the next generation of computer-assisted interventions: self-navigating capsule endoscopes that must map the stomach, surgical robots that need accurate tissue geometry, and diagnostic support systems that flag subtle early-stage lesions before they become advanced cancers. Because Lumina-Net operates at near-real-time speed and its complete code and pretrained weights are slated for release on GitHub upon acceptance, the barrier to integrating it into these pipelines is low. The study relied exclusively on publicly available, anonymized datasets, requiring no new ethics approval, and the authors report no conflicts of interest.</p>
<p>Caveats remain, as they always do with deep learning in medicine. The method was validated on two datasets and its generalization to unusual patient populations, atypical lighting hardware, or pathological tissue with markedly different reflectance properties will require further study. And like all generative restoration systems, inpainting networks must be used with care in diagnostic contexts, since any reconstructed pixel is by definition inferred rather than observed. Still, the combination of statistical rigor, expert validation and demonstrated downstream utility marks Lumina-Net as a serious step toward endoscopic video that is as clean as the underlying anatomy deserves. If the glare can be removed as reliably as this work suggests, both the eyes of the endoscopist and the algorithms of the robotic future may finally see the digestive tract clearly.</p>
<p><strong>Subject of Research:</strong> Deep learning-based specular reflection removal in gastrointestinal endoscopy video using temporal learning and feature fusion.</p>
<p><strong>Article Title:</strong> Lumina-Net: temporal learning with feature fusion for endoscopic artifact removal</p>
<p><strong>Article References:</strong> Lumina-Net: temporal learning with feature fusion for endoscopic artifact removal. (n.d.). <a href="https://doi.org/10.1007/s11517-026-03674-1" rel="noopener noreferrer">https://doi.org/10.1007/s11517-026-03674-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11517-026-03674-1" rel="noopener noreferrer">10.1007/s11517-026-03674-1</a></p>
<p><strong>Keywords:</strong> Gastrointestinal endoscopy, Specular reflection removal, Video inpainting, Temporal learning, Transformer, Feature fusion, Depth estimation, HyperKvasir, GastroHUN, Robotic-assisted intervention, Medical image computing, Real-time processing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204348</post-id>	</item>
	</channel>
</rss>
