<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>optical flow &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/optical-flow/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 12:57:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>optical flow &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Rebuilds Missing Radar Echoes Without Knowing Where the Gaps Are</title>
		<link>https://scienmag.com/ai-rebuilds-missing-radar-echoes-without-knowing-where-the-gaps-are/</link>
		
		<dc:creator><![CDATA[Russell Cooper]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 12:57:38 +0000</pubDate>
				<category><![CDATA[Athmospheric]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based weather radar hole filling]]></category>
		<category><![CDATA[Atmospheric Measurement Techniques]]></category>
		<category><![CDATA[atmospheric measurement techniques for radar data]]></category>
		<category><![CDATA[automatic radar echo gap detection]]></category>
		<category><![CDATA[BiConvLSTM-UNet]]></category>
		<category><![CDATA[ConvLSTM]]></category>
		<category><![CDATA[data reconstruction]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for radar echo reconstruction]]></category>
		<category><![CDATA[meteorological radar mosaic analysis]]></category>
		<category><![CDATA[multi-radar data fusion challenges]]></category>
		<category><![CDATA[nowcasting]]></category>
		<category><![CDATA[optical flow]]></category>
		<category><![CDATA[precipitation estimation]]></category>
		<category><![CDATA[radar hardware failure impact on weather monitoring]]></category>
		<category><![CDATA[radar mosaic]]></category>
		<category><![CDATA[real-time weather radar image enhancement]]></category>
		<category><![CDATA[severe convective weather]]></category>
		<category><![CDATA[severe weather forecasting with missing data]]></category>
		<category><![CDATA[storm tracking with incomplete radar data]]></category>
		<category><![CDATA[terrain and building interference in weather radar]]></category>
		<category><![CDATA[U-Net]]></category>
		<category><![CDATA[weather radar]]></category>
		<category><![CDATA[weather radar data gap filling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247850</guid>

					<description><![CDATA[A new deep learning model called BiConvLSTM-UNet can restore missing regions in weather radar mosaic data without requiring any map of where the data gaps are, outperforming optical flow and other deep learning approaches across multiple missing-data scenarios.]]></description>
										<content:encoded><![CDATA[<p>When a severe thunderstorm is bearing down on a city, forecasters rely on a seamless mosaic of radar echoes stitched together from multiple radars to track the storm&#8217;s every move. But that mosaic is only as good as its weakest link. A single radar&#8217;s hardware failure, a delayed file transfer, or a software glitch can carve a silent hole into the composite picture, leaving forecasters blind over an entire region at the worst possible moment. A new study published in Atmospheric Measurement Techniques by Husong Guo, Muyun Du, and colleagues presents a deep learning method that can fill those holes automatically, and it does so without ever being told where the holes are.</p>
<p>The problem is more consequential than it might first appear. A single weather radar can only see effectively out to roughly 200 to 300 kilometers, and terrain, buildings, and vegetation can block parts of its view. To monitor large-scale weather systems continuously, meteorological agencies fuse observations from networks of radars into mosaic products. In China, the operational Severe Weather Automatic Nowcasting system, run by the Hubei Meteorological Observatory, generates composite reflectivity mosaics from nine S-band radars at a spatial resolution of one kilometer and a temporal resolution of six minutes. These products underpin quantitative precipitation estimation and the nowcasting of severe convective weather. When a region of the mosaic goes dark, the accuracy of rainfall estimates and storm warnings degrades immediately.</p>
<p>Existing repair strategies have struggled to keep pace with the variety of ways data can go missing. Physical approaches, such as correcting beam blockage using dual-polarization measurements, address only specific obstruction problems. Optical flow methods estimate the motion field of radar echoes and extrapolate observed information into the gap, which works reasonably well when storms evolve smoothly and the missing region is spatially continuous. But during rapidly developing convection, when echoes can form, dissipate, or intensify abruptly, the assumption that echoes simply translate breaks down and reconstruction accuracy collapses. Meanwhile, deep learning approaches such as CNN-BiConvLSTM and DSA-UNet have shown strong modeling power, yet most of them depend on a missing-data mask, an explicit map of which pixels are absent. In real operational mosaics, especially when the spatial distribution of gaps is highly uncertain, such masks are difficult to obtain, which limits practical deployment.</p>
<p>The new method, called BiConvLSTM-UNet, sidesteps that requirement entirely. It treats missing echo restoration as a sequence reconstruction task: given a corrupted sequence of radar frames, the model learns the inherent spatiotemporal patterns of echo evolution and reconstructs the complete sequence, implicitly identifying and filling the gaps. The architecture follows a U-shaped encoder-decoder design. The encoder stacks depthwise separable convolution modules that progressively reduce spatial resolution while expanding the channel dimension, building hierarchical spatial features. At the bottleneck, a bidirectional ConvLSTM module models the feature sequence in both forward and backward time directions, so the model can draw on context from frames both before and after a gap. The decoder then restores spatial resolution using PixelShuffle upsampling combined with skip connections that merge high-level semantics with fine spatial detail, and a ReLU1 activation confines outputs to the physically valid range between zero and one.</p>
<p>Training the model required careful data engineering. The team used composite reflectivity mosaics from March to September of 2022 and 2023, covering two full rainy seasons over Hubei Province and totaling 102,720 frames. They focused on meteorologically significant echoes between 10 and 75 dBZ, clipped values outside that range, and discarded frames with almost no echo coverage. After testing sequence lengths of 5, 15, and 30 frames, they settled on 15 frames per sample as the best balance between informational richness and computational efficiency, yielding a final dataset of 60,615 frames split into training, validation, and test sets. To simulate realistic data loss, each frame was divided into five sectors, and random subregions of 100 by 100, 150 by 150, or 200 by 200 grid points were blanked out, with one to three frames per sample randomly affected so the model would learn to handle both isolated and consecutive gaps.</p>
<p>The loss function proved equally important. The researchers combined a weighted L1 loss, in which pixel weights increase with echo intensity so the model pays more attention to strong reflectivity, with a multi-scale structural similarity loss that preserves both global morphology and local texture. Ablation experiments confirmed that each component matters. Removing the structural similarity term dropped the peak signal-to-noise ratio from 25.608 to 24.445 and the structural similarity index from 0.765 to 0.741, and for intense echoes above 40 dBZ the critical success index fell to just 0.078. Replacing the weighted L1 term with a standard unweighted version reduced the critical success index at the 40 dBZ threshold from 0.192 to 0.163 and the probability of detection from 0.359 to 0.215, showing that intensity weighting is what enables the model to recover fierce convective cores rather than smoothing them away.</p>
<p>Head-to-head comparisons against ConvLSTM, DSA-UNet, and classical optical flow showed BiConvLSTM-UNet achieving the best overall performance across peak signal-to-noise ratio, structural similarity, critical success index, and probability of detection, while preserving sharper boundaries and more spatially coherent echo structures. ConvLSTM produced robust but locally unsmooth reconstructions, DSA-UNet recovered echoes proactively but tended toward over-smoothed edges and a high false alarm ratio, and optical flow, faithful to echo boundaries when storms moved steadily, failed badly under rapidly evolving convection. One honest caveat emerged: BiConvLSTM-UNet showed a moderately elevated false alarm ratio, reflecting a recall-oriented strategy that prefers to reconstruct plausible echoes rather than miss them, particularly in low-reflectivity or ambiguous areas.</p>
<p>Because the model reconstructs the entire sequence rather than only the gaps, it risks subtly altering pixels that were actually observed. To guard against this, the team added a post-processing step that keeps original observed values everywhere the input data exists and substitutes model predictions only where the input reads zero. The effect was dramatic: in non-missing regions, the root mean square error fell from 0.704 to 0.266 and the mean absolute error from 0.260 to 0.013, while a newly introduced boundary continuity metric showed that the smoothness of intensity transitions across gap edges was preserved rather than degraded. Generalization tests extended the picture further. When the missing length grew from one to three frames to four random frames, performance declined only marginally, but four consecutive missing frames caused consistent deterioration across all metrics, indicating that uninterrupted loss of temporal context fundamentally challenges the model, which then adopts a conservative bias favoring omission over false activation.</p>
<p>The authors also tested the model outside its comfort zone, evaluating it on data from the non-rainy season months of 2023. It maintained moderate performance, with a peak signal-to-noise ratio of 25.838 and a structural similarity of 0.749, but detection of strong echoes weakened as thresholds rose, a consequence of the distributional shift between sparse, dispersed cold-season echoes and the dense rainy-season data it was trained on. The team proposes expanding training data across seasons, incorporating complementary variables such as vertical wind shear and convective available potential energy, and refining loss formulations with spatially adaptive weighting to address these limits. The study does not claim mask-free reconstruction is inherently superior to mask-based methods; rather, it demonstrates for the first time that high-fidelity radar echo restoration is feasible using only the intrinsic spatiotemporal structure of the echo sequences themselves. For operational meteorology, where missing masks are rarely available and every minute of a storm&#8217;s evolution counts, that is a meaningful step toward weather radar mosaics that heal themselves in real time.</p>
<p><strong>Subject of Research:</strong> Deep learning-based restoration of missing radar echoes in weather radar mosaic data</p>
<p><strong>Article Title:</strong> Research on deep learning-based missing echo restoration method for weather radar mosaic data</p>
<p><strong>Article References:</strong> Guo, H., Du, M., Fan, X., Wu, C., Lai, A., &amp; Ma, H. (2026). Research on deep learning-based missing echo restoration method for weather radar mosaic data. <em>Atmospheric Measurement Techniques, 19</em>(19), 6341-6356. <a href="https://doi.org/10.5194/amt-19-6341-2026" rel="noopener noreferrer">https://doi.org/10.5194/amt-19-6341-2026</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.5194/amt-19-6341-2026" rel="noopener noreferrer">10.5194/amt-19-6341-2026</a></p>
<p><strong>Keywords:</strong> weather radar, radar mosaic, deep learning, BiConvLSTM-UNet, data reconstruction, nowcasting, severe convective weather, ConvLSTM, U-Net, optical flow, precipitation estimation, Atmospheric Measurement Techniques</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247850</post-id>	</item>
		<item>
		<title>Hybrid CNN-Transformer AI Spots Anomalies in Surveillance Video With Record Accuracy</title>
		<link>https://scienmag.com/hybrid-cnn-transformer-ai-spots-anomalies-in-surveillance-video-with-record-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 01:36:27 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI for surveillance footage analysis]]></category>
		<category><![CDATA[autoencoder]]></category>
		<category><![CDATA[autoencoder-based video anomaly detection]]></category>
		<category><![CDATA[CNN]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[Convolutional Neural Networks and Vision Transformers]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[detecting unusual activities in security footage]]></category>
		<category><![CDATA[Farneback]]></category>
		<category><![CDATA[hybrid CNN-transformer architecture]]></category>
		<category><![CDATA[innovative computer vision techniques]]></category>
		<category><![CDATA[large-scale surveillance data analysis]]></category>
		<category><![CDATA[memory module]]></category>
		<category><![CDATA[optical flow]]></category>
		<category><![CDATA[real-time surveillance monitoring]]></category>
		<category><![CDATA[record accuracy in anomaly detection]]></category>
		<category><![CDATA[surveillance]]></category>
		<category><![CDATA[surveillance video anomaly detection]]></category>
		<category><![CDATA[UCSD Ped2]]></category>
		<category><![CDATA[unsupervised anomaly detection in videos]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<category><![CDATA[unsupervised learning for security applications]]></category>
		<category><![CDATA[video anomaly detection]]></category>
		<category><![CDATA[vision transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=239870</guid>

					<description><![CDATA[Researchers have developed a memory-augmented CNN-vision transformer autoencoder that uses Farneback optical flow to detect anomalies in surveillance video with up to 98.97 percent AUC on standard benchmarks.]]></description>
										<content:encoded><![CDATA[<p>A world with roughly one billion surveillance cameras is no longer a distant prediction, and the sheer volume of footage those cameras generate has long outstripped the capacity of human operators to watch it all. Against that backdrop, a team of researchers in India has unveiled a new artificial intelligence framework that learns what normal behavior looks like in a video scene and flags anything that deviates from it, without ever being shown a single labeled example of an anomaly. The work, published in Cluster Computing, combines two of the most influential architectures in modern computer vision, convolutional neural networks and vision transformers, into a single autoencoder that reconstructs the everyday rhythm of a scene and stumbles visibly when something unusual happens.</p>
<p>The research, led by Vandana Pathak of Graphic Era Deemed to be University in Dehradun, together with Manoj Diwakar, Neeraj Kumar Pandey, Sanjay Roka and Prabhishek Singh, addresses a stubborn problem in video surveillance: anomalies are rare, varied and almost impossible to enumerate in advance. A cyclist cutting through a pedestrian zone, a person sprinting the wrong way down a crowded corridor, or an abandoned bag left on a plaza all look nothing alike, which makes supervised classification impractical. The dominant alternative, unsupervised anomaly detection, flips the task on its head. Instead of teaching a model to recognize trouble, it teaches the model to recognize normality so thoroughly that trouble becomes conspicuous by its absence.</p>
<p>The centerpiece of the new framework is a memory-augmented CNN-ConvViT autoencoder. Autoencoders compress an input into a compact latent representation and then attempt to reconstruct it, and when trained exclusively on normal footage they reconstruct familiar patterns well and unfamiliar ones poorly. The reconstruction error then serves as an anomaly score. The difficulty, well documented in prior work, is that a sufficiently powerful autoencoder learns to generalize too much, reconstructing even the anomalies it was never trained on, which erases the very signal the system depends on. Memory modules were introduced to counteract this by storing prototypical features of normal behavior and forcing the encoder to express every input as a combination of those stored prototypes, limiting the network&#8217;s ability to improvise.</p>
<p>What distinguishes the new approach is the design of its memory component, called the Temporal-Aware Prototype Memory Module, or TAPMM. Rather than treating memory as a static dictionary of spatial appearances, TAPMM explicitly learns prototypes of normal spatio-temporal behavior, capturing how scenes evolve over time rather than merely how they look in a single frame. This temporal awareness matters because many surveillance anomalies are defined by motion rather than appearance: a person walking calmly through a parking lot is unremarkable, while the same person running through it may warrant attention. By encoding temporal structure into the memory itself, the framework narrows the gap between what the model stores and what constitutes an anomaly in practice.</p>
<p>The architecture also tackles a complementary weakness. Convolutional neural networks excel at extracting local spatial detail, such as edges, textures and the shapes of individual objects, but their receptive fields limit their grasp of long-range relationships across a scene. Vision transformers, by contrast, use self-attention to model global context, letting every part of an image attend to every other part, but they can be less efficient at capturing fine-grained local structure. The proposed framework fuses convolutional blocks with Conv-ViT blocks, a hybrid in which convolutional operations and transformer attention are integrated so that local spatial details and global contextual dependencies are modeled jointly. This kind of hybridization reflects a broader trend in computer vision, following influential studies asking whether vision transformers see like convolutional networks and demonstrating that the two paradigms have complementary strengths.</p>
<p>Motion information enters the pipeline through a second clever design choice. Alongside the raw video frames, the researchers feed the network an additional input channel computed with Farneback optical flow, a dense optical flow algorithm that estimates the motion vector of every pixel between consecutive frames. Optical flow gives the model an explicit, pixel-level description of how everything in the scene is moving, independent of how it looks. By combining appearance information from the raw frames with motion information from the flow channel, the network can learn both appearance-based variations, such as an unexpected object, and motion-based variations, such as movement in the wrong direction or at an abnormal speed. This dual-stream strategy echoes earlier two-flow architectures but integrates the motion signal directly into a transformer-augmented reconstruction framework.</p>
<p>At inference time, the system scores each frame using the reconstruction error, complemented by the peak signal-to-noise ratio, a standard image quality metric that drops when a reconstruction deviates sharply from the original. Frames whose reconstructions are poor, or whose PSNR falls below the pattern established by normal footage, are flagged as anomalous. The evaluation was conducted on four widely used benchmarks: UCSD Ped1 and UCSD Ped2, which capture pedestrian walkways with cyclists, skaters and occasional vehicles intruding into the frame; the Avenue dataset, filmed in a campus entrance hall with loitering, throwing and running; and ShanghaiTech, a large and challenging collection of thirteen scenes with diverse camera angles and crowd conditions.</p>
<p>The results are striking. On UCSD Ped1 the framework achieved an area under the receiver operating characteristic curve of 98.97 percent with an equal error rate of 4.09 percent, meaning the point where its false alarm rate and miss rate cross sits below five percent. On UCSD Ped2 it reached 95.63 percent AUC with a 4.34 percent EER, and on Avenue it recorded 93.53 percent AUC with a 12.23 percent EER. On the hardest benchmark, ShanghaiTech, it attained 89.14 percent AUC with a 16.49 percent EER. The consistent performance across datasets with very different scene dynamics, lighting conditions and anomaly types suggests the hybrid design generalizes rather than overfitting to a single environment, and the reported equal error rates on the UCSD benchmarks place the method among the strongest reconstruction-based approaches described in the literature.</p>
<p>The implications extend well beyond academic benchmarks. Security operators, transit authorities and smart-city planners all face the same economics: footage is cheap, attention is expensive. Systems that can reliably narrow a human operator&#8217;s focus to the small fraction of video that actually deserves scrutiny could change how surveillance is staffed and reviewed. Because the framework is unsupervised, it also sidesteps the privacy and labeling burdens of supervised training, since it requires only examples of ordinary activity, which every camera already records in abundance. The authors note that the datasets used in the study are available from the first author on reasonable request, and the work was carried out without dedicated funding.</p>
<p>Challenges remain before such systems can be trusted in the wild. Real deployments must cope with camera shake, weather, gradual shifts in what counts as normal as seasons and crowds change, and the ethical questions that accompany any technology capable of deciding, autonomously, what counts as suspicious behavior. The equal error rates on the more crowded and heterogeneous benchmarks, while strong, still leave room for false alarms that could erode operator trust. Yet the trajectory is clear. By marrying the local precision of convolutions, the global reasoning of transformers, a memory that remembers how normal scenes unfold in time, and an explicit motion signal from optical flow, this work offers a blueprint for surveillance AI that watches quietly, learns the rhythm of a place, and speaks up only when the rhythm breaks.</p>
<p><strong>Subject of Research:</strong> Unsupervised spatio-temporal video anomaly detection using a memory-augmented CNN-ViT autoencoder with Farneback optical flow</p>
<p><strong>Article Title:</strong> Spatio-temporal video anomaly detection via CNN-ViT autoencoder and farneback optical flow</p>
<p><strong>Article References:</strong> Pathak, V., Diwakar, M., Pandey, N. K., Roka, S., &amp; Singh, P. (2026). Spatio-temporal video anomaly detection via CNN-ViT autoencoder and farneback optical flow. <em>Cluster Computing, 29</em>(13), Article 773. <a href="https://doi.org/10.1007/s10586-026-06598-5" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06598-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06598-5" rel="noopener noreferrer">10.1007/s10586-026-06598-5</a></p>
<p><strong>Keywords:</strong> video anomaly detection, surveillance, autoencoder, vision transformer, CNN, optical flow, Farneback, memory module, unsupervised learning, computer vision, deep learning, UCSD Ped2</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">239870</post-id>	</item>
		<item>
		<title>Scientists Teach Computers to Watch Rock Avalanches Move, Frame by Frame</title>
		<link>https://scienmag.com/scientists-teach-computers-to-watch-rock-avalanches-move-frame-by-frame/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:40:21 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[advanced methods in natural hazard detection]]></category>
		<category><![CDATA[China]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision for geological hazards]]></category>
		<category><![CDATA[deformation characterization]]></category>
		<category><![CDATA[disaster mitigation]]></category>
		<category><![CDATA[early warning]]></category>
		<category><![CDATA[frame-by-frame avalanche movement tracking]]></category>
		<category><![CDATA[innovative techniques in geohazard monitoring]]></category>
		<category><![CDATA[integrating video and seismic signals]]></category>
		<category><![CDATA[landslide dynamics]]></category>
		<category><![CDATA[large-scale rock avalanche case study China]]></category>
		<category><![CDATA[multi-source data fusion for landslide analysis]]></category>
		<category><![CDATA[natural hazards]]></category>
		<category><![CDATA[optical flow]]></category>
		<category><![CDATA[real-time rock avalanche observation]]></category>
		<category><![CDATA[rock avalanche]]></category>
		<category><![CDATA[rock avalanche monitoring]]></category>
		<category><![CDATA[seismic data analysis for landslides]]></category>
		<category><![CDATA[seismic signal analysis]]></category>
		<category><![CDATA[technological advances in slope failure analysis]]></category>
		<category><![CDATA[temporal alignment]]></category>
		<category><![CDATA[video-based motion analysis]]></category>
		<category><![CDATA[visual and seismic data integration in geology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216509</guid>

					<description><![CDATA[A joint framework combining optical-flow computer vision and seismic signal analysis reconstructs the moment-by-moment motion and deformation of large rock avalanches.]]></description>
										<content:encoded><![CDATA[<p>When a mountainside collapses, the transformation is terrifyingly fast. Millions of tons of rock detach from a slope, shatter, and race downhill at speeds that can exceed dozens of meters per second, burying villages and reshaping valleys in minutes. For decades, scientists studying these catastrophic rock avalanches have had to reconstruct their dynamics largely after the fact, relying on the geometry of deposits, eyewitness accounts, and the faint vibrations the events leave in the ground. Now, a research team in China has proposed a way to actually watch an avalanche move, using computer vision to extract frame-by-frame motion from video footage and seismic signals to keep the analysis honest.</p>
<p>The new study, published in the journal Natural Hazards by Yiwei Liu, Aiguo Xing, and colleagues at Shanghai Jiao Tong University, the China Institute of Geo-Environment Monitoring, and Hunan University, presents a joint analytical framework that fuses two very different data streams: ordinary video recordings and seismic observations. The work is a case-oriented, exploratory attempt to integrate visual and seismic information for rock avalanche analysis, and it demonstrates its approach on a representative large rock avalanche event in Southwest China. The goal is deceptively simple to state but technically demanding to achieve: to obtain velocity distributions and motion descriptions at any moment during an avalanche, something neither data source can deliver alone.</p>
<p>The core of the visual analysis is a technique called optical flow, a cornerstone of computer vision with a history stretching back to foundational work in the early 1980s. Optical flow algorithms estimate the apparent motion of pixels or features between successive frames of a video, producing a dense field of motion vectors that describes how every part of the image is shifting. In the context of a rock avalanche, those vectors translate directly into kinematic information: which parts of the flowing mass are moving fastest, where the front of the avalanche is advancing, and how the body of the flow stretches, compresses, and deforms as it descends. The researchers implemented their pipeline in Python, building on widely used open-source libraries including OpenCV, NumPy, and pandas, and they have released the code and data openly through GitHub and the Zenodo repository.</p>
<p>Video alone, however, has a fundamental weakness. A camera records only what is in its field of view, and the motion it captures is projected onto a two-dimensional image plane. Without independent confirmation, it can be difficult to know precisely when events in the video occurred in absolute time, or whether apparent motion reflects real ground movement or artifacts such as camera shake, changing illumination, or dust obscuring the scene. This is where the seismic half of the framework becomes essential. Large rock avalanches generate ground vibrations that are recorded by seismic instruments at considerable distances, and seismologists have developed increasingly sophisticated methods to invert those signals for landslide characteristics such as location, timing, momentum, and frictional behavior.</p>
<p>In the joint framework, seismic signal analysis serves as an independent reference for temporal alignment and validation. By matching prominent features in the seismic record, such as the abrupt onset of high-frequency energy when the mass impacts the valley floor, with corresponding moments in the video, the researchers can anchor the visual timeline to a precise, instrumentally measured clock. The seismic data then act as a cross-check: if the optical-flow analysis suggests a surge in motion at a particular moment, the seismic record should show a corresponding change in energy release. Agreement between the two independent streams increases confidence in the reconstructed dynamics; disagreement flags moments that deserve closer scrutiny.</p>
<p>The framework is organized into a four-stage architecture that integrates data acquisition, preprocessing, joint analysis, and result visualization. In the acquisition stage, video footage is gathered from monitoring systems and public sources, alongside seismic records covering the event window. Preprocessing addresses the practical realities of real-world imagery: stabilizing frames, correcting for lighting variation, and preparing the seismic traces through filtering and decomposition techniques. The joint analysis stage runs the optical-flow motion extraction and the seismic interpretation in parallel, aligning them in time. Finally, the visualization stage renders the results as velocity distributions and motion descriptions that can be examined at any moment during the avalanche, turning raw footage and seismograms into quantitative, interpretable pictures of the disaster in progress.</p>
<p>Applied to a large rock avalanche in Southwest China, the framework demonstrated its practicality. The results showed that the combined approach provides complementary insights to those obtained from seismic signal analysis alone. Seismic inversion can estimate bulk properties of a landslide, but it inherently averages the behavior of the entire moving mass. Video-based optical flow, by contrast, resolves spatial patterns: it can reveal that the front of the flow accelerates while the trailing mass decelerates, or that deformation concentrates along particular zones within the avalanche body. Merging the two perspectives yields a richer and more defensible reconstruction than either could offer independently.</p>
<p>The significance of this capability extends well beyond academic curiosity. Rock avalanches routinely cause severe casualties and economic losses, and improving the characterization of event-scale dynamic behavior and precursory deformation is considered essential for effective disaster mitigation and early warning. Landslides are among the deadliest geological hazards worldwide, and China, with its steep terrain, active tectonics, and dense population in mountainous regions, has experienced repeated catastrophic events, including the long-runout Pusa rock avalanche of 2017 in Guizhou Province and the massive Chamoli rock and ice avalanche in the Indian Himalaya in 2021, which killed hundreds. Understanding exactly how these masses accelerate, fragment, and spread is critical for designing warning systems, mapping hazard zones, and validating the numerical models engineers use to predict runout.</p>
<p>What makes the new framework particularly timely is the explosion of available video. Monitoring cameras now watch many hazardous slopes continuously, and when disasters strike, surveillance cameras, smartphones, and public platforms often capture the event from multiple angles. Until recently, this footage was mostly used qualitatively, as dramatic evidence of what happened. The authors argue that video-based analysis has become a valuable complementary data source for investigating landslide dynamics, and their framework provides a systematic, reproducible way to convert pixels into physics. Optical-flow methods have already proven useful in related domains, from measuring glacier surface velocities in satellite imagery to tracking slow slope creep in high-alpine settings, and this study extends that lineage to the most violent end of the landslide spectrum.</p>
<p>The researchers are careful to frame the work as a case-oriented and exploratory attempt rather than a finished operational system. Challenges remain, including the sensitivity of optical flow to dust, poor lighting, and low frame rates, and the difficulty of converting image-plane motion into true three-dimensional ground velocities without accurate camera calibration and terrain models. Nevertheless, the study highlights the potential of video-based computer vision techniques to enhance the interpretation of landslide dynamics and to support future monitoring and early-warning research. If the approach matures, the aftermath of the next catastrophic slope failure may be documented not just by trembling seismometers and stunned witnesses, but by algorithms that watched every frame, measured every surge, and helped explain one of nature&#8217;s most violent phenomena in unprecedented detail.</p>
<p><strong>Subject of Research:</strong> Integration of computer vision and seismic signal analysis to characterize the dynamics of large rock avalanches</p>
<p><strong>Article Title:</strong> Characterizing dynamic motion and deformation during large rock avalanches using a joint framework integrating computer vision and seismic signal analysis</p>
<p><strong>Article References:</strong> Liu, Y., Xing, A., Wang, W., Wang, Q., Zhuang, Y., Bilal, M., &amp; Zhu, K. (2026). Characterizing dynamic motion and deformation during large rock avalanches using a joint framework integrating computer vision and seismic signal analysis. <em>Natural Hazards, 122</em>(18), Article 628. <a href="https://doi.org/10.1007/s11069-026-08349-6" rel="noopener noreferrer">https://doi.org/10.1007/s11069-026-08349-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11069-026-08349-6" rel="noopener noreferrer">10.1007/s11069-026-08349-6</a></p>
<p><strong>Keywords:</strong> rock avalanche, optical flow, computer vision, seismic signal analysis, landslide dynamics, deformation characterization, early warning, disaster mitigation, video-based motion analysis, Natural Hazards, China, temporal alignment</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216509</post-id>	</item>
		<item>
		<title>AI Learns to Spot Eye Contact Between Autistic Children and Clinicians During Free Play</title>
		<link>https://scienmag.com/ai-learns-to-spot-eye-contact-between-autistic-children-and-clinicians-during-free-play/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:52:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced AI frameworks for social behavior analysis]]></category>
		<category><![CDATA[AI tools for measuring eye contact in children]]></category>
		<category><![CDATA[AI-based eye contact detection in autism assessment]]></category>
		<category><![CDATA[autism spectrum disorder]]></category>
		<category><![CDATA[automated social engagement analysis through deep learning]]></category>
		<category><![CDATA[automatic coding of eye contact during free play]]></category>
		<category><![CDATA[behavioral coding]]></category>
		<category><![CDATA[clinical assessment]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision in autism diagnosis]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[enhancing clinical autism assessments with AI technology]]></category>
		<category><![CDATA[eye contact]]></category>
		<category><![CDATA[free play assessment]]></category>
		<category><![CDATA[Grad-CAM]]></category>
		<category><![CDATA[machine learning for autism social behavior markers]]></category>
		<category><![CDATA[multi-modal video analysis for autism research]]></category>
		<category><![CDATA[multi-view video]]></category>
		<category><![CDATA[multi-view video analysis for autism diagnostics]]></category>
		<category><![CDATA[optical flow]]></category>
		<category><![CDATA[real-time eye contact recognition in clinical settings]]></category>
		<category><![CDATA[ResNet-50]]></category>
		<category><![CDATA[social behavior recognition]]></category>
		<category><![CDATA[unobtrusive assessment of social interactions in autism]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204172</guid>

					<description><![CDATA[A multi-view deep learning framework can objectively recognize clinically defined mutual eye contact between children and assessors during naturalistic free play, achieving F1 scores above 0.90.]]></description>
										<content:encoded><![CDATA[<p>Mutual eye contact is one of the most powerful signals in human social life. It conveys attention, intimacy, trust, and emotional state, and its absence early in development has long been recognized as a clinically meaningful marker of atypical social development, most notably in autism spectrum disorder. Yet the tools clinicians use to measure it remain stubbornly low-tech: trained observers watch recorded sessions and manually code when eye contact begins and ends, a process that is slow, expensive, and vulnerable to human bias and error. Now, a team of researchers from the University of Sydney, the Chinese University of Hong Kong, Shanghai Jiao Tong University, and collaborating institutions has unveiled an artificial intelligence framework that can automatically detect clinically defined eye contact between children and assessors during real, unscripted free play sessions—potentially transforming how social engagement is quantified in clinical assessments.</p>
<p>The study, published in the journal Cognitive Computation, introduces a multi-modal, multi-view deep learning framework called the Multi-Modal Social Behavior Recognition (MSBR) system. Unlike eye-tracking glasses or constrained single-camera setups, the framework relies on four synchronized, wall-mounted high-definition cameras capturing complementary views of a clinical assessment room. Children—103 diagnosed with autism and 33 typically developing controls, all aged 3 to 12 years—completed semi-structured 60-minute social interaction sessions that included a five-minute free play task. During free play, children could choose toys, move around the room, and interact spontaneously with an assessor and, at times, their caregivers. Crucially, no wearable devices were attached to the children, avoiding the calibration drift, facial obstruction, and sensory discomfort that plague head-mounted eye trackers, particularly in children with sensory sensitivities.</p>
<p>Defining what counts as eye contact was a careful clinical exercise. For this study, an eye contact event was defined as the child directing visual attention toward the eyes of the co-located assessor while the assessor simultaneously maintained eye contact with the child. The researchers stress that this is a video-observed behavioral event, not a measurement of gaze vectors, ocular fixation, or gaze angle. Two trained coders with domain knowledge annotated the start and end times of each mutual eye contact event using the MATLAB Video Labeler, achieving moderate to high inter-rater reliability, with mean Cohen&#8217;s Kappa values of 0.72 in the typically developing group and 0.84 in the autism group. Ambiguous cases—such as distinguishing genuine mutual eye contact from general face-looking—were reviewed and resolved with lead clinical researchers according to strict coding criteria.</p>
<p>The scale of the annotated dataset reflects the rarity of the target behavior. Across 136 videos totaling 733 minutes of recording, mutual eye contact events were identified in 77 autistic children and 25 typically developing children, totaling just 567 seconds of cumulative eye contact. The framework&#8217;s developers trimmed the videos into clips labeled either &#8220;Eye Contact&#8221; or &#8220;Other,&#8221; yielding 1,726 eye contact clips and 2,435 non-eye-contact clips. Each clip was then decoded into two complementary visual representations: RGB frames capturing static appearance, and optical flow images capturing motion patterns between consecutive frames. This dual representation feeds the framework&#8217;s two branches—a spatial-domain branch processing single RGB frames and a temporal-domain branch processing sequences of optical flow images—both built on the ResNet-50 convolutional backbone within a Temporal Segment Network architecture.</p>
<p>Training the model on such imbalanced data required careful engineering. Because most of the free play time did not involve eye contact, the researchers used randomly sampled, fixed-length sequences from long clips, increasing exposure to positive eye contact events while preventing the model from being dominated by lengthy &#8220;Other&#8221; segments. Data augmentation techniques, including MultiScaleCrop for the RGB branch and RandomResizedCrop for the optical flow branch, were applied only to training data, with all images resized to 224 by 224 pixels and horizontally flipped with 50 percent probability. The model was trained in PyTorch on two NVIDIA GeForce RTX 2080Ti GPUs using Adam optimization with a cosine annealing learning rate scheduler, for up to 300 epochs with categorical cross-entropy loss.</p>
<p>A methodological hallmark of the study is its rigorous participant-independent evaluation. The researchers used five-fold stratified cross-validation with participant-level partitioning: within each fold, 60 percent of unique participants were allocated to training, 20 percent to validation, and 20 percent to held-out testing, ensuring that no participant&#8217;s data appeared in more than one partition. Autistic and typically developing children were stratified across partitions to maintain group representation. Hyperparameters were tuned exclusively on validation sets, and held-out test data were reserved for final evaluation, providing a robust measure of how the model performs on children it has never seen.</p>
<p>The results were striking. A fused single-view model achieved F1 scores of 0.85 for eye contact behavior and 0.90 for non-eye-contact behavior. But the multi-view fusion of predictions from all four cameras pushed performance to an F1 score of 0.92 for eye contact and 0.94 for &#8220;Other&#8221; behaviors, with a Top-1 accuracy of 0.94 for the best fused configuration. Paired fold-level comparisons confirmed that multi-view prediction significantly outperformed single-view prediction across all modalities and diagnostic groups, with mean paired F1 improvements of 0.065 for eye contact and 0.044 for non-eye-contact behaviors and very large fold-level effect sizes. Notably, no statistically significant performance differences emerged between the autism and typically developing groups, and the spatial-domain RGB models outperformed temporal optical flow models—an outcome the authors attribute to the transient, small-amplitude nature of eye contact movements, which makes motion features difficult to extract.</p>
<p>Interpretability received dedicated attention. Using Gradient-weighted Class Activation Mapping, or Grad-CAM, the team generated heatmaps showing which image regions drove the model&#8217;s classifications. In the deeper convolutional layers, activation concentrated around faces, upper bodies, and interaction-relevant areas such as toys being manipulated—precisely the visual information clinicians and trained coders rely on when distinguishing mutual eye contact from its absence. The authors caution, however, that these visualizations are qualitative and do not demonstrate that the model estimates gaze vectors directly; they indicate associations, not causal explanations of the model&#8217;s decision process.</p>
<p>The study&#8217;s limitations are candidly acknowledged. The framework was trained and evaluated within a specific four-camera clinical setup and a particular age range, and external validation across independent sites has not yet been performed. The system recognizes video-observed behavioral events rather than measuring physiological gaze, so future work should validate these clinically defined behaviors against eye-tracking or gaze-estimation methods. The substantial manual annotation burden remains a challenge, motivating future development of weakly supervised or action localization approaches. Privacy, secure data governance, and deployment feasibility also demand attention, though the models&#8217; relatively modest footprint—about 23.5 million parameters and 89.7 MiB of FP32 storage per modality—suggests that deployment on GPU-enabled workstations or edge-computing platforms may be feasible, and privacy-preserving frameworks using de-identified features offer a promising direction.</p>
<p>Even with these caveats, the implications are considerable. An objective, video-derived behavioral marker of mutual eye contact could relieve clinicians of hours of manual video coding, reduce subjective bias, and enable scalable, quantitative assessment of social engagement consistent with established clinical criteria. Such a tool could support more detailed behavioral characterization in autism research and inform future clinical decision-support systems, while preserving the ecological richness of spontaneous multi-person interaction—toy play, variable body orientation, free movement, and unscripted engagement—that constrained laboratory paradigms sacrifice. The researchers envision extending the framework to additional behavioral modalities, including gesture dynamics, body movement, speech features, and audio-visual interaction patterns, moving toward a comprehensive digital characterization of children&#8217;s social behavior. If validated across independent clinical sites and integrated with complementary gaze-measurement technologies, multi-view AI frameworks of this kind could become a routine component of developmental assessment, turning ordinary clinical video into rigorous, quantifiable evidence about how children connect with the people around them.</p>
<p><strong>Subject of Research:</strong> Automated multi-view deep learning recognition of clinically defined eye contact between children and assessors during free play autism assessments</p>
<p><strong>Article Title:</strong> Naturalistic Social Dyads Assessment In Free Play: A Multi-view Framework For Recognizing Co-located Eye Contact Between Children and Their Assessors</p>
<p><strong>Article References:</strong> Sun, C., Guastella, A. J., Ouyang, W., Zhou, L., Thapa, R., Thomas, E. E., Zhao, H., &amp; McEwan, A. (2026). Naturalistic Social Dyads Assessment In Free Play: A Multi-view Framework For Recognizing Co-located Eye Contact Between Children and Their Assessors. <em>Cognitive Computation, 18</em>(1), Article 110. <a href="https://doi.org/10.1007/s12559-026-10659-7" rel="noopener noreferrer">https://doi.org/10.1007/s12559-026-10659-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12559-026-10659-7" rel="noopener noreferrer">10.1007/s12559-026-10659-7</a></p>
<p><strong>Keywords:</strong> eye contact, autism spectrum disorder, deep learning, multi-view video, free play assessment, social behavior recognition, computer vision, Grad-CAM, behavioral coding, clinical assessment, optical flow, ResNet-50</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204172</post-id>	</item>
		<item>
		<title>Robots Learn to Track Moving Objects by Watching Human Contact</title>
		<link>https://scienmag.com/robots-learn-to-track-moving-objects-by-watching-human-contact/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:36:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in autonomous robot localization]]></category>
		<category><![CDATA[challenges of moving objects in robot mapping]]></category>
		<category><![CDATA[contact experience]]></category>
		<category><![CDATA[Dyna-SLAM]]></category>
		<category><![CDATA[dynamic environments]]></category>
		<category><![CDATA[dynamic SLAM systems for mobile robots]]></category>
		<category><![CDATA[epipolar constraints]]></category>
		<category><![CDATA[handling moving furniture and objects in robotic navigation]]></category>
		<category><![CDATA[human contact-based object tracking for robots]]></category>
		<category><![CDATA[human-driven cues for robotic environment understanding]]></category>
		<category><![CDATA[improving robot navigation accuracy amidst moving obstacles]]></category>
		<category><![CDATA[leveraging human-object interactions for robot perception]]></category>
		<category><![CDATA[localization techniques for robots in busy households]]></category>
		<category><![CDATA[movable objects]]></category>
		<category><![CDATA[optical flow]]></category>
		<category><![CDATA[ORB-SLAM2]]></category>
		<category><![CDATA[RGB-D camera]]></category>
		<category><![CDATA[robot perception in cluttered and dynamic settings]]></category>
		<category><![CDATA[robotic object tracking in dynamic environments]]></category>
		<category><![CDATA[robotics]]></category>
		<category><![CDATA[semantic segmentation]]></category>
		<category><![CDATA[SLAM]]></category>
		<category><![CDATA[TUM RGB-D dataset]]></category>
		<category><![CDATA[visual SLAM in moving scenes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203124</guid>

					<description><![CDATA[A new RGB-D SLAM system from researchers in China keeps indoor robots accurately localized in dynamic environments by using human contact experience to predict the motion states of movable objects such as books and cups.]]></description>
										<content:encoded><![CDATA[<p>Robots navigating a busy living room face a deceptively hard problem: the world refuses to stay still. A companion robot&#8217;s camera sees people walking past, chairs pulled across the floor, and books or cups carried from one table to another. Every one of those moving things can corrupt the map the robot is quietly building of its surroundings. A new study published in Autonomous Robots proposes a way for a robot to lean on a distinctly human clue—the fact that people tend to hold certain objects—to keep its localization steady even when the furniture is on the move.</p>
<p>The research, led by Jilin Zhang of the University of Jinan with colleagues from Shandong Normal University, the University of Jinan and Lunan Technician College, targets a weak spot in modern visual Simultaneous Localization and Mapping, or SLAM. Classical SLAM systems assume the world they observe is static. When people walk through the frame, feature points attached to them move for reasons that have nothing to do with camera motion, and the system&#8217;s estimate of its own trajectory drifts. Recent dynamic SLAM methods, such as Dyna-SLAM and SaD-SLAM, attack this by detecting humans and other obviously dynamic objects and discarding their pixels. But the authors point out a stubborn category of uncertainty: movable objects like books and cups. Most of the time a book sits still on a desk, so treating it as static background is reasonable—until someone picks it up and carries it across the room. A SLAM system that blindly trusts those features inherits the object&#8217;s motion as phantom camera motion.</p>
<p>The team&#8217;s answer is a dynamic SLAM system built around what they call human contact experience. Rather than hard-coding which objects are dynamic and which are static, the system learns from observation how often humans come into contact with particular categories of objects. Objects that are frequently held, such as cups and books, receive a prior state that makes the system suspicious of their apparent motion; objects that people rarely touch keep their static status. When a human and a movable object are in contact, the object&#8217;s features are treated as unreliable and excluded from pose estimation. In effect, the robot accumulates a form of common-sense knowledge about indoor life and uses it to decide which of the things it sees can be trusted as reference points.</p>
<p>Technically, the system weaves together three modules. The first is an adaptive frame selection strategy driven by semantic segmentation results. Instead of feeding every RGB-D frame into the computationally expensive segmentation pipeline, the system adaptively chooses which frames to process, reducing computational resource consumption while improving the quality of the prior information available to later stages. This matters for real robots, which must localize in real time on hardware far less powerful than a laboratory workstation. The second module refines the geometric analysis: by combining optical flow with epipolar constraints, the system determines the motion states of both humans and movable objects. Optical flow captures how pixels shift between consecutive frames, while the epipolar constraint describes how a static point in a rigid scene should move given the camera&#8217;s own motion. A point that violates the epipolar geometry is almost certainly moving independently of the camera—an elegant, geometry-based way to flag dynamic content without relying on semantics alone.</p>
<p>The third and conceptually novel piece is the contact experience module itself. Drawing on information from multiple consecutive frames, the module records the contact frequency between humans and movable objects and uses that history to update the prior state of objects in the indoor environment. An object seen repeatedly in human hands shifts its prior toward dynamic; an object that has never been touched retains a static prior. Because this knowledge is updated continuously, the system adapts to a particular environment over time rather than relying on a fixed, hand-tuned list of dynamic classes. The authors describe this as using human contact experience with movable objects to predict their true states—a statistical prior grounded in the everyday physics of how people interact with their belongings.</p>
<p>Everything rests on accurate camera trajectories, so the researchers evaluated their system on the TUM RGB-D benchmark, the standard dataset for testing RGB-D SLAM under dynamic conditions. The benchmark includes sequences in which people walk, sit and interact with objects while the camera moves through the scene—precisely the conditions that break static-world assumptions. The proposed method was compared against ORB-SLAM2, the widely used open-source baseline for monocular, stereo and RGB-D cameras, and against two representative dynamic-environment systems, Dyna-SLAM and SaD-SLAM.</p>
<p>The reported results show the new system operating stably in dynamic environments and, crucially, handling state changes of indoor movable objects more effectively than its predecessors. Where Dyna-SLAM and SaD-SLAM can mask out walking people, they have no principled mechanism for the cup that was static in frame one and moving in frame ten. By combining adaptive frame selection, flow-and-epipolar geometry and contact-frequency priors, the new method covers both ends of the problem: it ignores pixels belonging to independently moving entities and reclassifies movable objects the moment their behavior changes. The authors note that the adaptive frame selection also keeps the computational cost in check, which matters for indoor companion robots that must run continuously.</p>
<p>The implications reach beyond a cleaner trajectory estimate. Indoor companion robots are expected to interact naturally with humans, and that requires knowing not just where the robot is, but what in the room is trustworthy as a landmark. A robot that understands that a person carrying a mug makes the mug&#8217;s features unreliable, but that the mug becomes a valid landmark again once set down, gains a more realistic model of its environment. The contact experience framework is also a small but suggestive step toward robots that learn everyday physics from observation—the kind of implicit knowledge humans use constantly without noticing.</p>
<p>The work was supported in part by the National Natural Science Foundation of China, the Taishan Scholar Foundation of Shandong Province and the Outstanding Youth Foundation of Shandong Province. As robots move from factory floors into homes, offices and hospitals, the ability to localize reliably amid human activity will stop being a research curiosity and become a baseline requirement. This study suggests that some of the best clues for separating a stable world from a shifting one may come from simply paying attention to what people are holding.</p>
<p><strong>Subject of Research:</strong> A dynamic RGB-D SLAM method that uses human contact experience to determine the motion states of movable objects for robot localization in dynamic indoor environments.</p>
<p><strong>Article Title:</strong> A RGB-D SLAM method based on contact experience in dynamic environment</p>
<p><strong>Article References:</strong> Zhang, J., Huang, K., Geng, H., Song, C., &amp; Zhang, M. (2026). A RGB-D SLAM method based on contact experience in dynamic environment. <em>Autonomous Robots, 50</em>(4), Article 40. <a href="https://doi.org/10.1007/s10514-026-10268-1" rel="noopener noreferrer">https://doi.org/10.1007/s10514-026-10268-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10514-026-10268-1" rel="noopener noreferrer">10.1007/s10514-026-10268-1</a></p>
<p><strong>Keywords:</strong> SLAM, RGB-D camera, dynamic environments, robotics, semantic segmentation, optical flow, epipolar constraints, contact experience, movable objects, ORB-SLAM2, Dyna-SLAM, TUM RGB-D dataset</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203124</post-id>	</item>
	</channel>
</rss>
