<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>detecting doctored videos using space and time &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/detecting-doctored-videos-using-space-and-time/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 14:36:43 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>detecting doctored videos using space and time &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Framework Spots Doctored Videos by Reading Both Space and Time</title>
		<link>https://scienmag.com/new-ai-framework-spots-doctored-videos-by-reading-both-space-and-time/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 14:36:43 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based video forgery identification]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[autoencoder]]></category>
		<category><![CDATA[comprehensive video forgery detection methods]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning video manipulation detection]]></category>
		<category><![CDATA[detecting doctored videos using space and time]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[forensic analysis of manipulated videos]]></category>
		<category><![CDATA[frame splicing]]></category>
		<category><![CDATA[hybrid neural network for video analysis]]></category>
		<category><![CDATA[identifying visual artifacts in fake videos]]></category>
		<category><![CDATA[inpainting]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Neural Processing Letters]]></category>
		<category><![CDATA[ResNet-50]]></category>
		<category><![CDATA[spatio-temporal analysis]]></category>
		<category><![CDATA[spatio-temporal video forensics]]></category>
		<category><![CDATA[tamper detection in videos]]></category>
		<category><![CDATA[tamperNet deep learning architecture]]></category>
		<category><![CDATA[unified video tampering detection framework]]></category>
		<category><![CDATA[video forensics]]></category>
		<category><![CDATA[video forgery detection]]></category>
		<category><![CDATA[video tampering detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223294</guid>

					<description><![CDATA[Researchers have developed TamperNet, a hybrid spatio-temporal deep learning framework that detects frame duplication, deletion, cloning, splicing, and inpainting in videos with 93.3 percent accuracy on a benchmark dataset.]]></description>
										<content:encoded><![CDATA[<p>Video has never been easier to fake. Consumer editing suites can duplicate a frame, splice in footage from another clip, or use inpainting algorithms to erase a person from a scene entirely, and the compression pipelines of social media platforms then wash away many of the subtle artifacts that forensic investigators traditionally relied upon. A research team led by Muhammad Hasaan Mujtaba, Muhammad Imran, and Saad Irfan of SZABIST University in Islamabad, together with Usman Muhammad of Aalto University in Finland and Hussain Dawood of Amity University Dubai and partner institutions, has now introduced a deep learning framework called TamperNet that confronts this problem head-on. Published open access in Neural Processing Letters, the work describes a hybrid spatio-temporal architecture designed to catch the most common families of video manipulation, including frame duplication, frame deletion, cloning, splicing, and inpainting, in a single unified pipeline rather than through a collection of narrow, forgery-specific detectors.</p>
<p>The central insight behind TamperNet is that video tampering almost always leaves two kinds of fingerprints at once. The first kind lives in the spatial domain: a spliced object may have inconsistent lighting, unnatural edges, or texture statistics that do not match its surroundings. The second kind lives in the temporal domain: duplicated or deleted frames disturb the smooth flow of motion, and inpainted regions often fail to reproduce the plausible dynamics that a real camera would capture. Most existing detection approaches emphasize one of these cues at the expense of the other, which makes them brittle when an editor deliberately suppresses the dominant artifact type. TamperNet instead fuses spatial and temporal evidence throughout the network, so that a manipulation must evade both channels simultaneously to slip past the detector.</p>
<p>Technically, the framework is built from several cooperating modules. A ResNet-50 backbone, a convolutional network originally developed for image recognition and widely reused as a feature extractor, serves as the spatio-temporal descriptor, converting raw frames into rich intermediate representations that encode edges, textures, and object structure. Because ResNet-50 processes frames individually, the authors pair it with a temporal representation module based on a long short-term memory network, or LSTM, a recurrent architecture specifically designed to capture dependencies over sequences. The LSTM reads the stream of frame-level features and learns what normal temporal evolution looks like, so that abrupt discontinuities, repeated segments, or missing intervals register as anomalies in the learned sequence representation.</p>
<p>Fusion is handled by concatenation, a straightforward but effective strategy in which the spatial feature vectors and the temporal feature vectors are joined end to end before being passed to the decision stages. This design choice matters because it preserves the full information content of both branches rather than compressing them prematurely into a single embedding. On top of the fused representation, TamperNet applies a spatio-temporal anomaly scoring strategy that combines two complementary error signals. The first is a temporal prediction error, which measures how badly the model&#8217;s learned dynamics fail to anticipate the actual next frames; manipulations that break continuity produce spikes in this error. The second is an autoencoder reconstruction error, in which a network trained to compress and reconstruct normal video content produces unusually large residuals when asked to reconstruct manipulated regions it has never effectively learned to model.</p>
<p>The combination of these two error signals gives the framework a degree of redundancy that single-cue detectors lack. A skilled forger might smooth out temporal discontinuities by re-encoding the video, degrading the usefulness of prediction error alone, but the altered content would still tend to reconstruct poorly, and vice versa. By scoring each candidate video against both criteria and fusing the results with the concatenated deep features, TamperNet can flag manipulations that span both the inter-frame domain, where the relationships between successive frames are disturbed, and the intra-frame domain, where the internal consistency of a single frame has been violated. This dual coverage is what allows one model to address such a heterogeneous list of forgery types, from crude frame duplication to sophisticated content-aware inpainting.</p>
<p>The team evaluated the framework on the Video Forgery Dataset, a benchmark abbreviated VFD that contains both authentic and manipulated video clips. The reported results are strong on the classification metrics that matter most for a binary detection task. TamperNet achieved an accuracy of 93.3 percent, meaning it correctly labeled nearly nineteen out of twenty videos. Precision, the fraction of flagged videos that were genuinely tampered, came in at 91.4 percent, while recall, the fraction of tampered videos that were successfully caught, reached 91.3 percent, indicating that the model&#8217;s errors are balanced rather than skewed toward false alarms or missed forgeries. The F1-score, the harmonic mean of precision and recall that penalizes imbalance between them, was 89.9 percent.</p>
<p>One metric, however, deserves careful attention. The area under the receiver operating characteristic curve, or ROC-AUC, which summarizes detection performance across all possible decision thresholds, was approximately 0.75. An ROC-AUC of 1.0 represents perfect ranking of tampered above authentic videos, while 0.5 represents chance. A value of 0.75 indicates discrimination clearly better than random but well short of the near-perfect separation implied by the headline accuracy figure, and the authors are candid about this gap. They note that the moderate ROC-AUC, together with the relatively small size of the tested benchmark, means that TamperNet&#8217;s demonstrated performance should be understood as evidence of feasibility on the tested set rather than proof of readiness for the messy conditions of real-world forensic deployment.</p>
<p>That candor is one of the more notable features of the paper. The authors explicitly state that further cross-dataset testing and compression-specific testing are required to assess how well the framework generalizes beyond the VFD benchmark. This matters because videos in the wild arrive heavily compressed by platform-specific codecs, resized, re-encoded multiple times, and captured by sensors with wildly different noise characteristics, and each of these transformations can mimic or mask the very artifacts a detector depends on. A model that excels on a curated benchmark can lose substantial accuracy when confronted with aggressive compression, and the research team&#8217;s decision to flag this limitation rather than bury it reflects a maturing attitude within the video forensics community, where benchmark overfitting has repeatedly inflated published claims.</p>
<p>The research itself is a genuinely international collaboration, with authors affiliated with SZABIST University&#8217;s Department of Robotics and Artificial Intelligence in Islamabad, the Department of Computer Science at Aalto University in Espoo, Finland, Amity University Dubai in the United Arab Emirates, Western Caspian University in Baku, Azerbaijan, and Jadara University in Irbid, Jordan. Open access funding was provided by Aalto University, and the corresponding author is Usman Muhammad. The article was received on 26 March 2026, accepted on 2 September 2026, and published on 1 October 2026 under a Creative Commons Attribution 4.0 license, meaning the full technical details, architecture descriptions, and experimental protocols are freely available to any researcher who wishes to build on or stress-test the work.</p>
<p>The road ahead, as the authors sketch it, involves three main directions. First, validation on larger benchmark sets is needed to establish that the reported accuracy holds when training and evaluation data are more diverse and abundant. Second, the framework must be hardened against the full range of compression and acquisition scenarios that real forensic casework involves, from low-bitrate smartphone uploads to professionally encoded broadcast footage. Third, and perhaps most ambitiously, the team plans to extend TamperNet from detection to localization, so that instead of merely issuing a verdict that a video has been manipulated, the system would highlight precisely which frames and which pixel regions were altered. Such spatially and temporally grounded output would transform the tool from an alarm into an evidentiary instrument, giving investigators, courts, and platform trust-and-safety teams a concrete map of where a video&#8217;s fabrications begin and end. As synthetic and manipulated media continue to erode the default assumption that footage can be trusted, frameworks like TamperNet mark an incremental but meaningful step toward restoring that trust, provided their limitations are tested as rigorously as their strengths are celebrated.</p>
<p><strong>Subject of Research:</strong> Spatio-temporal deep learning for detecting video manipulations such as frame duplication, splicing, and inpainting</p>
<p><strong>Article Title:</strong> TamperNet: A Spatio-Temporal Deep Learning Framework for Video Manipulation Detection</p>
<p><strong>Article References:</strong> Mujtaba, M. H., Imran, M., Irfan, S., Muhammad, U., &amp; Dawood, H. (2026). TamperNet: A Spatio-Temporal Deep Learning Framework for Video Manipulation Detection. <em>Neural Processing Letters</em>. <a href="https://doi.org/10.1007/s11063-026-11885-8" rel="noopener noreferrer">https://doi.org/10.1007/s11063-026-11885-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11063-026-11885-8" rel="noopener noreferrer">10.1007/s11063-026-11885-8</a></p>
<p><strong>Keywords:</strong> video tampering detection, deep learning, spatio-temporal analysis, LSTM, autoencoder, ResNet-50, video forensics, frame splicing, inpainting, anomaly detection, Neural Processing Letters, digital forensics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223294</post-id>	</item>
	</channel>
</rss>
