<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>digital forensics &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/digital-forensics/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 23:09:22 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>digital forensics &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Shadows Betray Fakes: Wedge-Based Analysis Exposes Doctored Images</title>
		<link>https://scienmag.com/shadows-betray-fakes-wedge-based-analysis-exposes-doctored-images/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 23:09:22 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced techniques for detecting image splicing]]></category>
		<category><![CDATA[composite image forensics]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[digital media manipulation detection]]></category>
		<category><![CDATA[DSO-1 dataset]]></category>
		<category><![CDATA[illumination direction]]></category>
		<category><![CDATA[image authentication]]></category>
		<category><![CDATA[image forgery detection]]></category>
		<category><![CDATA[image forgery detection using shadows]]></category>
		<category><![CDATA[lighting consistency]]></category>
		<category><![CDATA[multimedia forensics methods]]></category>
		<category><![CDATA[multimedia security]]></category>
		<category><![CDATA[optical principles in image verification]]></category>
		<category><![CDATA[photo manipulation]]></category>
		<category><![CDATA[physical signatures in image forensics]]></category>
		<category><![CDATA[physics-based image authenticity verification]]></category>
		<category><![CDATA[physics-based methods]]></category>
		<category><![CDATA[shadow analysis]]></category>
		<category><![CDATA[shadow analysis in digital forensics]]></category>
		<category><![CDATA[shadow geometry analysis]]></category>
		<category><![CDATA[shadow inconsistencies in doctored photos]]></category>
		<category><![CDATA[wedge-based analysis]]></category>
		<category><![CDATA[wedge-based shadow detection technique]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=250305</guid>

					<description><![CDATA[Researchers in Mumbai have developed a wedge-based shadow analysis technique that detects image forgeries by testing whether the shadows cast by scene objects converge on a single consistent illumination direction.]]></description>
										<content:encoded><![CDATA[<p>Every photograph carries an invisible witness to its own history. When light strikes objects in a scene, it casts shadows that obey the strict geometry of optics, and those shadows record precisely where the illuminating source must have been. If someone splices an object from one photograph into another, the shadows in the composite rarely agree with one another. A new study published in Multimedia Tools and Applications by Divya Surve and Anant Nimkar of the Sardar Patel Institute of Technology in Mumbai exploits exactly this physical signature, introducing a wedge-based shadow analysis technique designed to expose image forgeries even in the messy, real-world scenarios that have long defeated earlier forensic approaches.</p>
<p>The problem the researchers set out to address is one of the most pressing in modern digital media. Forged images are now a critical issue in the transmission of multimedia information, and detection strategies have traditionally split into two broad families. Statistical pixel-based models look for mathematical fingerprints left by editing operations, such as anomalies in compression artifacts, sensor noise patterns, or resampling traces. Physics-based strategies, by contrast, depend on physical indications like illumination direction, shading, reflection, and shadows. The second family has a powerful advantage: a forger can remove statistical traces with enough care, but it is far harder to fabricate a shadow geometry that is fully consistent with the lighting of a scene the object never actually occupied.</p>
<p>Yet physics-based methods have carried a well-known weakness. Standard gradient-based illumination assessment, which estimates the direction of light from the shading gradients on object surfaces, runs into serious limitations and challenges once shadows are present in the image. The difficulty becomes especially severe when objects cast shadows onto one another. In a cluttered scene, a person&#8217;s shadow may fall across a wall, a car, and another person simultaneously, and these overlapping shadows lead gradient-based estimators to produce varying or confusing outcomes. The illumination directions inferred from different objects can appear inconsistent even in a completely genuine photograph, generating false alarms that undermine the credibility of the entire forensic analysis.</p>
<p>The new technique tackles this failure mode head-on with a geometric construction the authors call wedge-based shadow analysis. The method begins by selecting key shadow points in the image, the salient locations where shadows attach to or extend from the objects that formed them. Angular wedges are then drawn between these shadow points and their probable forming objects. Each wedge is a cone of possible directions in the image plane, capturing the geometric relationship between an object, the shadow it casts, and the position of the light source that must connect them. Because a single light source must lie within every wedge derived from every consistent object-shadow pair in the scene, the wedges collectively constrain where the illumination can be.</p>
<p>The decision rule that follows is elegantly simple. Linear constraints are applied using these wedges, and the presence of an intersection between their standard limits determines the illumination direction. If the wedges all share a common region, the lighting implied by every shadow in the image agrees, and the image is judged consistent with a single physical light source. Consistency across the wedges therefore represents a genuine image, whereas inconsistency, the failure of the wedges to converge on a shared illumination direction, demonstrates a forgery sample. In other words, a spliced object whose shadow was copied from a differently lit scene will generate a wedge that points somewhere no other wedge in the image points, and the intersection test catches the contradiction.</p>
<p>This approach directly addresses the scenario that breaks gradient-based methods. When objects cast shadows onto each other, the wedges provide a way to reason about the collective geometry of the scene rather than relying on local shading estimates that overlapping shadows contaminate. The angular construction tolerates the ambiguity inherent in any single shadow, which spans a range of possible light positions, while exploiting the fact that a forger&#8217;s mistakes accumulate across multiple shadows. A genuine scene with many interlocking shadows still yields wedges that overlap; a composite does not, because the attacker would need to re-render every shadow in the image under a single coherent lighting model to pass the test.</p>
<p>The experimental evaluation used the DSO-1 dataset, a publicly available benchmark containing 44 images, both indoor and outdoor, all featuring shadows. The dataset includes both genuine photographs and manipulated ones, making it a standard proving ground for shadow-based forensics. On this benchmark, the wedge-based model achieved a detection accuracy of 58.73 percent for outdoor images and 67.58 percent for indoor images. The indoor advantage is intuitive: interior scenes tend to have more controlled, single-source lighting, which makes the wedge intersections cleaner and the consistency test more discriminating. Outdoor scenes, with diffuse skylight and multiple environmental light contributions, present a harder geometric puzzle.</p>
<p>Perhaps the most striking result concerns a specific and practically important subset of the data. For genuine outdoor images featuring shadows among objects, the scenes where shadows overlap and objects shade one another, the method reached a potential performance of 82 percent. That figure matters because these overlapping-shadow scenes are precisely the cases where the technique was designed to outperform gradient-based illumination assessment, which tends to produce confusing or contradictory estimates under the same conditions. The result suggests that the wedge formulation converts the very complication that defeats earlier methods, mutual shadow casting, into a source of additional geometric constraints that strengthen the authenticity verdict.</p>
<p>The study situates itself within a rich lineage of physics-based forensics research. Earlier work established that inconsistencies in lighting expose digital forgeries, that shading and shadows can be analyzed jointly to detect manipulation, and that reflection inconsistencies offer a complementary signal. Gradient-based illumination description was developed specifically for forgery detection, and other researchers have pursued optimized three-dimensional lighting environment estimation, linear constraints based on shading and shadows, Lambert model analysis with shadows, and shadow consistency checks using color-space features such as HSV. The wedge technique extends this tradition by targeting the multi-object, mutually shadowing scenes that constrained many of its predecessors, and the authors position it as a viable solution for handling the challenges of gradient-based illumination assessment in difficult real-world scenarios.</p>
<p>The broader significance of the work lies in the ongoing arms race between image manipulation and image verification. As generative tools make convincing composites trivially easy to produce, forensic science increasingly depends on cues that are expensive to fake. Shadows are among the most demanding of these cues, because getting them right requires not just artistic skill but a physically accurate understanding of the scene&#8217;s lighting geometry. A method that reads the angular relationships between objects and their shadows, and that remains robust when those shadows interlock, adds a meaningful layer of defense. The authors&#8217; findings, drawn from a modest but well-established benchmark, indicate that wedge-based shadow analysis can serve as a practical component in the forensic toolkit, complementing statistical detectors and metadata analysis. For editors, journalists, courts, and platforms confronting a daily flood of questionable imagery, the message is a compelling one: the light in a photograph always tells the truth, and now there is a more reliable way to make the shadows testify.</p>
<p><strong>Subject of Research:</strong> Physics-based image forgery detection using wedge-based analysis of object shadows and illumination consistency</p>
<p><strong>Article Title:</strong> Illuminating authenticity using deciphering image forgery through scene object shadow analysis</p>
<p><strong>Article References:</strong> Surve, D., &amp; Nimkar, A. (2026). Illuminating authenticity using deciphering image forgery through scene object shadow analysis. <em>Multimedia Tools and Applications, 85</em>(10), Article 801. <a href="https://doi.org/10.1007/s11042-026-21956-6" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21956-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21956-6" rel="noopener noreferrer">10.1007/s11042-026-21956-6</a></p>
<p><strong>Keywords:</strong> image forgery detection, shadow analysis, digital forensics, illumination direction, physics-based methods, wedge-based analysis, DSO-1 dataset, photo manipulation, computer vision, image authentication, multimedia security, lighting consistency</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">250305</post-id>	</item>
		<item>
		<title>Courts in Cameroon Still Try Cybercriminals With Outdated Penal Code</title>
		<link>https://scienmag.com/courts-in-cameroon-still-try-cybercriminals-with-outdated-penal-code/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 12:18:29 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[adoption and implementation of cyberlaw in Cameroon]]></category>
		<category><![CDATA[ANTIC]]></category>
		<category><![CDATA[Budapest Convention]]></category>
		<category><![CDATA[Cameroon]]></category>
		<category><![CDATA[Cameroon cybercrime law enforcement]]></category>
		<category><![CDATA[cyber law]]></category>
		<category><![CDATA[cybercrime]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[effectiveness of cybersecurity legislation in Cameroon]]></category>
		<category><![CDATA[electronic evidence]]></category>
		<category><![CDATA[gaps in Cameroon]]></category>
		<category><![CDATA[impact of outdated laws on cybercrime sentencing]]></category>
		<category><![CDATA[judicial handling of cybercrimes in Cameroon]]></category>
		<category><![CDATA[judicial reliance on general penal code for digital offenses]]></category>
		<category><![CDATA[judiciary]]></category>
		<category><![CDATA[legal certainty in Cameroon cybercrime prosecutions]]></category>
		<category><![CDATA[legal challenges in prosecuting digital crimes in Cameroon]]></category>
		<category><![CDATA[legal reform]]></category>
		<category><![CDATA[mobile money fraud]]></category>
		<category><![CDATA[outdated penal code in Cameroon]]></category>
		<category><![CDATA[Penal Code]]></category>
		<category><![CDATA[prosecution]]></category>
		<category><![CDATA[role of Budapest Convention in Cameroon's cybercrime laws]]></category>
		<category><![CDATA[technological sophistication of cybercrimes in Cameroon]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=235002</guid>

					<description><![CDATA[A new study finds that despite Cameroon's 2010 cybercrime law, its courts continue to prosecute digital offences under the ordinary Penal Code, producing lighter sentences and legal uncertainty.]]></description>
										<content:encoded><![CDATA[<p>Cameroon&#8217;s courts are prosecuting some of the most technologically sophisticated crimes of the digital age with legal tools designed for a world of physical theft and paper documents. A new doctrinal study of cybercrime prosecution in the Central African nation reveals a striking paradox: fifteen years after the country adopted a dedicated cybercrime law, its judges and prosecutors continue to rely overwhelmingly on the ordinary Penal Code, undermining legal certainty and handing down sentences far lighter than legislators intended. The research, published in Discover Global Society, analyzes judicial decisions from trial courts across both of Cameroon&#8217;s legal traditions and finds a judiciary that, in most cases, simply sidesteps the specialized statute.</p>
<p>The study, authored by legal scholar Boris Awa, examines how Law No. 2010/012 of 21 December 2010 on Cybersecurity and Cybercrime, known as the Cyber-law, has fared in practice. Adopted to fill a legal vacuum created by rapid technological change, the law criminalizes a comprehensive range of digital offences modeled on the Budapest Convention on Cybercrime. These include unauthorized interception of communication networks, disruption of information systems, fraudulent access to networks, misuse of personal data, publication of intimate images, and child pornography. The statute was meant to guarantee access to justice in an era when the internet had become the primary vector for criminal activity. Yet the study&#8217;s central finding is that, from a litany of judicial decisions examined, the courts continue to apply ordinary penal code offences when determining cybercrime charges.</p>
<p>The scale of the problem is considerable. Crypto scams alone account for more than 54 percent of cybercrime in Cameroon, with damages estimated at over two million US dollars in 2023, and no comprehensive data exists on the full scale of the various cybercrimes committed in the country. INTERPOL&#8217;s African Cyberthreat Assessment Report identifies online scams, digital extortion, business email compromise, botnets and ransomware as the driving forces behind cyberthreats across the continent. Against this backdrop, the gap between the specialized law on the books and the law applied in courtrooms has real consequences for victims, defendants and the credibility of the justice system itself.</p>
<p>The research rests on a qualitative, doctrinal methodology grounded in primary and secondary sources, including legislation, constitutional provisions and case law obtained from three trial courts: the Court of First Instance Bonanjo in the economic capital Douala, the Court of First Instance Bamenda in the North West Region, and the Buea Court of First Instance in the South West Region. This selection allowed a comparative view across Cameroon&#8217;s French-speaking civil law and English-speaking common law traditions, a dual legal heritage enshrined in Article 68 of the Constitution. The author acknowledges that judgments from seven other regions could not be examined due to financial and logistical constraints, so the analysis does not purport to be exhaustive, but the selected courts carry significant caseloads and offer a credible window into contemporary judicial practice.</p>
<p>The case studies are revealing. In the 2024 Tresor case, an employee of a company partnering with MTN Cameroon used his professional credentials to reset a customer&#8217;s mobile money PIN and withdrew 3,200,000 CFA francs, roughly 3,100 US dollars, from the complainant&#8217;s account. The court convicted him of simple theft under section 318 of the Penal Code and imposed a five-month prison term, suspended for one year. The study argues that the proper charge was hacking under section 65 of the Cyber-law, which punishes unauthorized access to an electronic communication network. Similarly, in the 2021 Joly case, an administrator of a WhatsApp group used for currency trading was convicted of theft by false pretence after failing to deliver currency worth 399,000 CFA francs, receiving a three-year sentence suspended for three years, when the study contends the sanction should have been sourced from the electronic commerce legislation. In a third case, an advance fee fraud conducted entirely by telephone and Western Union transfers was again prosecuted as ordinary false pretence rather than under the Cyber-law.</p>
<p>The consequences of this reliance on ordinary law are structural, not merely cosmetic. Under the Penal Code, judges enjoy broad discretion to grant mitigating circumstances and suspended sentences, and in practice most cybercriminals either serve no imprisonment or receive markedly lighter penalties. The Cyber-law, by contrast, expressly forbids suspended sentences for the offences it defines, and the study argues that mitigating circumstances cannot apply because they are not formally provided for in the statute, in line with the principle of legality. Moreover, key Penal Code concepts translate poorly to the digital realm: aggravated theft requires material elements such as force, weapons, breaking in, climbing in or the use of a false key or motor vehicle, none of which corresponds to the mechanics of a cyber intrusion. The result is that offences which the Cyber-law treats as serious crimes, some punishable by more than twenty years of imprisonment, are reduced to misdemeanours or simple offences under ordinary law, a leniency the study warns may itself encourage the culture of committing cybercrime.</p>
<p>Evidence is the other half of the problem. The Criminal Procedure Code of 2005 admits electronic proof only tacitly and narrowly, through provisions on wiretapping and electronic listening devices that the Court of Appeal in the Tchoffo Jonas case interpreted as investigative tools for capturing live communications, not as a framework for authenticating stored digital data. The Cyber-law contains its own procedural rules, empowering criminal investigation officers and sworn officials of the National Agency for Information and Communication Technologies, ANTIC, to conduct searches and seizures of electronic data. But in practice, most judicial police officers lack the training to collect electronic evidence themselves, forcing prosecutors to depend on ANTIC for expert analysis. Because ANTIC is centralized in Yaoundé while Cameroon has 282 courts of first instance and 39 high courts, requests pile up, trials stall, and suspects are released on bail or acquitted for insufficient evidence, as occurred in several of the cases studied.</p>
<p>The study also documents a second, smaller current of jurisprudence in which courts do engage with the specialized law. The Buea Court of First Instance held in one case that a cybercrime is consummated so long as the medium used is electronic communication, whether a website, a text message or a phone call, and advised that doubtful emails be examined by a forensic IT expert. In a 2019 Bamenda ruling, Justice Bikelle Theresia set out the elements the prosecution must establish under section 65 of the Cyber-law for unauthorized access, and the accused was convicted. Yet such decisions remain scarce. The study points to a paucity of cybercrime cases overall, attributing it to trivialization by investigators and prosecutors, settlements between parties, and even alleged connivance with so-called scammers, whose foreign victims rarely file complaints or follow up because of the cost involved.</p>
<p>Why does the judiciary keep reaching for the old codes? The study identifies a web of causes: a shortage of local legal scholarship, with only three publications having analyzed the role of judicial actors in cybercrime prosecution; a weak law reporting system, with a single reporter covering English-speaking Cameroon and none for the Francophone regions; and, above all, inadequate training. Most Cameroonian law faculties offer information and communication technology law only as a final-year module, if at all, and criminal law courses are often taught without any reference to cyber-enabled offending. The Chief Justice of the Supreme Court, Daniel Mekobe Sone, acknowledged the problem in his 2023 address at the solemn reopening of the court, decrying the non-application of cybercrime provisions, the peculiarities of electronic evidence and general public ignorance. The problem is not unique to Cameroon, with similar dynamics documented in Nigeria and South Africa, though the study notes that Cameroon&#8217;s over-reliance on its ordinary Penal Code sets it apart.</p>
<p>The prescriptions are concrete. The study recommends that ANTIC, in partnership with the Cameroon Bar Association, mount massive sensitization and training campaigns for lawyers, investigators, prosecutors, court registrars, bailiffs and judges, covering the qualification of offences and the handling of electronic evidence. It calls for revising the Cyber-law itself to address emerging threats such as deepfakes, artificial intelligence and cryptocurrencies, and for removing the criminalization of libel, which comparable jurisdictions have struck down as an affront to free speech. It urges Cameroon to ratify the Budapest Convention and the 2024 United Nations Convention against Cybercrime to ease cross-border cooperation, to sanction judicial corruption, to accelerate law reporting, and to stimulate scholarly output. Without these reforms, the study warns, the gap between the digital crimes being committed and the analogue justice being dispensed will only widen, eroding trust in the electronic networks on which the country&#8217;s future depends.</p>
<p><strong>Subject of Research:</strong> Judicial prosecution of cybercrime and the application of specialized cyber law versus ordinary penal law in Cameroon</p>
<p><strong>Article Title:</strong> Rethinking cybercrime prosecution in Cameroon</p>
<p><strong>Article References:</strong> Awa, B. (2026). Rethinking cybercrime prosecution in Cameroon. <em>Discover Global Society, 4</em>(1), Article 215. <a href="https://doi.org/10.1007/s44282-026-00582-5" rel="noopener noreferrer">https://doi.org/10.1007/s44282-026-00582-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44282-026-00582-5" rel="noopener noreferrer">10.1007/s44282-026-00582-5</a></p>
<p><strong>Keywords:</strong> cybercrime, Cameroon, cyber law, prosecution, electronic evidence, Penal Code, judiciary, ANTIC, mobile money fraud, Budapest Convention, legal reform, digital forensics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">235002</post-id>	</item>
		<item>
		<title>Hybrid AI Reaches Near-Perfect Accuracy in the Hunt for Doctored Images</title>
		<link>https://scienmag.com/hybrid-ai-reaches-near-perfect-accuracy-in-the-hunt-for-doctored-images/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 11:57:48 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in computer vision for image forgery]]></category>
		<category><![CDATA[benchmark datasets]]></category>
		<category><![CDATA[CASIA dataset]]></category>
		<category><![CDATA[challenges and solutions in fake image identification]]></category>
		<category><![CDATA[CNN]]></category>
		<category><![CDATA[convolutional neural networks in digital forensics]]></category>
		<category><![CDATA[copy-move forgery]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Deep learning for image forgery detection]]></category>
		<category><![CDATA[deepfake detection]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[digital watermarking and tampering detection methods]]></category>
		<category><![CDATA[forensic techniques for verifying image authenticity]]></category>
		<category><![CDATA[future]]></category>
		<category><![CDATA[GaN]]></category>
		<category><![CDATA[high-accuracy AI models for fake image recognition]]></category>
		<category><![CDATA[hybrid AI systems for doctored image identification]]></category>
		<category><![CDATA[image forgery detection]]></category>
		<category><![CDATA[metaheuristic optimization]]></category>
		<category><![CDATA[metaheuristic optimization algorithms in image manipulation detection]]></category>
		<category><![CDATA[particle swarm optimization]]></category>
		<category><![CDATA[state-of-the-art AI accuracy in doctored image detection]]></category>
		<category><![CDATA[systematic review of AI-driven image verification]]></category>
		<category><![CDATA[transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234938</guid>

					<description><![CDATA[A new systematic review finds that hybrid deep learning and metaheuristic optimization frameworks achieve 98 to 100 percent accuracy in detecting copy-move image forgeries, while identifying robustness and generalization as the field's key remaining challenges.]]></description>
										<content:encoded><![CDATA[<p>A doctored photograph can travel around the world before anyone notices the seam. In an era when generative tools can fabricate convincing scenes in seconds, the science of proving that an image has been manipulated has become one of the most consequential frontiers in computer vision. A new systematic review published in Multimedia Tools and Applications by Archana M R of GITAM University, Dayanand J of Guru Dev Nanak Engineering College, and R N Kulkarni of Ballari Institute of Technology and Management takes stock of that frontier, mapping how deep learning frameworks and metaheuristic optimization algorithms are being combined to catch forgers at their own game. The survey&#8217;s headline finding is striking: hybrid systems that fuse convolutional neural networks with optimization techniques are achieving detection accuracies of 98 to 100 percent on standard benchmarks, outperforming their non-hybrid counterparts by a wide margin.</p>
<p>The review begins by grounding readers in the two great families of forensic techniques. Active methods embed protective information into an image before it ever leaves the camera or the editing suite. Digital watermarking is the classic example: a fragile watermark is woven into pixel data so that any subsequent alteration disturbs the embedded pattern and reveals the tampering. Watermarking has proven valuable in domains such as medical imaging, where authentication of every pixel is a legal and clinical necessity, and researchers have even developed lossless deep-learning-based watermarking schemes for that purpose. The weakness of active approaches, however, is structural. They require cooperation at the point of capture, and the overwhelming majority of images circulating online carry no such protection.</p>
<p>Passive, or blind, forensics asks the opposite question: what traces does manipulation itself leave behind? Every act of forgery, however careful, disturbs the statistical fabric of an image. Resampling introduces periodic correlations that betray scaling or rotation, a phenomenon first formalized by Popescu and Farid in 2005. JPEG compression leaves ghosts and block-grained artifacts that can be analyzed to localize edits. Camera characteristics such as chromatic aberration, sensor noise patterns, and camera response functions create internal consistency checks, because a spliced-in region often arrives with the fingerprint of a different device. Even the physics of light plays a role: inconsistent shading and shadows across a composite scene can expose a fabrication, an approach pioneered by Farid&#8217;s group and later extended into perceptual metrics for detecting retouching.</p>
<p>Within passive forensics, the review devotes particular attention to copy-move forgery, the deceptively simple act of copying a region of an image and pasting it elsewhere to conceal or duplicate content. Because the pasted patch comes from the same image, it inherits the same noise, lighting, and compression history, defeating many classic consistency checks. Early defenses relied on exhaustive block matching, notably Fridrich and colleagues&#8217; 2003 block-based method, and on transform-domain tricks such as combining the discrete cosine transform with singular value decomposition. The field then shifted toward keypoint detectors: SIFT, SURF, ORB, KAZE, and Harris corners each offered distinctive, scale-tolerant descriptors that could be matched across an image to reveal duplicated regions, with clustering schemes such as J-Linkage and mDBSCAN used to separate genuine matches from coincidental ones.</p>
<p>Keypoint methods, however, struggle with small or extremely smooth tampered regions that yield too few distinctive features, and they can be defeated by geometric warping of the copied patch. This is where deep learning entered the picture. Convolutional neural networks learn manipulation traces directly from data, and architectures such as TamperNet, BusterNet, and Noiseprint demonstrated that networks could localize tampered areas, distinguish source from target regions in copy-move forgeries, and even recover camera-model fingerprints. Dual-branch CNNs, coarse-to-refined networks paired with adaptive clustering, and self-consistency learning, in which a network learns that spliced regions fail to cohere with their surroundings, all pushed accuracy upward. The review catalogs these architectures and compares their training regimes and performance across the field&#8217;s standard benchmarks: CASIA, MICC, GRIP, and CoMoFoD.</p>
<p>The most distinctive contribution of the survey is its focus on optimization, a layer of the pipeline that often goes unexamined in mainstream coverage. Deep networks are riddled with hyperparameters, from learning rates to layer configurations, and feature-matching pipelines involve thresholds and clustering parameters that are notoriously difficult to set by hand. Metaheuristic algorithms, nature-inspired search procedures such as particle swarm optimization, ant colony optimization, genetic algorithms, the multi-verse optimizer, the Archimedes optimization algorithm, and even more exotic entrants like the political optimizer and battle royale optimization, offer a way to tune these choices automatically. The review documents systems that pair CNNs with a hybrid spotted hyena and grasshopper optimizer, superpixel clustering refined by emperor penguin optimization, and golden ball optimization applied to medical image forensics, among many others.</p>
<p>The payoff for this hybridization is quantifiable. Across the studies surveyed, CNN-plus-optimization frameworks consistently outperform non-hybrid methods, reaching the 98 to 100 percent accuracy band on benchmark datasets. Optimization also improves robustness, helping detectors cope with post-forgery operations such as compression, noise addition, and blurring that would otherwise degrade matching. Yet the authors are candid about the limits. Robustness to geometrical distortions, rotations, scalings, and reflections remains a persistent weak point, since many detectors assume that copied regions appear in near-original form. Computational inefficiency is a second concern: some pipelines demand processing power incompatible with real-time or mobile deployment, a gap that lightweight architectures such as MobileNets and U-Net-style segmentation backbones are only beginning to close. Poor generalization across diverse forgery sources is the third weakness, as models trained on one dataset or manipulation type often falter on another.</p>
<p>Looking forward, the review identifies three research directions that could reshape the field. Transformer-based architectures, following the ViT and Swin Transformer lineage, are already being adapted for forgery localization, with recent work on multi-exit vision transformers suggesting gains in both robustness and efficiency. GAN-based detection is a double-edged frontier: generative adversarial networks can both synthesize deepfakes and be trained to recognize the artificial fingerprints that generators inadvertently leave behind, a question explored since researchers first asked whether GANs leave detectable traces. And real-time forensic systems, capable of operating across varied imaging conditions on everything from smartphones to newsroom verification desks, represent the practical destination toward which much of this research is pointing. Semi-supervised localization methods and meta-learning approaches that let models adapt quickly to new manipulation types also feature among the promising avenues.</p>
<p>What makes this survey timely is the asymmetry it highlights between the tools of deception and the tools of detection. Generative models improve at a pace set by commercial competition, while forensic systems must generalize to manipulations they have never seen, often under compression and resizing imposed by social media platforms. The authors&#8217; synthesis suggests that neither raw deep learning nor clever optimization alone will win that race; it is the disciplined combination of the two, with metaheuristic search tuning learned models and learned models supplying the representational power that handcrafted features lack, that currently defines the state of the art. For journalists, courts, and platforms that increasingly depend on image provenance, the message is both reassuring and cautionary: near-perfect detection is now achievable in the laboratory, but translating that performance into robust, fast, universally trustworthy verification remains the field&#8217;s unfinished work.</p>
<p><strong>Subject of Research:</strong> Image forgery detection using deep learning frameworks and metaheuristic optimization techniques</p>
<p><strong>Article Title:</strong> Advancing image forgery detection: An investigation into optimization techniques and deep learning frameworks</p>
<p><strong>Article References:</strong> R, A. M., J, D., &amp; Kulkarni, R. N. (2026). Advancing image forgery detection: An investigation into optimization techniques and deep learning frameworks. <em>Multimedia Tools and Applications, 85</em>(9), Article 746. <a href="https://doi.org/10.1007/s11042-026-21916-0" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21916-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21916-0" rel="noopener noreferrer">10.1007/s11042-026-21916-0</a></p>
<p><strong>Keywords:</strong> image forgery detection, copy-move forgery, deep learning, CNN, metaheuristic optimization, particle swarm optimization, digital forensics, transformers, GAN, deepfake detection, CASIA dataset, benchmark datasets</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234938</post-id>	</item>
		<item>
		<title>New AI Framework Spots Doctored Videos by Reading Both Space and Time</title>
		<link>https://scienmag.com/new-ai-framework-spots-doctored-videos-by-reading-both-space-and-time/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 14:36:43 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based video forgery identification]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[autoencoder]]></category>
		<category><![CDATA[comprehensive video forgery detection methods]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning video manipulation detection]]></category>
		<category><![CDATA[detecting doctored videos using space and time]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[forensic analysis of manipulated videos]]></category>
		<category><![CDATA[frame splicing]]></category>
		<category><![CDATA[hybrid neural network for video analysis]]></category>
		<category><![CDATA[identifying visual artifacts in fake videos]]></category>
		<category><![CDATA[inpainting]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Neural Processing Letters]]></category>
		<category><![CDATA[ResNet-50]]></category>
		<category><![CDATA[spatio-temporal analysis]]></category>
		<category><![CDATA[spatio-temporal video forensics]]></category>
		<category><![CDATA[tamper detection in videos]]></category>
		<category><![CDATA[tamperNet deep learning architecture]]></category>
		<category><![CDATA[unified video tampering detection framework]]></category>
		<category><![CDATA[video forensics]]></category>
		<category><![CDATA[video forgery detection]]></category>
		<category><![CDATA[video tampering detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223294</guid>

					<description><![CDATA[Researchers have developed TamperNet, a hybrid spatio-temporal deep learning framework that detects frame duplication, deletion, cloning, splicing, and inpainting in videos with 93.3 percent accuracy on a benchmark dataset.]]></description>
										<content:encoded><![CDATA[<p>Video has never been easier to fake. Consumer editing suites can duplicate a frame, splice in footage from another clip, or use inpainting algorithms to erase a person from a scene entirely, and the compression pipelines of social media platforms then wash away many of the subtle artifacts that forensic investigators traditionally relied upon. A research team led by Muhammad Hasaan Mujtaba, Muhammad Imran, and Saad Irfan of SZABIST University in Islamabad, together with Usman Muhammad of Aalto University in Finland and Hussain Dawood of Amity University Dubai and partner institutions, has now introduced a deep learning framework called TamperNet that confronts this problem head-on. Published open access in Neural Processing Letters, the work describes a hybrid spatio-temporal architecture designed to catch the most common families of video manipulation, including frame duplication, frame deletion, cloning, splicing, and inpainting, in a single unified pipeline rather than through a collection of narrow, forgery-specific detectors.</p>
<p>The central insight behind TamperNet is that video tampering almost always leaves two kinds of fingerprints at once. The first kind lives in the spatial domain: a spliced object may have inconsistent lighting, unnatural edges, or texture statistics that do not match its surroundings. The second kind lives in the temporal domain: duplicated or deleted frames disturb the smooth flow of motion, and inpainted regions often fail to reproduce the plausible dynamics that a real camera would capture. Most existing detection approaches emphasize one of these cues at the expense of the other, which makes them brittle when an editor deliberately suppresses the dominant artifact type. TamperNet instead fuses spatial and temporal evidence throughout the network, so that a manipulation must evade both channels simultaneously to slip past the detector.</p>
<p>Technically, the framework is built from several cooperating modules. A ResNet-50 backbone, a convolutional network originally developed for image recognition and widely reused as a feature extractor, serves as the spatio-temporal descriptor, converting raw frames into rich intermediate representations that encode edges, textures, and object structure. Because ResNet-50 processes frames individually, the authors pair it with a temporal representation module based on a long short-term memory network, or LSTM, a recurrent architecture specifically designed to capture dependencies over sequences. The LSTM reads the stream of frame-level features and learns what normal temporal evolution looks like, so that abrupt discontinuities, repeated segments, or missing intervals register as anomalies in the learned sequence representation.</p>
<p>Fusion is handled by concatenation, a straightforward but effective strategy in which the spatial feature vectors and the temporal feature vectors are joined end to end before being passed to the decision stages. This design choice matters because it preserves the full information content of both branches rather than compressing them prematurely into a single embedding. On top of the fused representation, TamperNet applies a spatio-temporal anomaly scoring strategy that combines two complementary error signals. The first is a temporal prediction error, which measures how badly the model&#8217;s learned dynamics fail to anticipate the actual next frames; manipulations that break continuity produce spikes in this error. The second is an autoencoder reconstruction error, in which a network trained to compress and reconstruct normal video content produces unusually large residuals when asked to reconstruct manipulated regions it has never effectively learned to model.</p>
<p>The combination of these two error signals gives the framework a degree of redundancy that single-cue detectors lack. A skilled forger might smooth out temporal discontinuities by re-encoding the video, degrading the usefulness of prediction error alone, but the altered content would still tend to reconstruct poorly, and vice versa. By scoring each candidate video against both criteria and fusing the results with the concatenated deep features, TamperNet can flag manipulations that span both the inter-frame domain, where the relationships between successive frames are disturbed, and the intra-frame domain, where the internal consistency of a single frame has been violated. This dual coverage is what allows one model to address such a heterogeneous list of forgery types, from crude frame duplication to sophisticated content-aware inpainting.</p>
<p>The team evaluated the framework on the Video Forgery Dataset, a benchmark abbreviated VFD that contains both authentic and manipulated video clips. The reported results are strong on the classification metrics that matter most for a binary detection task. TamperNet achieved an accuracy of 93.3 percent, meaning it correctly labeled nearly nineteen out of twenty videos. Precision, the fraction of flagged videos that were genuinely tampered, came in at 91.4 percent, while recall, the fraction of tampered videos that were successfully caught, reached 91.3 percent, indicating that the model&#8217;s errors are balanced rather than skewed toward false alarms or missed forgeries. The F1-score, the harmonic mean of precision and recall that penalizes imbalance between them, was 89.9 percent.</p>
<p>One metric, however, deserves careful attention. The area under the receiver operating characteristic curve, or ROC-AUC, which summarizes detection performance across all possible decision thresholds, was approximately 0.75. An ROC-AUC of 1.0 represents perfect ranking of tampered above authentic videos, while 0.5 represents chance. A value of 0.75 indicates discrimination clearly better than random but well short of the near-perfect separation implied by the headline accuracy figure, and the authors are candid about this gap. They note that the moderate ROC-AUC, together with the relatively small size of the tested benchmark, means that TamperNet&#8217;s demonstrated performance should be understood as evidence of feasibility on the tested set rather than proof of readiness for the messy conditions of real-world forensic deployment.</p>
<p>That candor is one of the more notable features of the paper. The authors explicitly state that further cross-dataset testing and compression-specific testing are required to assess how well the framework generalizes beyond the VFD benchmark. This matters because videos in the wild arrive heavily compressed by platform-specific codecs, resized, re-encoded multiple times, and captured by sensors with wildly different noise characteristics, and each of these transformations can mimic or mask the very artifacts a detector depends on. A model that excels on a curated benchmark can lose substantial accuracy when confronted with aggressive compression, and the research team&#8217;s decision to flag this limitation rather than bury it reflects a maturing attitude within the video forensics community, where benchmark overfitting has repeatedly inflated published claims.</p>
<p>The research itself is a genuinely international collaboration, with authors affiliated with SZABIST University&#8217;s Department of Robotics and Artificial Intelligence in Islamabad, the Department of Computer Science at Aalto University in Espoo, Finland, Amity University Dubai in the United Arab Emirates, Western Caspian University in Baku, Azerbaijan, and Jadara University in Irbid, Jordan. Open access funding was provided by Aalto University, and the corresponding author is Usman Muhammad. The article was received on 26 March 2026, accepted on 2 September 2026, and published on 1 October 2026 under a Creative Commons Attribution 4.0 license, meaning the full technical details, architecture descriptions, and experimental protocols are freely available to any researcher who wishes to build on or stress-test the work.</p>
<p>The road ahead, as the authors sketch it, involves three main directions. First, validation on larger benchmark sets is needed to establish that the reported accuracy holds when training and evaluation data are more diverse and abundant. Second, the framework must be hardened against the full range of compression and acquisition scenarios that real forensic casework involves, from low-bitrate smartphone uploads to professionally encoded broadcast footage. Third, and perhaps most ambitiously, the team plans to extend TamperNet from detection to localization, so that instead of merely issuing a verdict that a video has been manipulated, the system would highlight precisely which frames and which pixel regions were altered. Such spatially and temporally grounded output would transform the tool from an alarm into an evidentiary instrument, giving investigators, courts, and platform trust-and-safety teams a concrete map of where a video&#8217;s fabrications begin and end. As synthetic and manipulated media continue to erode the default assumption that footage can be trusted, frameworks like TamperNet mark an incremental but meaningful step toward restoring that trust, provided their limitations are tested as rigorously as their strengths are celebrated.</p>
<p><strong>Subject of Research:</strong> Spatio-temporal deep learning for detecting video manipulations such as frame duplication, splicing, and inpainting</p>
<p><strong>Article Title:</strong> TamperNet: A Spatio-Temporal Deep Learning Framework for Video Manipulation Detection</p>
<p><strong>Article References:</strong> Mujtaba, M. H., Imran, M., Irfan, S., Muhammad, U., &amp; Dawood, H. (2026). TamperNet: A Spatio-Temporal Deep Learning Framework for Video Manipulation Detection. <em>Neural Processing Letters</em>. <a href="https://doi.org/10.1007/s11063-026-11885-8" rel="noopener noreferrer">https://doi.org/10.1007/s11063-026-11885-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11063-026-11885-8" rel="noopener noreferrer">10.1007/s11063-026-11885-8</a></p>
<p><strong>Keywords:</strong> video tampering detection, deep learning, spatio-temporal analysis, LSTM, autoencoder, ResNet-50, video forensics, frame splicing, inpainting, anomaly detection, Neural Processing Letters, digital forensics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223294</post-id>	</item>
		<item>
		<title>New AI Model Catches Deepfakes by Listening and Watching Simultaneously</title>
		<link>https://scienmag.com/new-ai-model-catches-deepfakes-by-listening-and-watching-simultaneously/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 23:04:46 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advancements in deepfake detection competitions]]></category>
		<category><![CDATA[AI-based fake media identification]]></category>
		<category><![CDATA[audio-visual fusion]]></category>
		<category><![CDATA[audio-visual synchronization]]></category>
		<category><![CDATA[audiovisual forensics]]></category>
		<category><![CDATA[BYOL-A]]></category>
		<category><![CDATA[challenges in detecting synthetic videos with cloned voices]]></category>
		<category><![CDATA[cross-attention]]></category>
		<category><![CDATA[cross-modal forgery detection techniques]]></category>
		<category><![CDATA[DDL-AV dataset]]></category>
		<category><![CDATA[deepfake detection]]></category>
		<category><![CDATA[deepfake localization and interpretability]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[ERF-BA-TFD+ model]]></category>
		<category><![CDATA[generative adversarial networks]]></category>
		<category><![CDATA[LAV-DF dataset]]></category>
		<category><![CDATA[multimodal AI models]]></category>
		<category><![CDATA[multimodal deepfake detection systems]]></category>
		<category><![CDATA[multimodal learning]]></category>
		<category><![CDATA[MViTv2]]></category>
		<category><![CDATA[synthetic media forgeries]]></category>
		<category><![CDATA[temporal localization]]></category>
		<category><![CDATA[transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215172</guid>

					<description><![CDATA[A multimodal AI model called ERF-BA-TFD+ detects and localizes audio-visual deepfakes by cross-reconstructing audio and video features and reasoning over full-length footage, winning first place in an international deepfake detection competition.]]></description>
										<content:encoded><![CDATA[<p>Deepfakes have long been treated as a visual problem: a swapped face, a warped expression, a flicker of unnatural texture around the mouth. But the most dangerous forgeries circulating today are multimodal, blending synthetic video with cloned voices, generated speech, and mismatched audio tracks that would fool a casual viewer completely. A team of researchers led by Leyan Wang, Jian Zhao, and Zhaofeng He, working at the Institute of Artificial Intelligence (TeleAI) at China Telecom along with collaborators at Beijing University of Posts and Telecommunications and several Chinese universities, has now unveiled a detection system built specifically for this messier reality. Their model, called ERF-BA-TFD+, was described in the open-access journal Vicinagearth in November 2025, and it recently took first place in the Audio-Visual Detection and Localization track of the Workshop on Deepfake Detection, Localization, and Interpretability, a competition centered on the demanding DDL-AV dataset.</p>
<p>The core insight behind the new system is that real audio and real video are locked together in ways that are extremely difficult to forge consistently. When a person speaks, lip movements, facial expressions, prosody, and timing all co-vary. Generators can produce convincing audio and convincing video separately, but keeping the two streams synchronized and semantically consistent across an entire clip is far harder. ERF-BA-TFD+ exploits exactly this gap. Rather than analyzing each modality in isolation, as most earlier detectors did, it processes audio and video features simultaneously and hunts for the subtle discrepancies that emerge when one stream has been manipulated and the other has not, or when both have been forged but their temporal alignment betrays the fabrication.</p>
<p>Technically, the framework rests on two powerful feature extractors. For the visual stream, the researchers chose MViTv2, a hierarchical multiscale vision transformer. Unlike standard vision transformers that process every token at a fixed scale, MViTv2 progressively pools features, letting it capture anomalies at multiple granularities, from pixel-level texture irregularities to coarse, motion-based artifacts such as inconsistent facial kinematics. For audio, the team employed BYOL-A, a self-supervised model pre-trained on a vast and diverse corpus of sound. The choice of a self-supervised encoder is strategic: supervised detectors trained to spot specific artifact types tend to overfit and fail when confronted with novel forgery methods, whereas BYOL-A learns a rich general-purpose representation that can flag unnatural prosody, lip-sync mismatches, and distortions that are often imperceptible to human ears.</p>
<p>The heart of the temporal forgery detection pipeline is a module called CRATrans, the Cross-Reconstruction Attention Transformer. Instead of simply concatenating audio and video features, which can dilute the distinctive signals in each stream, CRATrans adopts a cross-reconstruction paradigm. During training, the model is forced to reconstruct the feature sequence of one modality using features from the other as context: visual features are rebuilt from audio cues and vice versa. This adversarial exercise teaches the network fine-grained inter-modal temporal dependencies. In genuine footage, where the streams are aligned and semantically consistent, reconstruction error stays low. In manipulated content, the attempt to reconstruct one stream from the other produces a markedly higher error, a direct and robust indicator of forgery. Multi-head self-attention models long-range dependencies within each modality, while cross-attention handles information exchange between them, and the resulting attention weights and reconstruction errors help pinpoint anomalous regions at inference time.</p>
<p>Localization proceeds in a coarse-to-fine hierarchy. Independent frame-level classification modules screen each modality separately, producing per-frame anomaly scores that capture modality-specific inconsistencies before any fusion can obscure them. A dedicated Boundary Localization Module, inspired by the proposal-relation mechanism of BSN++, then converts anomalous frames into precise temporal boundaries. It builds a confidence matrix over all possible start-end frame pairs, sharpened by two complementary attention mechanisms: position-aware attention for global temporal context and channel-aware attention for relationships among feature channels. Two separate boundary modules, one per modality, ensure that precision in localizing a forged audio segment is not compromised by imprecision in video, and the final outputs are refined through soft non-maximum suppression, with parameters empirically tuned to merge overlapping proposals and suppress low-confidence detections.</p>
<p>Perhaps the most inventive component is the Evidence-based Reasoning Framework, or ERF, which addresses a blind spot that plagues nearly all temporal detectors: the full-fake video. A model with a finite temporal receptive field learns to spot forgeries by contrasting suspicious segments with authentic ones. But when an entire video is fabricated, there is no clean baseline to contrast against, and such clips can slip through undetected. ERF reasons statistically over the whole video&#8217;s detection output. Localized forgeries produce some segments with high confidence; fully real videos yield uniformly low scores. A full-fake video, by contrast, often produces moderately confident scores everywhere with no single peak. If the maximum confidence across predicted segments fails to exceed a threshold, ERF re-evaluates the video at a global level and can reclassify it as a potential full-fake. The module is lightweight, essentially a corrective rule layer, yet it closed a genuine gap in the model&#8217;s logic.</p>
<p>The experimental path to the final system was itself revealing, unfolding in three phases on the DDL-AV and LAV-DF datasets, both of which contain segmented clips and full-length videos and confront models with text-to-speech, voice cloning, voice swapping, face swapping, facial animation, and text-to-video generation. In the first phase, the baseline model scored impressively on LAV-DF, achieving an average precision of 0.9630 at an intersection-over-union threshold of 0.5, but training on DDL-AV actually degraded those lenient-threshold scores, a counter-intuitive result the authors attribute to hyperspecialization on subtle artifacts. Tellingly, the strictest metric, AP at 0.95, improved slightly, suggesting the fine-tuned model had grown more sensitive to forgeries requiring very precise localization even as it lost confidence on easy cases.</p>
<p>The second phase exposed and then fixed a deeper weakness. A bad-case analysis showed the fusion model was almost blind to audio-only forgeries, posting a near-zero AP of 0.0163 on that subset, apparently because the strong visual stream drowned out faint auditory manipulation signals. Integrating the Unified Multimodal Attention framework, which uses cross-attention to force a more judicious weighting of both streams, produced a dramatic rebound: AP at 0.5 on the audio-forgery subset jumped to 0.9243. The final phase added the ERF module and lifted the overall competition score to 0.78, with average recall in the top-100 detections on a long-video validation set rising from 0.6513 to 0.7886. The authors report that ERF-BA-TFD+ achieved state-of-the-art results on DDL-AV and outperformed most competing models on LAV-DF while also offering superior processing speed.</p>
<p>For a field racing to keep pace with generative models built on GANs, diffusion architectures, and variational autoencoders, the study offers both a practical tool and a design philosophy. Its lessons, that balanced cross-modal attention is essential for resisting single-modality attacks, and that global statistical reasoning can rescue detectors from the full-fake blind spot, form what the authors describe as a blueprint for future systems. Challenges remain, particularly severe or non-linear temporal desynchronization between audio and video and the approach of next-generation forgeries with fewer low-level artifacts. The team points toward zero-shot and few-shot learning as a route to rapid adaptation against unseen manipulation techniques, and both the LAV-DF dataset, available on Hugging Face, and the forthcoming full release of DDL-AV should give the wider community the means to build on this work as the arms race between fabrication and detection continues.</p>
<p><strong>Subject of Research:</strong> Multimodal audio-visual deepfake detection using cross-modal reconstruction attention and evidence-based reasoning</p>
<p><strong>Article Title:</strong> ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection</p>
<p><strong>Article References:</strong> Wang, L., Zhao, J., Zhang, X., Guo, X., Yuan, Y., Zhang, T., Chu, J., Jiang, Y., Yang, X., Jin, L., Zhang, C., &amp; He, Z. (2025). ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection. <em>Vicinagearth, 2</em>(1), Article 10. <a href="https://doi.org/10.1007/s44336-025-00021-0" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00021-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00021-0" rel="noopener noreferrer">10.1007/s44336-025-00021-0</a></p>
<p><strong>Keywords:</strong> deepfake detection, multimodal learning, audio-visual fusion, transformers, MViTv2, BYOL-A, temporal localization, DDL-AV dataset, LAV-DF dataset, digital forensics, generative adversarial networks, cross-attention</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215172</post-id>	</item>
		<item>
		<title>Diffusion Models Get a Forensic Upgrade: Two-Stage AI Pinpoints Doctored Pixels in Photos</title>
		<link>https://scienmag.com/diffusion-models-get-a-forensic-upgrade-two-stage-ai-pinpoints-doctored-pixels-in-photos/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 22:32:47 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced techniques for detecting manipulated pixels]]></category>
		<category><![CDATA[AI-based photo forgery detection]]></category>
		<category><![CDATA[conditional diffusion models]]></category>
		<category><![CDATA[copy-move forgery]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for image authenticity]]></category>
		<category><![CDATA[DF2023 dataset]]></category>
		<category><![CDATA[diffusion model for image forensics]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[dual-stream classifier]]></category>
		<category><![CDATA[forensic classifier for doctored images]]></category>
		<category><![CDATA[forensic image manipulation detection]]></category>
		<category><![CDATA[generalization challenges in image forgery detection]]></category>
		<category><![CDATA[Generative Models]]></category>
		<category><![CDATA[identifying subtle image manipulations]]></category>
		<category><![CDATA[image forensics]]></category>
		<category><![CDATA[image manipulation localization]]></category>
		<category><![CDATA[inpainting localization]]></category>
		<category><![CDATA[multimedia forensics using diffusion models]]></category>
		<category><![CDATA[pixel-level image tampering localization]]></category>
		<category><![CDATA[real-world application of AI in image authenticity]]></category>
		<category><![CDATA[splicing detection]]></category>
		<category><![CDATA[steganalysis rich model]]></category>
		<category><![CDATA[two-stage AI framework for image forensics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212847</guid>

					<description><![CDATA[Researchers at the National Institute of Technology Goa have built a two-stage framework that first classifies image forgeries with a dual-stream forensic classifier and then uses conditional diffusion models to generate precise pixel-level manipulation masks.]]></description>
										<content:encoded><![CDATA[<p>Every day, millions of images circulate through social media, news outlets, and courtrooms, and a growing share of them have been quietly altered. A cloned patch of sky, a spliced-in face, an airbrushed-out bystander — these edits are often invisible to the human eye, yet they can shape elections, damage reputations, and even sway legal verdicts. Researchers at the National Institute of Technology Goa have now proposed a fresh way to catch such forgeries, described in the journal Multimedia Tools and Applications, that pairs a specialized forensic classifier with the most talked-about technology in modern artificial intelligence: diffusion models.</p>
<p>The new framework, developed by Mohammad Zohaib Hamdule and Venkatanareshbabu Kuppili, tackles a task known as image manipulation localization, or IML. Detection alone is not enough for most real-world applications; investigators need to know exactly which pixels in a photograph were tampered with. That is a far harder problem, because manipulation traces are subtle, varied, and constantly evolving as editing tools improve. Conventional deep learning approaches, the authors note, often struggle to generalize beyond the specific forgery techniques they were trained on, faltering when confronted with new manipulation schemes.</p>
<p>The team&#8217;s answer is a two-stage system that splits the problem in two. Rather than asking a single network to simultaneously figure out whether an image is fake, what kind of fakery was used, and where it happened, the framework first classifies and then localizes. This modular design mirrors the way a human forensic analyst works: identify the type of edit first, then apply the right analytical lens to trace its boundaries. Modularity also brings a practical bonus — each stage can be upgraded independently as new techniques emerge.</p>
<p>The first stage is a Dual-Stream Manipulation Classifier, and its architecture reveals a deep understanding of how digital forgeries leave fingerprints. One stream processes the image in its ordinary RGB form, capturing semantic content — textures, objects, edges. The second stream is more forensic in spirit: it passes the image through Steganalysis Rich Model filters, a family of high-pass filters borrowed from the field of steganalysis, where researchers have long used them to expose hidden data embedded in images. These SRM filters suppress the natural content of the photograph and amplify low-level noise artifacts — the microscopic inconsistencies left behind whenever pixels are copied, spliced, erased, or enhanced.</p>
<p>Both streams feed into a ResNet-style four-stage backbone, the workhorse convolutional architecture that has underpinned computer vision for nearly a decade. By fusing standard visual features with these noise residuals, the classifier learns to recognize four of the most common manipulation categories: Copy-Move, where a region is duplicated and pasted elsewhere in the same image; Splicing, where content from one photograph is inserted into another; Removal, also called inpainting, where an object is erased and the hole filled in; and Enhancement, where attributes such as color, brightness, or fine detail are adjusted to deceive. On the DF2023 dataset, a benchmark for digital forensics, this classifier reached an accuracy of 89 percent — a strong result given how visually different the four categories can be.</p>
<p>Once the manipulation type is known, the image is routed to the second stage: a set of specialized Conditional Diffusion Models, one for each manipulation class. Diffusion models, the same family of generative networks behind today&#8217;s most impressive text-to-image systems, work by learning to reverse a gradual noising process. In this framework they are repurposed for an entirely different goal: instead of generating photorealistic pictures, they generate masks — binary maps that paint the manipulated region white and the untouched background black. The localization task is thereby reframed as an image-to-mask generation problem, with the suspect photograph serving as the conditioning input that guides the denoising process toward the correct answer.</p>
<p>Training such generative models for precise localization demanded a technical innovation of its own. The standard training objective for image generation, mean squared error, treats every pixel equally and tends to wash out small targets. A tiny spliced region or a narrow inpainted stroke occupies only a handful of pixels, and a model trained purely on squared error can learn to predict a blank mask and still score decently. The researchers therefore modified the loss function to combine mean squared error with Intersection over Union, the standard overlap metric in segmentation. This hybrid objective pushes the model to reproduce not just approximate shading but the exact spatial structure of the manipulation, with particular benefit for smaller masks that would otherwise be smoothed away.</p>
<p>The numbers back up the design. Across the DF2023 dataset, the localization diffusion models achieved an average Intersection over Union of 0.70 and an F1 score of 0.77 — metrics that balance precision and recall when judging how faithfully the predicted mask matches the true tampered region. The system also demonstrated competitive performance on well-established benchmark datasets including IMD2020, CoMoFoD, CASIA, and COVERAGE, which collectively span realistic splices, copy-move forgeries, and controlled manipulation scenarios. Consistency across these heterogeneous collections suggests the approach is not merely memorizing the quirks of one dataset, a persistent weakness in the field.</p>
<p>What makes the work especially timely is the central paradox it highlights: generative models now create the forgeries, and generative models can also expose them. Earlier attempts to bring generative machinery to forensics leaned on generative adversarial networks, which produce output in a single pass and can be unstable to train. Diffusion models, by contrast, refine their predictions over many iterative denoising steps, an approach that has recently proven effective in segmentation tasks from medical imaging to remote sensing. The Goa team&#8217;s results add image forensics to that growing list, joining related efforts that use diffusion-based models for inpainting localization and forgery localization more broadly.</p>
<p>The implications extend well beyond the laboratory. Investigators and prosecutors increasingly rely on digital images as evidence, and studies have shown that people are surprisingly poor at spotting manipulated photos of real-world scenes. A tool that can automatically classify the type of forgery and trace its pixel-level boundaries could strengthen fact-checking workflows, support media authentication desks, and give courts a more rigorous basis for judging photographic evidence. The modular architecture also offers a pragmatic path forward: as AI-generated and AI-edited imagery grows more sophisticated, individual components of the pipeline — new filters, new classifiers, new generative backbones — can be swapped in without rebuilding the entire system. For now, the framework&#8217;s 89 percent classification accuracy and 0.70 average IoU represent a meaningful step toward forensic tools that can keep pace with the editing software they are built to catch, confirming that a modular, generative strategy has real promise in the escalating contest between image manipulation and image verification.</p>
<p><strong>Subject of Research:</strong> Image manipulation localization using dual-stream classification and conditional diffusion models</p>
<p><strong>Article Title:</strong> A modular image manipulation localization framework using a dual-stream classifier and conditional diffusion models</p>
<p><strong>Article References:</strong> Hamdule, M. Z., &amp; Kuppili, V. (2026). A modular image manipulation localization framework using a dual-stream classifier and conditional diffusion models. <em>Multimedia Tools and Applications, 85</em>(10), Article 778. <a href="https://doi.org/10.1007/s11042-026-21940-0" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21940-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21940-0" rel="noopener noreferrer">10.1007/s11042-026-21940-0</a></p>
<p><strong>Keywords:</strong> image forensics, image manipulation localization, conditional diffusion models, dual-stream classifier, deep learning, steganalysis rich model, copy-move forgery, splicing detection, inpainting localization, DF2023 dataset, digital forensics, generative models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212847</post-id>	</item>
		<item>
		<title>AI, Drones and Blockchain Reshape Disaster Victim Identification</title>
		<link>https://scienmag.com/ai-drones-and-blockchain-reshape-disaster-victim-identification/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:57:03 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[3D printing]]></category>
		<category><![CDATA[AI in forensic analysis]]></category>
		<category><![CDATA[algorithmic bias]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[blockchain]]></category>
		<category><![CDATA[blockchain for data security in DVI]]></category>
		<category><![CDATA[challenges of identifying human remains after disasters]]></category>
		<category><![CDATA[Data Privacy]]></category>
		<category><![CDATA[digital forensics]]></category>
		<category><![CDATA[disaster victim identification]]></category>
		<category><![CDATA[DNA phenotyping]]></category>
		<category><![CDATA[DNA profiling in disaster victim ID]]></category>
		<category><![CDATA[drone technology for disaster recovery]]></category>
		<category><![CDATA[drones]]></category>
		<category><![CDATA[ethical considerations in forensic technology]]></category>
		<category><![CDATA[forensic anthropology techniques]]></category>
		<category><![CDATA[forensic odontology methods]]></category>
		<category><![CDATA[forensic science]]></category>
		<category><![CDATA[forensic science advancements]]></category>
		<category><![CDATA[mass disasters]]></category>
		<category><![CDATA[remote sensing]]></category>
		<category><![CDATA[technological innovations in forensic investigations]]></category>
		<category><![CDATA[use of AI and drones in mass casualty events]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204800</guid>

					<description><![CDATA[A new narrative review maps how artificial intelligence, drones, DNA phenotyping and blockchain could transform disaster victim identification while raising serious ethical concerns about privacy, bias and accountability.]]></description>
										<content:encoded><![CDATA[<p>When a catastrophic earthquake, tsunami or aircraft crash claims hundreds or thousands of lives, the grim work of identifying the dead becomes one of the most demanding tasks in all of forensic science. Disaster victim identification, known universally in the field as DVI, exists to ensure that human remains are matched to names with rigor and dignity, allowing families to bury their loved ones and legal systems to close the record. A narrative review published in the Journal of Emergency and Disaster Medicine by Doaa Tawfik of Cairo University&#8217;s Department of Forensic Medicine and Clinical Toxicology surveys the technological wave now breaking over this solemn discipline, and delivers a clear warning: the same tools that promise speed and accuracy also carry profound ethical risks that forensic teams are only beginning to confront.</p>
<p>Traditional DVI rests on three established pillars. Forensic anthropology applies skeletal analysis and archaeological technique to remains that may be fragmented, burned or decomposed, guiding recovery and interpretation. Forensic odontology compares dental records, which often survive conditions that destroy other identifiers. DNA profiling, widely regarded as the most reliable method, analyzes genetic material from remains and compares it against reference samples donated by relatives or recovered from personal items. These techniques work, the review notes, but they are time-consuming, resource-intensive and mentally taxing, particularly in mass casualty scenarios where thousands of data points must be manually compared, all while a strict chain of custody is maintained to protect the legal and ethical rights of the deceased.</p>
<p>The review identifies a paradigm shift underway across three stages of the DVI workflow: scene management and recovery, victim identification itself, and data integrity and management. At the disaster scene, the first stage, drones and remote sensing are emerging as force multipliers. Equipped with real-time aerial imaging and thermal scanning, drones can survey affected areas that ground teams cannot reach safely or quickly, a critical advantage because delays in reaching remains accelerate post-mortem DNA degradation and complicate identification. Unmanned aircraft can deliver sampling kits and rapid-DNA devices, and recent studies demonstrate the feasibility of aerial environmental DNA sampling and surface swabbing, recovering trace human DNA from vegetation and surfaces as a supplementary, non-contact approach when direct recovery is delayed. Validation studies and operational protocols are still required before routine integration, but the strategic role of drones in extending sampling capacity in protracted or inaccessible disaster environments is now well supported by the literature.</p>
<p>Yet the aerial revolution arrives with baggage. Drones raise safety issues for first responders in the event of malfunction, confidentiality concerns about data collected over surveilled neighborhoods, and questions about algorithmic decision-making bias. The review flags a deeper structural problem: the majority of drone studies have been conducted in or by nations that develop and own these advanced technologies, so reported success rates, cost-benefit analyses and logistical frameworks may not translate to resource-constrained regions without technical expertise. The literature also reveals a paucity of validated evidence on drones&#8217; actual capacity to identify disaster victims, partly because conducting research during real-world disasters is ethically and logistically fraught. Simulation studies, meanwhile, suffer from such heterogeneity in design that the review calls for a standardized disaster simulation checklist to reduce bias and improve methodological consistency.</p>
<p>Inside mortuaries and identification units, 3D printing is reshaping forensic reconstruction. The technology can produce lifelike facial models based on skeletal remains, aiding visual identification and increasing the chances of recognition. But accuracy remains a concern, because bone density and surface characteristics cannot be fully replicated and modeling parameters affect print quality. There is also a uniquely modern hazard: the open-source culture of the 3D community means any model can be easily shared, downloaded and printed, potentially compromising evidence integrity. These issues have kept 3D-printed evidence on shaky admissibility footing in courts. In 2023, researchers in the UK made the first attempt to create an ethical framework for 3D reconstruction, articulating nine principles including transparency, beneficence, context, non-maleficence and anonymity.</p>
<p>Digital forensics has opened another identification channel. Smartphones accumulate extensive personal and behavioral metadata, including contacts, messaging logs, geolocation traces, gait data and app usage history, all of which analysts can extract and correlate with external records to support identity hypotheses. When victims carried implanted medical devices or wearables connected to phone applications, communication artifacts such as timestamps, device IDs and telemetry logs stored on the phone can serve as a digital bridge linking the device to its owner. Encryption and data deletion remain significant hurdles, and severe physical damage to devices in disasters limits usefulness, so the review emphasizes that smartphones should aid DVI efforts rather than stand alone. The field also faces mounting ethical risks around data accuracy, standardization, confidentiality and accountability, compounded by non-technical factors like inadequate training, cognitive bias and poor case management, particularly where speed is prioritized over accuracy.</p>
<p>The most transformative and most ethically charged technology is artificial intelligence. Machine learning is already applied to DNA mixture deconvolution, ancestry prediction, kinship matching and assessing the forensic value of complex samples. Deep learning shows promise in estimating age and sex from skeletal remains, dental records and medical images, and in predicting post-mortem interval and cause of death. AI-driven face restoration using diffusion models and Generative Adversarial Networks has been shown to improve forensic face recognition accuracy by reconstructing degraded images of the deceased to a more identifiable state, and AI software can compare vast datasets efficiently, reducing human error and accelerating identification. The caveats, however, are substantial. AI can generate inaccurate information, a phenomenon known as AI hallucination, which must be scientifically scrutinized and documented. Both human and algorithmic bias can enter through candidate selection and through demographic composition of training datasets, meaning models may perform poorly on populations outside their training data and misidentification could compound the tragedy for affected families.</p>
<p>The review draws instructive parallels from commercial deployments. Analysis of facial recognition cases such as Clearview AI and airport biometric surveillance reveals fundamental tensions between technological innovation and privacy, confidentiality and informed consent, exposing systemic gaps in governance. The implication for forensic science is stark: if using biometric data without explicit consent in public spaces raises serious societal and regulatory challenges, applying such technologies to vulnerable deceased populations, where consent can never be obtained, demands even more rigorous scrutiny and restrictive governance. The well-known Gender Shades study demonstrated significant accuracy disparities across race and gender intersections in commercial classification systems, showing that biased datasets and opaque model design can perpetuate systemic discrimination, and that these risks are not theoretical but have already manifested in practice. The review also stresses cultural sensitivity: fairness and accountability in AI for disaster risk management require local stakeholder inclusion and transparent decision pathways, and generative AI systems must be culturally tailored to different racial and ethnic communities to maintain trust during crisis communication.</p>
<p>DNA phenotyping and predictive biometrics extend the frontier further, allowing scientists to generate probable facial structures, eye color and ancestry information from genetic material alone, which is valuable when no missing-persons list or reference sample exists. But the accuracy of phenotyping remains a challenge, especially with mixed DNA samples, many countries lack legal frameworks governing its responsible use, and no new DNA markers have been established to support accurate measurement and validation. Predictions are probabilistic and subject to error, so misclassification may produce misleading leads or unfair targeting of individuals or groups. Privacy and informed consent are at stake when samples are used without explicit permission to infer traits or ancestry that individuals may consider sensitive, and the literature increasingly calls for privacy impact assessment frameworks before laboratories and law enforcement adopt the technology.</p>
<p>For data integrity, blockchain offers a potential revolution in chain of custody. As a distributed ledger producing immutable, time-stamped, cryptographically secured records shared across multiple nodes, it prevents unilateral modification of stored data and could enable secure, unified platforms for storing, managing and cross-jurisdictionally comparing sensitive identifying data such as DNA, dental and medical records. Reviews support blockchain&#8217;s usefulness for preserving evidence integrity and enabling real-time global collaboration among healthcare teams, but concerns remain over data collection, confidentiality, sharing, ownership, cost, privacy and unauthorized access. The review concludes with a set of recommendations: training disaster teams on ethically sound and culturally appropriate technologies, developing checklists and guidelines aligned with local, national and international regulations, building internationally recognized blockchain forensic databases, exploring bioethical policies for predictive biometrics and genetic privacy with compensation mechanisms for bias, and operationalizing the proposed ethical framework through pilot programs in disaster-prone regions. Accountability, the review insists, must ultimately remain with human and institutional actors, preserving what it calls attributability so that identification decisions express human values. Forensic science, it argues, must evolve with a dual focus, embracing cutting-edge technology while upholding the highest ethical standards, so that identification becomes not only faster and more reliable but fair, transparent and respectful of victims&#8217; dignity.</p>
<p><strong>Subject of Research:</strong> A narrative review of emerging technologies and their ethical implications in disaster victim identification</p>
<p><strong>Article Title:</strong> Disaster victim identification: a narrative review of innovations and ethical considerations</p>
<p><strong>Article References:</strong> Tawfik, D. (2026). Disaster victim identification: a narrative review of innovations and ethical considerations. <em>Journal of Emergency and Disaster Medicine, 2</em>(1), Article 7. <a href="https://doi.org/10.1007/s44467-026-00010-3" rel="noopener noreferrer">https://doi.org/10.1007/s44467-026-00010-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44467-026-00010-3" rel="noopener noreferrer">10.1007/s44467-026-00010-3</a></p>
<p><strong>Keywords:</strong> disaster victim identification, forensic science, artificial intelligence, DNA phenotyping, blockchain, drones, remote sensing, 3D printing, digital forensics, algorithmic bias, data privacy, mass disasters</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204800</post-id>	</item>
	</channel>
</rss>
