<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>privacy and copyright concerns in AI &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/privacy-and-copyright-concerns-in-ai/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 16:26:10 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>privacy and copyright concerns in AI &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Edge Maps and Depth Images Can Unlock What AI Safety Filters Erased</title>
		<link>https://scienmag.com/edge-maps-and-depth-images-can-unlock-what-ai-safety-filters-erased/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 16:26:10 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial attacks]]></category>
		<category><![CDATA[adversarial attacks on AI safety]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[AI safety filters]]></category>
		<category><![CDATA[concept erasure]]></category>
		<category><![CDATA[concept erasure in generative models]]></category>
		<category><![CDATA[concept suppression in AI]]></category>
		<category><![CDATA[ControlNet]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[evaluation of AI content filters]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[generative AI risk assessment]]></category>
		<category><![CDATA[grey-box attack]]></category>
		<category><![CDATA[image generation safety]]></category>
		<category><![CDATA[machine unlearning]]></category>
		<category><![CDATA[model robustness]]></category>
		<category><![CDATA[model safety testing methods]]></category>
		<category><![CDATA[privacy and copyright concerns in AI]]></category>
		<category><![CDATA[reappearance of erased concepts]]></category>
		<category><![CDATA[structural conditioning]]></category>
		<category><![CDATA[structural control tools in AI]]></category>
		<category><![CDATA[text-to-image diffusion models]]></category>
		<category><![CDATA[text-to-image generation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248589</guid>

					<description><![CDATA[Researchers have shown that concepts erased from text-to-image diffusion models can reappear when structural controls like edge maps and depth maps guide generation, exposing a critical gap in AI safety evaluation.]]></description>
										<content:encoded><![CDATA[<p>Text-to-image diffusion models have become one of the most widely deployed generative technologies in the world, powering creative design tools, entertainment platforms, and content production pipelines. Because these models are trained on enormous collections of web-scraped images, they inevitably retain concepts that operators would rather they never produce: explicit sexual content, the copyrighted styles of living artists, and sensitive portraits of private individuals. In response, a thriving research field known as concept erasure has emerged, promising to surgically suppress specific ideas from a model while leaving its general creative abilities intact. But a new study published in the journal Cybersecurity reveals a striking blind spot in how these safeguards are tested, and the findings could force a fundamental rethink of how generative AI safety is measured.</p>
<p>The research, led by Qiqi Bao and Jiaoling Li of Zhejiang University of Science and Technology together with colleagues at Harbin Institute of Technology Shenzhen, Zhejiang University, and Université Paris Cité, demonstrates that concepts which appear thoroughly erased under standard text-based testing can reappear with alarming ease once the model is combined with the structural control tools that define modern image generation. The team calls the attack Structure-Guided Concept Reappearance, or SGCR, and its central insight is deceptively simple: safety evaluations focus almost exclusively on the text pathway, while real-world diffusion systems increasingly generate images through a second, non-textual route that defenses never touch.</p>
<p>To understand why this matters, it helps to look at how contemporary diffusion pipelines actually work. A latent diffusion model consists of an encoder that compresses images into a compact latent representation, a decoder that reconstructs them, and a denoising network that gradually transforms random noise into a coherent image. Text prompts influence this process mainly through cross-attention modules that bind semantic content to spatial locations. Concept erasure methods, whether they edit model weights as ESD, UCE, MACE, and TRCE do, or intervene at inference time as Negative Prompting, Safe Latent Diffusion, and AdaVD do, all share the same underlying objective: they weaken the association between a textual trigger and its corresponding visual concept. The implicit assumption has always been that if the text route is blocked, the concept is effectively gone.</p>
<p>That assumption collapses when a structural controller enters the picture. Tools such as ControlNet have made it routine for users to steer generation with Canny edge maps, depth maps, or semantic segmentation maps derived from reference images, all without modifying the backbone model. These structural representations discard color, texture, and most appearance information, which is precisely why they were assumed to be semantically inert. The researchers show otherwise. Using a zero-shot CLIP probe, they demonstrated that edge, depth, and segmentation maps extracted from target-concept reference images retain enough category-related geometry, spatial layout, and instance-specific cues to reliably distinguish target conditions from benign ones. A line drawing stripped of every pixel of explicit content still carries the silhouette of what it depicts.</p>
<p>The SGCR attack exploits this residual information under a deliberately practical grey-box threat model. The attacker needs no access to training data, no gradients from the defended model, and no ability to modify its parameters. Instead, the attacker simply knows the deployed backbone-controller configuration, supplies a completely benign text prompt containing no target semantics, and attaches a pre-trained structural controller fed with a condition extracted from a publicly available reference image of the target concept. The controller injects geometry-consistent residual features into the U-Net through zero-convolution connections, and these structurally modulated representations are progressively integrated across downstream blocks during denoising, steering the latent trajectory toward spatial configurations that conform to the supplied scaffold.</p>
<p>The experimental results are stark. Across twelve target concepts spanning explicit content from the I2P benchmark, copyrighted artistic styles including Van Gogh, Picasso, and Kelly McKernan, and generic objects drawn from CIFAR-10, SGCR markedly increased attack success rates against representative defenses from all three major paradigms. MACE, which reduced the text-only attack success rate for nudity to zero percent, saw that figure climb to 72 percent under Canny guidance, while its success rate on Picasso reached a full 100 percent. UCE&#8217;s nudity rate rose from 22 to 49 percent, and its Cat success rate jumped from 2 to 74 percent. Notably, the Fréchet Inception Distance of MACE&#8217;s nudity outputs fell from 48.61 to 7.37, meaning the attacked images were not degraded artifacts but high-quality generations closely matching the original model&#8217;s output distribution.</p>
<p>The team went to considerable lengths to rule out alternative explanations. In a controlled prompt-structure analysis with seven distinct conditions, they showed that adding target-consistent structure while holding the text fixed raised success rates by 20 to 51 percentage points across all eight models tested, whereas pairing the benign prompt with a non-target human structure produced far weaker effects. This establishes that the reappearance is driven by the semantic content of the structural condition itself, not merely by the activation of a controller. They also found that exact prompt-reference pairing was unnecessary: cyclically shifting target-category structures across prompts yielded statistically indistinguishable results, indicating category-level transfer across reference instances. A multi-reference, multi-seed extension covering three concepts and four defenses confirmed that the behavior persists across 100 distinct reference images and three random seeds, with success rates ranging from 58 to 84.3 percent.</p>
<p>Mechanistic probes added further texture to the picture. When the researchers replaced attention maps with those from paired benign generations, swapping self-attention produced larger drops in attack success than swapping cross-attention, suggesting that image-side feature interactions play a greater role than the text-conditioning pathway in propagating structural information. Activation-energy visualizations showed self-attention responses aligning with the supplied structural contours even as text-related cross-attention remained weak, and intermediate denoising trajectories revealed target-consistent spatial organization emerging progressively in the mid-to-late sampling stages. An ablation over guidance strength found the attack most effective in an intermediate range around a scale of 1.0, with effects saturating beyond that point. The authors are careful to frame all of this as a system-level property of the composed backbone-controller pipeline rather than proof that the erased backbone secretly harbors the concept in a single recoverable module.</p>
<p>The practical implications are scoped precisely but seriously. The demonstrated threat applies to local modular systems and to services that expose compatible structural-control interfaces, not to closed text-only APIs, since an attacker must be able to attach a controller and supply a structural condition. Within that scope, the consequences extend beyond regenerating a single reference image: structural conditions can be reused with different benign prompts and seeds to produce endless target-consistent variants, potentially defeating exact-match and duplicate-detection safeguards. The authors propose three directions for future work: evaluation protocols that include controller-augmented generation, defenses that jointly suppress text-conditioned generation, image-side feature interactions, and attached controllers, and safety alignment studied at the system level. The parallel they draw is instructive: just as visual inputs have jailbroken aligned large language models, multimodal deployments of diffusion models can expose gaps that no text-focused evaluation will ever catch. For an industry racing to certify the safety of generative systems, the message is uncomfortable but clear: erasing a concept from what a model hears does not mean erasing it from what a model can be shown.</p>
<p><strong>Subject of Research:</strong> Evasion attacks on concept erasure safeguards in text-to-image diffusion models using structural conditioning</p>
<p><strong>Article Title:</strong> Evasion attacks on generative safeguards: target-concept reappearance under structural control in grey-box settings</p>
<p><strong>Article References:</strong> Bao, Q., Li, J., Zhang, Y., Qian, Y., Gu, Z., Ji, S., Wang, B., &amp; Naït-Abdesselam, F. (2026). Evasion attacks on generative safeguards: target-concept reappearance under structural control in grey-box settings. <em>Cybersecurity, 9</em>(1), Article 227. <a href="https://doi.org/10.1186/s42400-026-00664-6" rel="noopener noreferrer">https://doi.org/10.1186/s42400-026-00664-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s42400-026-00664-6" rel="noopener noreferrer">10.1186/s42400-026-00664-6</a></p>
<p><strong>Keywords:</strong> diffusion models, concept erasure, AI safety, ControlNet, adversarial attacks, text-to-image generation, structural conditioning, machine unlearning, cybersecurity, generative AI, grey-box attack, model robustness</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248589</post-id>	</item>
	</channel>
</rss>
