<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>shadow detection and elimination in images &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/shadow-detection-and-elimination-in-images/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 13:25:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>shadow detection and elimination in images &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>How AI Is Learning to Erase Shadows From the Documents We Photograph Every Day</title>
		<link>https://scienmag.com/how-ai-is-learning-to-erase-shadows-from-the-documents-we-photograph-every-day/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 13:25:14 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advancements in AI for archiving and digitization]]></category>
		<category><![CDATA[AI-powered document shadow removal]]></category>
		<category><![CDATA[benchmark datasets]]></category>
		<category><![CDATA[challenges in capturing clear digital documents]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[datasets]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for shadow removal]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[document image processing]]></category>
		<category><![CDATA[generative adversarial networks]]></category>
		<category><![CDATA[image enhancement]]></category>
		<category><![CDATA[image processing for document enhancement]]></category>
		<category><![CDATA[impact of shadows on readability of photographed documents]]></category>
		<category><![CDATA[international research on document image enhancement]]></category>
		<category><![CDATA[OCR]]></category>
		<category><![CDATA[open-source datasets for document image processing]]></category>
		<category><![CDATA[Retinex theory]]></category>
		<category><![CDATA[shadow detection and elimination in images]]></category>
		<category><![CDATA[shadow removal]]></category>
		<category><![CDATA[smartphone document photography]]></category>
		<category><![CDATA[survey of AI techniques in document image restoration]]></category>
		<category><![CDATA[transformer and diffusion models for image correction]]></category>
		<category><![CDATA[transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222978</guid>

					<description><![CDATA[The first comprehensive survey of document shadow removal catalogs nine datasets and 41 studies, mapping the field's evolution from physics-based illumination models to transformers and diffusion networks.]]></description>
										<content:encoded><![CDATA[<p>Every day, billions of photographs of documents are taken with smartphones: receipts, contracts, homework pages, newspaper clippings, and archival records. Yet a simple act of physics conspires against this digital convenience. Whenever a hand, a phone, or an object blocks ambient light, a shadow falls across the page, darkening text, distorting colors, and degrading the readability of what should be a clean digital copy. A newly published comprehensive survey in the open-access journal Vicinagearth, led by Bingshu Wang and Changping Li of Northwestern Polytechnical University together with colleagues at Shenzhen University, Guizhou University, Huizhou University, and South China University of Technology, offers the first systematic map of this deceptively hard problem, cataloging nine open-source datasets, forty-one research works, and a taxonomy of methods that spans classical image processing to modern transformers and diffusion models.</p>
<p>The scale of the literature review itself tells a story about how fast this field is moving. The authors searched databases such as Google Scholar and IEEE Xplore using terms including document shadow removal, shadow detection and removal, and document image enhancement, collecting work published up to September 2024 from authors in more than ten countries. Of the 41 selected publications, 30, or 73.2 percent, appeared at conferences while 11, or 26.8 percent, appeared in journals, a distribution that reflects the field&#8217;s preference for rapid dissemination at venues such as the Conference on Computer Vision and Pattern Recognition, the International Conference on Computer Vision, and ICASSP. China accounts for the largest share of output, with six journal and thirteen conference papers representing 46.3 percent of the total. The temporal trend is equally striking: for journal papers, 72.7 percent of publications are concentrated in 2023 and 2024, and conference output has climbed steadily since 2022, signaling a research area that has only recently reached critical mass.</p>
<p>To understand why shadow removal from documents is harder than it looks, consider what a shadow actually does to a page. It attenuates illumination non-uniformly, often with soft penumbral edges, and it interacts with the underlying content: thin strokes of text, colored figures, paper texture, and even the curvature of a thick bound book. Early benchmark datasets captured this difficulty only partially. The Adobe dataset introduced by Bako and colleagues in 2016 contains just 81 pairs of shadowed and shadow-free images captured with a Canon 5D Mark II, drawn from 11 documents with 5 to 9 variations in shadow intensity and shape, and it includes only light shadows on plain text pages. Jung and colleagues added 87 pairs featuring multiple cast shadows created by deliberately obstructing the light source, while Shah and colleagues produced a dataset of 100 images with deliberately harder shadows and sharp transitions using obstacles at varying distances from a Moto G5 Plus camera. Each of these early collections was small and limited to narrow document types, which constrained how well models trained on them could generalize.</p>
<p>The dataset landscape has since matured considerably. Kligler and colleagues expanded diversity with 300 image pairs spanning handwritten documents, printed documents, posters, and fonts, each category containing 7 to 10 documents with 8 to 12 shadow variations. Wang and colleagues built the Optical Shadow Removal dataset of 237 manually captured images of books, newspapers, and brochures shot on an iPhone XR specifically to address the scarcity of strong shadows. The most significant recent contributions tackle the scale problem head-on. Lin and colleagues created the Synthetic Document Shadow Removal Dataset with 8,309 pairs of synthetic shadow images, shadow-free images, and shadow masks generated from 970 documents under varied lighting and occluders, alongside a real-world companion set of 540 captured images. Matsuo and colleagues pushed synthesis further with a fully synthetic dataset that generates documents from text, graphics, and textures and renders shadows with a graphics engine. Most ambitiously, Zhang and colleagues assembled the Real Document Dataset, 4,916 real image pairs split into 4,371 for training and 545 for testing, covering papers, books, and brochures across complex scenes. The survey evaluates all nine datasets on scale, authenticity, and diversity, noting that synthetic data solves the pairing problem, since true shadow-free and shadowed captures of the same page under identical conditions are nearly impossible to collect in the wild, but that synthesizing realistic multi-source lighting remains an open challenge.</p>
<p>On the algorithmic side, the survey draws a fundamental line between conventional methods, which rely on manually designed features and physical models, and neural network-based methods, which learn features directly from data. Conventional approaches split into two families. Shadow map-based methods first estimate where the shadow lies, often by comparing local background colors to global ones, then use that map to correct the image. Bako&#8217;s pioneering algorithm estimated local text and background colors and generated a shadow map from their ratio, achieving an average mean squared error of 22.26, though its assumption of a constant background limited performance on varied pages. Kligler&#8217;s group reinterpreted image points as 3D point clouds to separate foreground from background and reached a structural similarity score of 0.943 on printed documents. Later work by Wang and colleagues refined background estimation with iterative shadow scaling, and a joint water-filling algorithm with adaptive chroma adjustment maintained brightness and color consistency across shadow boundaries.</p>
<p>Illumination-based methods take the opposite tack: rather than detecting shadow pixels directly, they model the lighting itself. Rooted in the physics of light transport and often built on Retinex theory, which separates an image into reflectance and illumination components, these methods correct the uneven illumination field so that shadows disappear as a byproduct. Classic examples include Zhang&#8217;s restoration system for thick bound documents, which used vertical projection profiles and connected component analysis to find shadow boundaries and was later deployed in a text retrieval project for the National Archives of Singapore. Jung&#8217;s water-filling algorithm treated pixel brightness as a terrain surface, simulated water flooding the valleys to estimate shadow artifacts, and removed them using a Lambertian surface model, achieving an average peak signal-to-noise ratio of 23.38. Other illumination approaches have delivered measurable downstream benefits: a method designed for the Tesseract optical character recognition engine raised character recognition accuracy by an average of 8.1 percent, and a black top-hat transform approach reached 99.54 percent recognition accuracy in OCR experiments, demonstrating that cleaner images translate directly into better machine reading.</p>
<p>The deep learning era has reorganized the field around two architectural philosophies. Single-stage methods perform everything inside one end-to-end network, simplifying workflows and enabling real-time inference. Generative adversarial networks dominated early efforts: DE-GAN used an encoder-decoder generator with a fully convolutional discriminator for document enhancement, UDoc-GAN tackled illumination correction in unpaired images by predicting background light features and improved character error rate by 3 percent, and a lightweight Laplacian pyramid network with input and output attention achieved a 35 percent relative improvement in mean average error while running in real time on mobile hardware. More recent single-stage systems bring in transformers and frequency analysis. FSENet used transformer modules for global color correction and cascaded sparse convolutions to restore high-frequency detail, reaching a PSNR of 28.67 and SSIM of 0.96 on a dataset of 7,620 high-resolution pairs, while ShaDocFormer integrated Otsu thresholding with transformer attention to detect shadow masks and achieved a PSNR of 29.46 on the Real Document Dataset. Diffusion models have entered the arena as well, with frameworks that treat binarization as a reverse diffusion process and others that preserve text features through differentiable modules during deblurring.</p>
<p>Multi-stage methods, by contrast, decompose the problem into sequential subtasks, typically detecting the shadow first and then removing it, which allows each stage to specialize. BEDSR-Net exemplifies the strategy, using a background estimation module to extract global background color and spatial distribution information before removal, and achieving average PSNR and SSIM scores of 37.55 and 0.9534 on synthetic and real images. Zhang&#8217;s color-aware background extraction network paired with a background-guided removal network posted an RMSE of 2.219, PSNR of 37.585, and SSIM of 0.983 on the Real Document Dataset, the strongest combined figures reported in the survey. ShadocNet employed transformers for global shadow remapping followed by pixel-level refinement, and DocDeshadower decomposed images into low- and high-frequency components with Laplacian pyramids before processing each with dedicated transformer branches. Across the benchmark tables, BGShadowNet recorded the highest average SSIM at 0.9431, while DCShadow-Net led on PSNR at 27.9608 and RMSE at 11.3178, though the authors caution that no single method dominates every dataset and that fair comparison on a unified benchmark remains an unfinished task.</p>
<p>Perhaps the survey&#8217;s most valuable contribution is its forward agenda, distilled into nine concrete recommendations. The authors call for large-scale benchmark datasets that better reflect real-world lighting, including multiple light sources and complex occluder shapes, along with multilingual document coverage and fine-grained shadow classification so that removal strategies can be matched to shadow types. They advocate multimodal fusion that combines OCR signals and 3D depth information with image processing, domain adaptation techniques to transfer knowledge from abundant synthetic data to real scenarios, and self-supervised or unsupervised learning to reduce dependence on costly annotations. On the deployment side, they highlight lightweight models for edge devices using compression and clustering, user-guided interactive removal, generative formulations that treat shadow regions as noise for GANs and diffusion models, and explainable methods that reveal why a model decided a region was shadowed. For a problem that most people encounter daily without realizing it has a name, the message of this first comprehensive survey is clear: the humble shadow on a photographed page has become a serious test bed for how well modern computer vision can understand light, physics, and text all at once, and the tools it produces will quietly shape how reliably the world&#8217;s paper heritage enters the digital age.</p>
<p><strong>Subject of Research:</strong> Shadow removal from document images using conventional and neural network-based computer vision methods</p>
<p><strong>Article Title:</strong> A comprehensive survey on shadow removal from document images: datasets, methods, and opportunities</p>
<p><strong>Article References:</strong> Wang, B., Li, C., Zou, W., Zhang, Y., Chen, X., &amp; Chen, C. P. (2025). A comprehensive survey on shadow removal from document images: datasets, methods, and opportunities. <em>Vicinagearth, 2</em>(1), Article 1. <a href="https://doi.org/10.1007/s44336-024-00010-9" rel="noopener noreferrer">https://doi.org/10.1007/s44336-024-00010-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-024-00010-9" rel="noopener noreferrer">10.1007/s44336-024-00010-9</a></p>
<p><strong>Keywords:</strong> document image processing, shadow removal, computer vision, deep learning, datasets, transformers, generative adversarial networks, diffusion models, OCR, image enhancement, Retinex theory, benchmark datasets</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222978</post-id>	</item>
	</channel>
</rss>
