Henna has colored human skin and hair for thousands of years, and its natural dye, lawsone, is prized precisely because it is gentle, plant-derived, and safe. Yet the modern cosmetics market has quietly corrupted this ancient product. To meet consumer demand for faster, darker results, suppliers increasingly blend henna with para-phenylenediamine, or PPD, a synthetic aromatic amine that produces the deep black shade impossible to achieve with pure henna alone. The problem is that PPD is a potent skin sensitizer, and so-called black henna has been linked to severe allergic reactions, blistering contact dermatitis, permanent scarring, and in extreme cases systemic toxicity. A new study published in Smart Agricultural Technology now offers a way to catch this adulteration in seconds, using nothing more than light and machine learning.
Researchers led by Omid Farhangi and Ahmad Banakar of Tarbiat Modares University, together with colleagues in Iran, set out to replace the slow, destructive laboratory tests currently used to detect PPD with a fully non-destructive optical method. Conventional approaches rely on high-performance liquid chromatography or gas chromatography–mass spectrometry, which deliver excellent accuracy but demand lengthy sample preparation, hazardous solvents, expensive instruments, and trained specialists. For market surveillance, customs checkpoints, and factory quality control, that combination is impractical. The team’s alternative was to shine ultraviolet, visible, and near-infrared light onto henna powder, record the reflected spectrum, and let algorithms translate the optical fingerprint into an exact PPD concentration.
The experimental design was deliberately rigorous. Pure henna leaves were collected from the major cultivation regions of Kerman and Sistan-Baluchestan in Iran, ground under controlled laboratory conditions, and supplemented with market samples from Yazd province. Laboratory-grade PPD was then mixed into the powder at seven levels ranging from 0 to 12 percent by weight, a range chosen to mirror the contamination levels actually reported in adulterated commercial henna. Eleven independently prepared samples were made at each concentration, yielding 77 physical samples in total. From each one, five reflectance spectra were recorded at different spots and averaged before any data splitting, a crucial step that prevented replicate measurements from leaking information between training and test sets.
The spectroscopy itself relied on an Ocean Optics USB2000 spectrometer covering 179 to 872 nanometers, illuminated by a UV-B lamp for the ultraviolet region and a halogen source for the visible and near-infrared. Light was delivered and collected through a bifurcated fiber probe, with samples held at a fixed thickness inside a black box to exclude ambient light. The raw spectra revealed a striking feature: a sharp, strong reflectance peak at 380 to 390 nanometers that grew steadily in intensity as PPD concentration increased. This peak arises from the conjugated pi-electron system of PPD’s aromatic amine structure, which interacts strongly with ultraviolet radiation, whereas lawsone’s quinonoid structure responds far more weakly in this window.
Principal component analysis confirmed what the spectra suggested. In the UV region, samples clustered cleanly by PPD concentration along the first principal component, demonstrating that the optical features were chemically specific and systematically tied to contamination. In the visible and near-infrared region, by contrast, no meaningful clustering appeared; the samples lined up along a single trend driven by overall reflectance offset, a physical scattering effect rather than a chemical signature. This distinction proved predictive of the modeling results: the UV region would reward nonlinear machine learning, while the visible region would struggle across the board.
The team compared five regression algorithms: partial least squares regression, decision tree, random forest, support vector regression, and an artificial neural network. On full-spectrum UV data, the decision tree with no preprocessing achieved a test coefficient of determination of 0.9376 with a residual prediction deviation of 4.14, edging out PLSR’s 0.9339. Random forest and support vector regression also performed respectably. In the visible region, performance dropped sharply, with the best model, PLSR with detrending, reaching only 0.7583 on the test set, still above the conventional threshold of 2.0 RPD that indicates a very good model but far below UV-region accuracy.
The study’s most innovative stage involved six metaheuristic optimization algorithms, including genetic algorithms, particle swarm optimization, ant colony optimization, and learning automata, each tasked with selecting the 15 most informative wavelengths from hundreds of spectral variables. Learning automata emerged as the clear winner in both spectral regions, achieving the lowest error and highest correlation in the UV range. Its advantage stems from an adaptive probability-update mechanism that rewards useful wavelengths and punishes uninformative ones, allowing it to escape local optima that trapped faster competitors. The genetic algorithm, despite the shortest runtime of nine seconds, performed worst in the UV region, highlighting the trade-off between computational speed and convergence quality.
Feature selection delivered a dramatic payoff. After reducing the data to just 15 wavelengths, more than 90 percent of the original spectral variables discarded, random forest with Gaussian filter preprocessing leapt from a test R-squared of 0.9186 to 0.9779, with an outstanding RPD of 6.99, the best result of the entire study. The artificial neural network followed closely with an R-squared of 0.9476 and RPD of 4.51. The improvement shows that ensemble methods, often destabilized by noisy and collinear spectral variables, can unlock their full potential when fed only the most diagnostic features. Interestingly, the learning automata selected wavelengths above 400 nanometers in addition to the characteristic UV peak, indicating that scattering-related changes in the visible range carry complementary information about PPD content.
The practical implications extend well beyond henna. Because the entire analysis depends on only 15 key wavelengths, manufacturers could build dedicated, low-cost multispectral sensors that operate at those specific optical points, deploying them on production lines, at customs checkpoints, and in retail settings for real-time screening. Such devices would align with the broader trend toward miniaturized analytical systems in the food and cosmetic industries. The authors caution, however, that their dataset was generated under controlled laboratory conditions, and real commercial samples may contain interfering compounds, variable particle sizes, and moisture that could affect spectral responses. Validation against chromatographic reference methods on diverse commercial products, calibration transfer protocols for portable instruments, and cross-dataset testing of the selected wavelengths remain necessary next steps.
The work also opens the door to broader applications of the same framework. Future research could extend the approach to simultaneous detection of multiple adulterants, including other aromatic amines, heavy metals, and plant-based fillers, or employ deep learning architectures such as one-dimensional convolutional neural networks and data fusion strategies combining UV-Vis-NIR spectroscopy with Raman techniques. For now, the study stands as a compelling demonstration that an ancient cosmetic product can be protected by thoroughly modern tools: a beam of light, a fiber-optic probe, and algorithms smart enough to read the chemical secrets hidden in a reflection. In a market where a single contaminated tattoo can cause lifelong injury, that combination may prove invaluable.
Subject of Research: Non-destructive detection of para-phenylenediamine adulteration in henna using UV–Vis–NIR spectroscopy and machine learning
Article Title: Non-destructive quantification of para-phenylenediamine in henna using UV–Vis–NIR spectroscopy and machine learning
Article References: Farhangi, O., Banakar, A., Ebadi, M.-T., & Mahdavian, A. (2026). Non-destructive quantification of para-phenylenediamine in henna using UV–Vis–NIR spectroscopy and machine learning. Smart Agricultural Technology, 15, Article 102621. https://doi.org/10.1016/j.atech.2026.102621
Image Credits: AI Generated
DOI: Not provided
Keywords: henna, para-phenylenediamine, UV-Vis-NIR spectroscopy, machine learning, food adulteration, cosmetic safety, chemometrics, wavelength selection, learning automata, random forest, artificial neural network, quality control
Cite Scienmag News
Blake Davidson. (October 10, 2026). Machine Learning Spots Toxic Dye in Henna With Light Alone. Scienmag. https://scienmag.com/machine-learning-spots-toxic-dye-in-henna-with-light-alone/
Blake Davidson. "Machine Learning Spots Toxic Dye in Henna With Light Alone." Scienmag, 10 October 2026, https://scienmag.com/machine-learning-spots-toxic-dye-in-henna-with-light-alone/. Accessed 10 October 2026.
Blake Davidson. "Machine Learning Spots Toxic Dye in Henna With Light Alone." Scienmag. October 10, 2026. https://scienmag.com/machine-learning-spots-toxic-dye-in-henna-with-light-alone/

