When ionizing radiation tears through a cell, it can snap both strands of the DNA double helix at once. These double-strand breaks are among the most dangerous injuries a genome can sustain, and cells respond by summoning repair proteins to the damage sites, which appear under the fluorescence microscope as tiny bright spots known as ionizing radiation-induced repair foci, or IRIFs. Counting these foci has long been a cornerstone of radiation biology: the number of foci per nucleus reflects the dose received and the pace at which DNA is stitched back together. But the counting itself has been a stubborn bottleneck, traditionally performed either by hand or with classical image-analysis pipelines that are notoriously sensitive to variations in illumination, noise, and cell morphology. A new study published in BMC Bioinformatics by Saeed Ullah, Gianluca Valentino, and colleagues at the University of Malta, together with Sylvain V. Costes of the University of Pittsburgh, now shows that a deep convolutional neural network trained entirely on synthetic images can reliably detect and quantify these repair foci in real microscopy data, opening a path toward high-throughput biodosimetry.
The team’s central challenge was one that plagues much of deep learning in biology: annotated training data are scarce and expensive. Manually labeling repair foci in fluorescence micrographs demands expert judgment, and even experts disagree, particularly when foci cluster densely or sit against bright background texture. Rather than curating thousands of painstakingly annotated real images, the researchers took a different route. They generated 95,000 synthetic image-mask pairs designed to replicate the signal characteristics, noise profiles, and morphological features of genuine fluorescence microscopy data. In these simulated images, the ground truth — the exact location of every focus — is known perfectly by construction, because the images are rendered from the labels rather than annotated afterward. This strategy, sometimes called simulation-to-reality transfer, sidesteps the annotation bottleneck entirely, provided the synthetic data capture enough of the statistical structure of real images for a network to generalize.
The architecture at the heart of the method is a modified U-Net, a convolutional network originally designed for biomedical image segmentation. U-Net’s signature design is its U-shaped encoder-decoder structure: a contracting path of convolutional and pooling layers that captures context by progressively downsampling the image, followed by an expanding path that upsamples feature maps back to full resolution, with skip connections shuttling fine-grained spatial information from encoder to decoder. The result is a network that can label every pixel — a task called semantic segmentation — while preserving the precise boundaries of small objects. For repair foci, which are often only a few pixels across, that pixel-level precision is essential. The modifications the authors introduced were aimed at the specific demands of foci detection, where the balance between catching faint, dim spots and avoiding spurious detections on noise determines the quality of the final count.
The performance figures on synthetic test data are striking. The trained model achieved a mean Dice coefficient of 0.995 and an Intersection over Union of 0.993. Both metrics measure the overlap between predicted and true segmented regions: the Dice coefficient is the harmonic mean of precision and recall at the pixel level, while IoU is the ratio of the intersection of the predicted and true masks to their union. Values this close to 1.0 indicate near-perfect agreement on the simulated images. But the authors were careful not to rest on synthetic laurels, since flawless performance on data drawn from the same distribution as training material says little about usefulness in the real world. The decisive test came from real fluorescence microscopy images of mouse fibroblasts, drawn from the NASA GeneLab OSD-366 experiment, encompassing multiple radiation types and doses, with 345 images manually annotated by experts to serve as ground truth.
On that real-world dataset, the model’s segmentation performance remained strong, with Dice coefficients ranging from 0.72 to 0.88 on average across the radiation categories. The drop from the near-perfect synthetic scores is expected and, in context, encouraging: it reflects the domain gap between simulated and genuine microscopy, including projection artifacts and biological variability that no simulator fully reproduces. More important for practical biodosimetry is the foci detection stage, where segmented regions are converted into counts per nucleus. Here the model demonstrated high sensitivity, with recall values between 0.95 and 0.98 across categories, and precision between 0.82 and 0.97. Aggregated over the whole dataset, the system achieved an overall recall of 0.97 and precision of 0.89 — meaning it missed very few genuine foci while keeping false positives within acceptable limits. In a dosimetry context, that asymmetry matters: missing a focus underestimates dose, while a modest rate of false positives is easier to tolerate and statistically correct for.
The study also probed how well the network’s counts track the true numbers of foci per nucleus, using confusion matrices that compare predicted counts against ground-truth counts. For X-ray exposures, the agreement was tight, with Pearson correlation coefficients of 0.94 for 0.1 Gy and 0.97 for 1.0 Gy, and exact matches between predicted and true counts in 73 to 92 percent of images even when the full, untruncated count range was analyzed. The picture was more complicated for iron ions, a form of heavy-ion radiation of particular interest in space radiation protection. For 0.3 Gy and 0.82 Gy iron exposures, correlations were 0.85 and 0.93 respectively, but exact-match fractions fell to 35 to 57 percent, with predictions spreading into neighboring count categories. The authors traced this to the character of heavy-ion damage: iron particles produce dense, clustered tracks of damage that generate larger and more irregular foci, along with long tails of high counts — up to 20 foci per nucleus at 0.82 Gy — that make precise counting intrinsically harder. Crucially, the researchers verified that this leakage was a genuine property of the data rather than an artifact of the count cap they applied for analysis.
The implications extend well beyond a single laboratory’s image analysis pipeline. Biodosimetry — estimating the radiation dose a person has absorbed by measuring biological endpoints — is a critical capability for radiation protection, for triage after radiological accidents, and for astronaut health during deep-space missions, where cosmic rays and solar particle events deliver mixed radiation fields that include the heavy ions that proved hardest here. Traditional foci counting, whether manual or script-based, is sensitive to image variability and difficult to reproduce across laboratories, which has limited throughput and comparability. A deep learning model that generalizes from synthetic training data to real images offers reproducibility by construction: once trained and validated, it applies the same criteria to every image, and it can process large image sets at a scale no human counter could match. The work was financed by Xjenza Malta through the FUSION R&I Space Upstream Programme under Project DeepAFQ, underscoring the space-health motivation behind the research.
The authors are candid about the limits of what they have shown. The validation rested on a real-world dataset of mouse fibroblast images from a single experiment, covering specific radiation types and doses at defined time points after exposure. Generalization across additional time points, cell types, imaging platforms, and data sources remains to be demonstrated before the approach can be deployed as a validated biodosimetry tool. The synthetic-to-real gap, while bridged impressively here, is not closed — the appendix shows representative cases where projection artifacts, such as ringing halos introduced by maximum intensity projection of image stacks, caused the model to predict structures absent from the ground truth. Such failure modes are informative: they define exactly where future training data and simulation refinements should focus, and they remind practitioners that automated counts still benefit from expert oversight in ambiguous cases.
Even with those caveats, the study adds to a growing body of evidence that synthetic training data can unlock deep learning applications where real annotated data are the limiting resource — a pattern increasingly seen across microscopy, pathology, and astronomy. For radiation biology specifically, the prospect of robust, automated IRIF quantification could accelerate studies of DNA repair kinetics, improve the statistical power of radiobiological experiments, and eventually support personalized medicine decisions where individual radiation sensitivity matters. It could also feed directly into the operational needs of space agencies assessing crew exposure. The Malta-led team has shown that a network that has never seen a real cell can still learn to read the marks radiation leaves on the genome — and count them with an accuracy that begins to rival the experts who trained the ground truth it was judged against.
Subject of Research: Deep learning-based detection and quantification of ionizing radiation-induced DNA repair foci in fluorescence microscopy images
Article Title: Ionizing radiation-induced repair foci identification using deep convolutional neural networks
Article References: Ullah, S., Anu, R. I., Borg, J., Costes, S. V., Borg, J., & Valentino, G. (2026). Ionizing radiation-induced repair foci identification using deep convolutional neural networks. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06662-2
Image Credits: AI Generated
DOI: 10.1186/s12859-026-06662-2
Keywords: DNA double-strand breaks, ionizing radiation-induced repair foci, deep learning, convolutional neural networks, U-Net, semantic segmentation, fluorescence microscopy, biodosimetry, synthetic training data, radiation biology, space health, NASA GeneLab
Cite Scienmag News
Drew Townsend. (October 7, 2026). AI Trained on Fake Images Learns to Spot Real DNA Damage from Radiation. Scienmag. https://scienmag.com/ai-trained-on-fake-images-learns-to-spot-real-dna-damage-from-radiation/
Drew Townsend. "AI Trained on Fake Images Learns to Spot Real DNA Damage from Radiation." Scienmag, 7 October 2026, https://scienmag.com/ai-trained-on-fake-images-learns-to-spot-real-dna-damage-from-radiation/. Accessed 7 October 2026.
Drew Townsend. "AI Trained on Fake Images Learns to Spot Real DNA Damage from Radiation." Scienmag. October 7, 2026. https://scienmag.com/ai-trained-on-fake-images-learns-to-spot-real-dna-damage-from-radiation/

