Computed tomography has transformed modern medicine by allowing clinicians to see inside the body without surgery, but the technology has a persistent weakness: metal can turn a detailed scan into a maze of bright streaks and dark bands. Dental fillings, hip replacements, surgical screws, vascular stents and other high-density objects can severely disrupt the X-ray measurements used to reconstruct an image. The resulting metal artifacts may conceal tissue, distort anatomical boundaries and complicate diagnosis or image-guided treatment. Now, researchers at ShanghaiTech University in China have introduced a new artificial-intelligence framework designed to reduce these artifacts without relying on paired examples of corrupted and pristine scans. Their method, called INR-DR, combines the physical rules of CT imaging with the image-forming capabilities of diffusion models, offering a new approach to one of medical imaging’s most difficult reconstruction problems.
The work, published in Quantitative Biology, addresses a central challenge in conventional metal artifact reduction. When X-rays pass through a metal object, they can undergo beam hardening, meaning that lower-energy photons are absorbed more strongly than higher-energy photons. At the same time, photon starvation can occur when so few X-rays reach the detector that the measurements become extremely noisy or unreliable. These effects create missing, distorted or inconsistent regions in the projection data, often called sinograms. Standard algorithms may fill metal-affected regions through interpolation, while deep-learning systems can learn to replace corrupted measurements from large collections of training examples. Yet interpolation may create artificial discontinuities, and supervised networks generally require matched metal-corrupted and metal-free images—datasets that are difficult to acquire from real patients.
INR-DR takes a different route. Instead of directly repairing the damaged projection measurements, it represents the scanned anatomy as a continuous function of spatial coordinates. This approach belongs to the family of implicit neural representations, in which a neural network learns to map a location in space to an estimated image intensity. Rather than storing the reconstruction only as a fixed grid of pixels or voxels, the model describes the underlying image as a coordinate-based field that can be queried throughout the volume. The researchers combine this representation with multiresolution hash encoding, a technique that allows the network to capture both broad anatomical structure and fine details while keeping the computational representation relatively compact. The result is an adaptable image model optimized for each individual scan.
The most important constraint comes from CT physics. A reconstructed image is useful only if, when passed through a forward projection model, it produces measurements that agree with the reliable observations collected by the scanner. INR-DR therefore uses a differentiable forward model to simulate the CT acquisition process. During optimization, the image representation is adjusted so that its projections remain consistent with measurements considered trustworthy, while the most severely corrupted metal-trace regions are not treated as if they were accurate evidence. This strategy avoids the abrupt boundary between measured and artificially filled projection data that can generate secondary streaks. Instead of asking a network to guess missing information in isolation, the framework searches for an image that satisfies the available measurements and remains plausible under the imaging system’s geometry.
To guide that search, the team incorporated a pretrained unconditional diffusion model. Diffusion models have become widely known for generating realistic images by gradually transforming noise into structured visual content, but in INR-DR the model is not used as a direct image generator. It acts as a regularizer, providing a statistical preference for anatomically plausible structures during reconstruction. The implicit neural representation is periodically guided through one-step denoising, encouraging it to move away from implausible patterns created by the incomplete or corrupted data. This distinction is important: the diffusion model does not simply overwrite the scan with an imagined anatomy. Instead, it works alongside the CT data-fidelity term, helping stabilize the ill-posed reconstruction while the physics-based component protects information supported by the measurements.
The framework consequently alternates between two forms of reasoning. One stage emphasizes data fidelity, updating the implicit image so that it agrees with the reliable CT observations and the scanner’s forward model. The other applies diffusion-based regularization, steering the reconstruction toward the kinds of structures learned from image data. These objectives can pull in different directions. Strictly following damaged measurements may preserve artifacts, while relying too heavily on a learned prior could smooth away real anatomical details or introduce structures that were never present. INR-DR is designed to balance those risks. Its multiresolution encoding helps retain local edges and small features, while the diffusion prior supplies broader anatomical plausibility when the measurements alone are insufficient.
In experiments using simulated scans and clinical dental CT data, the method reduced conspicuous streak artifacts associated with metal implants while preserving surrounding tissue structure. The researchers evaluated cases involving large, medium and small metal objects, situations that can produce very different patterns of projection corruption. According to the study, INR-DR outperformed traditional reconstruction and metal-artifact-reduction techniques as well as supervised learning baselines across these settings. The reported findings suggest that the method can remain effective when implant size and configuration change, an important consideration because real-world patients rarely match the standardized examples used to train many supervised systems. The clinical dental results are particularly relevant because fillings, crowns and other dental materials often occupy complex positions close to bone and soft tissue.
Ablation experiments further indicated that the method’s components make complementary contributions. Removing or weakening the physics-based data-consistency mechanism can allow the reconstruction to drift from the measurements, while omitting the diffusion prior can leave the model more vulnerable to noise and severe missing information. The combination provides a form of mutual restraint: CT physics limits unsupported invention, and the learned prior helps fill in ambiguity without resorting to crude interpolation. This is also why the approach may be attractive for other inverse problems. Sparse-view CT, low-dose imaging and related reconstruction tasks all require the recovery of an image from incomplete, noisy or degraded measurements. A framework that can adapt to a single case without requiring a perfectly paired training dataset could be useful in settings where collecting comprehensive reference scans is impractical or ethically difficult.
The study does not, however, eliminate the barriers between a promising reconstruction method and routine clinical use. INR-DR is optimized for each scan, and that per-case process may be slower than applying a conventional trained neural network in a single forward pass. The method must also be tested across broader patient populations, scanner manufacturers, acquisition protocols, anatomical regions and implant materials. In medical imaging, visual improvement alone is not enough: researchers must establish that diagnostically important structures are preserved, that pathological findings are not altered and that the method behaves reliably when the data differ from the conditions used during development. Even so, the ShanghaiTech team’s work points toward a potentially powerful direction for artifact reduction—one in which artificial intelligence does not replace the physics of imaging, but works within it. By combining a continuous neural image representation, differentiable CT modeling and diffusion-based anatomical guidance, INR-DR offers a new blueprint for reconstructing clearer images from measurements that metal has made unreliable.
Subject of Research: Not applicable
Article Title: Diffusion model-regularized implicit neural representation for computed tomography metal artifact reduction
Web References: https://doi.org/10.1002/qub2.70031
References: Quantitative Biology, “Diffusion model-regularized implicit neural representation for computed tomography metal artifact reduction,” DOI: 10.1002/qub2.70031
Image Credits: Higher Education Press
Keywords: computed tomography, metal artifact reduction, diffusion models, implicit neural representation, medical imaging, deep learning, image reconstruction, CT physics, dental CT, artificial intelligence

