Every year, ischemic stroke strikes millions of people, and one of its most treacherous warning signs hides in plain sight: the carotid arteries, the paired vessels that carry blood up each side of the neck to the brain. When fatty plaques build up in these arteries, doctors can see them on ultrasound or MRI scans. But the real question is not whether a plaque exists — it is whether that plaque is stable and harmless or fragile and primed to rupture, showering debris into the brain. Now, a systematic review and meta-analysis published in BMC Medical Imaging by Dong Ma of the Fifth Affiliated Hospital of Southern Medical University and colleagues has taken stock of how well artificial intelligence can make that life-or-death distinction, and the results are both encouraging and sobering.
The research team systematically searched the literature through June 12, 2026, ultimately including twenty-eight studies that applied either radiomics or deep learning to carotid plaque imaging. Radiomics is the practice of extracting hundreds of quantitative features from medical images — texture patterns, shape descriptors, intensity statistics — that are invisible to the human eye. Deep learning goes a step further: convolutional neural networks learn their own hierarchical representations directly from pixel data, bypassing handcrafted features entirely. Both approaches promise something traditional imaging assessment cannot deliver: an objective, reproducible, automated verdict on plaque vulnerability, free from the inter-observer variability that plagues subjective visual readings.
The headline numbers are striking. For assessing plaque vulnerability itself, radiomics-based models achieved a pooled area under the curve (AUC) of 0.84, with a 95 percent confidence interval of 0.80 to 0.88. Deep learning models performed even better, reaching a pooled AUC of 0.92 (95 percent CI: 0.89 to 0.94). For the ultimate clinical question — predicting which patients will actually suffer a stroke — pooled performance was an AUC of 0.84 (95 percent CI: 0.74 to 0.95), though the wide confidence interval hints at substantial heterogeneity across studies. An AUC of 0.92 means that in roughly nine out of ten paired comparisons, the model correctly ranks a vulnerable plaque as riskier than a stable one, a level of discrimination that would be genuinely transformative if it held up in everyday clinical practice.
The meta-analysis also revealed important differences between imaging modalities. MRI-based radiomics showed the most consistent diagnostic performance across studies, with an I² statistic of 0.00 percent — meaning essentially none of the observed variation between studies could be attributed to heterogeneity rather than chance. That kind of homogeneity is rare in medical AI literature and suggests that MRI features of plaque composition, such as intraplaque hemorrhage and lipid-rich necrotic cores, translate reliably across scanners and patient populations. Ultrasound, the most widely available and cheapest modality, posted the highest numerical performance among radiomics subgroups, with an AUC of 0.87. Given that carotid ultrasound is already the first-line screening tool in most of the world, embedding validated AI analysis into that existing workflow could be the fastest route to the bedside.
Methodological rigor was a central concern of the review. The authors appraised study quality using QUADAS-AI, a tool specifically designed to assess diagnostic accuracy studies involving artificial intelligence, and the Radiomics Quality Score 2.0 (RQS 2.0), which grades how thoroughly radiomics studies address feature reproducibility, biological validation, and clinical usefulness. These frameworks matter because the AI diagnostic literature is littered with studies that look impressive on paper but suffer from small samples, optimistic internal validation, and no external testing. The review’s use of hierarchical summary receiver operating characteristic (HSROC) curves and leave-one-out sensitivity analyses reflects a growing insistence that pooled estimates of AI performance be interrogated as carefully as the underlying models themselves.
Behind the technical vocabulary lies a straightforward clinical story. A vulnerable carotid plaque is one with a thin or ruptured fibrous cap, a large lipid core, intraplaque hemorrhage, or active inflammation — features that make rupture and downstream embolization likely. Conventional imaging grading, such as stenosis percentage, captures only how narrowed an artery is, yet many strokes arise from plaques that never severely narrow the vessel. Quantitative image analysis can, in principle, detect these compositional red flags automatically. Machine learning classifiers commonly deployed in the reviewed studies — logistic regression, support vector machines, random forests for radiomics features, and convolutional or Bayesian convolutional neural networks for raw image data — were trained to map these image signatures onto histologically confirmed plaque vulnerability or observed clinical outcomes.
Interpretability is where the field’s enthusiasm meets its hardest test. A neural network that flags a plaque as dangerous is of limited use to a vascular surgeon if it cannot explain why. The review notes that techniques such as SHAP (SHapley Additive exPlanations) are increasingly being used to attribute model decisions to specific image features, and Bayesian convolutional neural networks offer a way to quantify the model’s own uncertainty. Decision curve analysis, another tool highlighted in the literature, evaluates whether using a model to guide clinical decisions actually yields net benefit compared with treating everyone or no one. These are exactly the kinds of validation steps that regulatory bodies and clinicians will demand before an algorithm earns a place in the reading room.
Standardization poses an equally formidable barrier. Radiomics features are notoriously sensitive to scanner manufacturer, acquisition protocol, image resolution, and reconstruction settings — a texture feature measured on one hospital’s MRI machine may not mean the same thing on another’s. Initiatives such as the Image Biomarker Standardization Initiative (IBSI) aim to harmonize feature definitions and processing pipelines, and reporting standards like TRIPOD-AI push for transparent documentation of model development and validation. The review’s authors are explicit that despite good diagnostic efficacy in research settings, challenges in standardization and model interpretability currently limit clinical translation. In other words, the gap between a promising retrospective study and a deployable clinical tool remains wide.
What would it take to close that gap? The trajectory suggested by this review points toward large, prospective, multi-center trials in which AI-derived plaque scores are tested against real stroke outcomes in consecutive patients, across diverse scanners and populations. It points toward externally validated models rather than internally cross-validated ones, and toward integration with existing clinical risk tools rather than replacement of them. The wide confidence interval around the stroke prediction AUC of 0.84 signals that this prognostic application, the most clinically consequential of all, is the least mature. Plaque characterization is closer to readiness; predicting an individual patient’s future stroke remains the field’s grand challenge.
Still, the direction of travel is unmistakable. Twenty-eight studies, pooled AUCs approaching 0.9 for deep learning, and a perfectly homogeneous MRI radiomics signal together describe a field that has moved from proof-of-concept to the threshold of clinical relevance. For patients walking around with ticking time bombs in their neck arteries, the promise is profound: a routine scan, read by an algorithm, that identifies the plaque most likely to cause a stroke years before it does. The review by Ma and colleagues, funded in part by the Southern Medical Imaging Alliance scientific research fund of the Guangdong Radiation Protection Association, makes clear that the algorithms are nearly there. The harder work — harmonizing the images, opening the black boxes, and proving the predictions in prospective trials — is what stands between today’s impressive research numbers and tomorrow’s saved lives.
Subject of Research: AI-based radiomics and deep learning assessment of carotid plaque vulnerability for ischemic stroke risk prediction
Article Title: Carotid plaque vulnerability assessment and stroke risk prediction based on radiomics and deep learning: a systematic review and meta-analysis
Article References: Ma, D., Chen, J., Li, Z., Lu, F., Cui, Y., Chen, J., Wang, S., & Zhou, T. (2026). Carotid plaque vulnerability assessment and stroke risk prediction based on radiomics and deep learning: a systematic review and meta-analysis. BMC Medical Imaging. https://doi.org/10.1186/s12880-026-02824-z
Image Credits: AI Generated
DOI: 10.1186/s12880-026-02824-z
Keywords: radiomics, deep learning, carotid plaque, stroke, ischemic stroke, plaque vulnerability, medical imaging, artificial intelligence, MRI, ultrasound, meta-analysis, predictive medicine
Cite Scienmag News
Ophelia Keating. (October 9, 2026). AI Reads the Neck’s Deadliest Plaques, but the Clinic Isn’t Ready Yet. Scienmag. https://scienmag.com/ai-reads-the-necks-deadliest-plaques-but-the-clinic-isnt-ready-yet/
Ophelia Keating. "AI Reads the Neck’s Deadliest Plaques, but the Clinic Isn’t Ready Yet." Scienmag, 9 October 2026, https://scienmag.com/ai-reads-the-necks-deadliest-plaques-but-the-clinic-isnt-ready-yet/. Accessed 9 October 2026.
Ophelia Keating. "AI Reads the Neck’s Deadliest Plaques, but the Clinic Isn’t Ready Yet." Scienmag. October 9, 2026. https://scienmag.com/ai-reads-the-necks-deadliest-plaques-but-the-clinic-isnt-ready-yet/

