A routine complete blood count, one of the most common medical tests in the world, may soon do far more than flag anaemia or infection. Researchers report in eClinicalMedicine that they have built and validated a machine learning model, called PRISM, which uses standard blood count parameters to identify adults who are likely to carry high-risk clonal haematopoiesis, a silent precursor state that can evolve into leukaemia and other blood cancers. The work addresses a stubborn bottleneck in precision medicine: clonal haematopoiesis is common, but the dangerous form is rare, and until now there has been no validated way to decide who should even be sequenced.
Clonal haematopoiesis occurs when a blood stem cell acquires a somatic mutation and its descendants expand to form a detectable clone, often without any symptoms. Population sequencing shows it in as many as one in five adults over 65, yet transformation to an overt myeloid neoplasm such as acute myeloid leukaemia or myelodysplastic syndrome happens at only 0.5 to 1.0 percent per year. The clinical risk landscape is starkly uneven. According to the clonal haematopoiesis risk score, or CHRS, only about one percent of carriers detected in population-scale sequencing fall into the high-risk category, facing a 10-year blood cancer incidence of 52 percent, while roughly 90 percent are low risk with a 10-year incidence below one percent. Sequencing everyone, therefore, would generate enormous volumes of clinically meaningless results.
The research team, led by investigators drawing on the UK Biobank, hypothesised that the molecular behaviour of dangerous clones leaves fingerprints in ordinary blood measurements. Mutant stem cells alter the production, size, and morphology of red cells, white cells, and platelets, and those changes are captured in the two dozen parameters reported with every complete blood count. To test the idea, they assembled an enormous training dataset of 369,260 UK Biobank participants aged 40 to 69 who had both whole-exome sequencing and blood counts drawn on the same day, then held back 92,316 participants for internal testing.
Defining the target was itself a technical exercise. The team computed a molecular clonal haematopoiesis score for each person, summing points for high-risk mutations in genes such as TP53, JAK2, SRSF2, and RUNX1, for a solitary DNMT3A mutation, for multiple mutations, and for a variant allele fraction of 20 percent or more. A score above 4 captured nearly everyone who later developed myeloid neoplasia or who was high or intermediate risk by CHRS. Participants were then labelled positive if they scored above the threshold and developed blood cancer, negative if they scored at or below it, and equivocal if they scored high but remained cancer-free during follow-up.
The model itself is a balanced random forest, a tree-based classifier chosen after comparing nine architectures on a resampled training set of 3,000 individuals, 1,000 per class. Synthetic minority over-sampling and random under-sampling corrected the extreme class imbalance, since true positives made up less than one tenth of one percent of the cohort. Feature selection was aggressive: highly correlated parameters such as white cell count and haematocrit were removed, the Boruta algorithm discarded sex and several less informative measures, and features inconsistently annotated across datasets were dropped. The final inputs are strikingly mundane: age plus absolute counts of red cells, lymphocytes, monocytes, neutrophils, eosinophils, reticulocytes, and platelets, along with haemoglobin, red cell distribution width, platelet crit, mean corpuscular volume, and giant platelets.
Performance in the UK Biobank test set was strong where it mattered most. The area under the receiver operating characteristic curve for the positive class reached 0.89, and sensitivity for high-risk clonal haematopoiesis was 0.73, with the three missed cases all under 60 with entirely normal blood counts. More strikingly, positive classification captured 100 percent of the 42 individuals deemed high risk by CHRS, and the overall negative predictive value was 0.993, meaning a negative result reliably rules out clinically significant clonal haematopoiesis. The model was intentionally permissive of false positives, accepting that sequencing would resolve them, and achieved one detected high-risk case for every 247 people flagged, compared with one per 2,500 under a screen-everyone strategy.
External validation demonstrated genuine portability. In a clinical cohort of 2,233 patients evaluated in Dana-Farber haematology clinics, where cytopenias were far more common, PRISM correctly classified 16 of 18 positive-label individuals, an AUC of 0.74 and sensitivity of 0.89, with a negative predictive value of 0.968. In the All of Us Research Program cohort of 14,674 Americans, the positive-class AUC was 0.87 with perfect sensitivity across ten cases and a negative predictive value of 0.998. Across all three cohorts, every CHRS-defined high-risk individual received a positive PRISM label, and intermediate-risk carriers almost never received a negative one. Benchmarking against the clonal cytopenia risk score showed that 18 of 19 high-risk clonal cytopenia cases were also flagged as positive.
The downstream consequences were equally telling. Among UK Biobank participants with clonal haematopoiesis who later developed myeloid neoplasia, 63 percent had been labelled positive and 25 percent equivocal, and the misclassified negatives were overwhelmingly low risk by post-sequencing CHRS. Ten-year cumulative incidence of myeloid neoplasia was 1.44 percent in the positive class versus 0.11 percent in the negative class, a hazard ratio of 8.81 that outperformed age, cytopenia, mean corpuscular volume, and red cell distribution width in multivariable models. Because clonal haematopoiesis also drives cardiovascular disease, the team checked whether negative labels would exclude people at cardiac risk; reassuringly, 10-year cardiovascular risk was lowest, at 4.6 percent, in the PRISM-negative group, compared with 15.7 percent in the positive class.
The authors are careful about scope. PRISM is a pre-sequencing triage tool, not a diagnostic test, and its equivocal class, which includes people with abnormal counts but no dangerous clones, cannot be resolved from a single blood draw; serial sampling and clinical correlation remain essential. The UK Biobank’s predominantly European ancestry limits generalisability, and specialised settings such as chemotherapy exposure or inherited bone marrow failure would need context-specific recalibration. Still, if prospective validation succeeds, the implications are considerable: haematology clinics could reassure worried patients with a negative label, transplant programmes could screen donors cheaply before sequencing, and health systems could cut unnecessary molecular testing by 14 to 55 percent while sparing low-risk individuals the psychological and financial burden of overdiagnosis. A model built from the humble blood count may become the gatekeeper for the genomic era of blood cancer prevention.
Subject of Research: A machine learning model using complete blood count parameters to prioritise sequencing for high-risk clonal haematopoiesis
Article Title: Derivation of a blood count-based machine learning model to prioritise sequencing for high-risk clonal haematopoiesis: a retrospective cohort study
Article References: Nandi, R., Zhao, K., Pershad, Y., Khondoker, N. N., Vlasschaert, C., Uddin, M. M., Natarajan, P., DeAngelo, D. J., Luskin, M. R., Lindsley, R. C., Uno, H., Shimony, S., Bick, A. G., & Weeks, L. D. (2026). Derivation of a blood count-based machine learning model to prioritise sequencing for high-risk clonal haematopoiesis: a retrospective cohort study. eClinicalMedicine, 100, Article 104238. https://doi.org/10.1016/j.eclinm.2026.104238
Image Credits: AI Generated
DOI: 10.1016/j.eclinm.2026.104238
Keywords: clonal haematopoiesis, machine learning, complete blood count, myeloid neoplasia, leukaemia risk, CHRS, UK Biobank, risk prediction, sequencing, premalignant conditions, cardiovascular disease, precision medicine
Cite Scienmag News
Nathaniel Bowman. (October 2, 2026). Routine Blood Counts Could Reveal Who Carries a Hidden Blood Cancer Risk. Scienmag. https://scienmag.com/routine-blood-counts-could-reveal-who-carries-a-hidden-blood-cancer-risk/
Nathaniel Bowman. "Routine Blood Counts Could Reveal Who Carries a Hidden Blood Cancer Risk." Scienmag, 2 October 2026, https://scienmag.com/routine-blood-counts-could-reveal-who-carries-a-hidden-blood-cancer-risk/. Accessed 2 October 2026.
Nathaniel Bowman. "Routine Blood Counts Could Reveal Who Carries a Hidden Blood Cancer Risk." Scienmag. October 2, 2026. https://scienmag.com/routine-blood-counts-could-reveal-who-carries-a-hidden-blood-cancer-risk/

