<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>cytomorphology &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/cytomorphology/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 12:54:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>cytomorphology &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Can Spot Leukemia Cells on Blood Smears, But One Blast Type Keeps Slipping Through</title>
		<link>https://scienmag.com/ai-can-spot-leukemia-cells-on-blood-smears-but-one-blast-type-keeps-slipping-through/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 12:54:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[acute myeloid leukemia]]></category>
		<category><![CDATA[acute myeloid leukemia diagnosis]]></category>
		<category><![CDATA[AI sensitivity and specificity in blood diagnostics]]></category>
		<category><![CDATA[AML subtypes detection accuracy]]></category>
		<category><![CDATA[automated blood cell classification]]></category>
		<category><![CDATA[blood smear]]></category>
		<category><![CDATA[blood smear analysis]]></category>
		<category><![CDATA[blood smear image analysis]]></category>
		<category><![CDATA[challenges in leukemia cell identification]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[cytomorphology]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Deep learning in hematology]]></category>
		<category><![CDATA[diagnostic accuracy]]></category>
		<category><![CDATA[immature white blood cell classification]]></category>
		<category><![CDATA[leukemia cell detection using AI]]></category>
		<category><![CDATA[limitations of AI in leukemia diagnosis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for hematological disorders]]></category>
		<category><![CDATA[meta-analysis]]></category>
		<category><![CDATA[promyelocytes]]></category>
		<category><![CDATA[QUADAS-AI]]></category>
		<category><![CDATA[white blood cells]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222762</guid>

					<description><![CDATA[A new systematic review and meta-analysis finds AI classifiers achieve high specificity for immature white blood cell subtypes in acute myeloid leukemia, but promyelocyte sensitivity remains a critical weakness across all model paradigms.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has been quietly transforming the way pathologists read blood smears, and a new systematic review has now delivered the most rigorous accounting yet of how well these algorithms actually perform when asked to identify the immature white blood cells that define acute myeloid leukemia. The verdict is a fascinating mix of triumph and caution. Across six eligible studies, AI classifiers achieved remarkably high specificity for all four clinically critical immature cell subtypes examined, meaning they rarely mistake a healthy mature cell for a dangerous leukemic blast. Yet sensitivity told a more troubling story, particularly for one subtype that appears to be systematically slipping past even the most sophisticated deep learning architectures.</p>
<p>The review, published in Machine Learning with Applications, focused on the four immature white blood cell subtypes central to AML diagnosis under the French-American-British classification system: erythroblasts, monoblasts, myeloblasts, and promyelocytes. These cells are notoriously difficult to distinguish even for trained human experts, because adjacent maturation stages share overlapping nuclear size, chromatin texture, and cytoplasmic characteristics. Manual classification is subjective, labor-intensive, and vulnerable to significant inter-observer variability, which is precisely why automated computer vision approaches have attracted such intense interest. The problem is that until now, no one had quantitatively synthesized the scattered performance data across different AI paradigms to establish reliable benchmarks.</p>
<p>The researchers behind the meta-analysis, led by Tusneem Elhassan and colleagues, searched Web of Science and PubMed for studies published between January 2019 and March 2024, ultimately identifying six that met their strict inclusion criteria. Each study was classified into one of three feature-extraction paradigms. Hand-crafted pipelines engineer explicit morphological, fractal, and textural descriptors for input to traditional classifiers such as Random Forest and Support Vector Machines. End-to-end deep learning pipelines jointly optimize feature extraction and classification directly from raw pixel data, exemplified by a pre-trained ResNeXt-50 model that achieved an overall AUC of 98.6 percent across fifteen cell subtypes. Hybrid pipelines combine unsupervised deep feature extraction through convolutional autoencoders with supervised classification, with one such model, the CAE-ResVGG FusionNet, reporting a striking 99.9 percent overall accuracy.</p>
<p>To pool these heterogeneous results, the team performed random-effects meta-analyses separately for each cell subtype, converting multiclass model outputs into one-versus-all binary tasks and reconstructing the underlying confusion matrices where necessary. The pooled results revealed a striking asymmetry. Myeloblasts, the most clinically prominent blast type, were detected with 98.7 percent pooled sensitivity, while monoblasts and erythroblasts achieved 88.2 and 93.5 percent respectively. Promyelocytes, however, managed only 65.0 percent pooled sensitivity, with a 95 percent prediction interval stretching from 23.4 to 91.8 percent, meaning a new model in this evidence base could plausibly miss more than three-quarters of these cells. Specificity, by contrast, exceeded 94 percent for every subtype and reached 99.9 percent for both erythroblasts and monoblasts.</p>
<p>The promyelocyte problem is not a statistical artifact but a biological one. Promyelocytes share nuclear size, chromatin texture, and nucleolar prominence with myeloblasts at one boundary, while at the other boundary their defining feature, prominent primary azurophilic granules, progressively diminishes as secondary granules emerge in early myelocytes. Primary studies cited in the review confirmed that the inter-class distance between promyelocytes and monoblasts was the smallest among all evaluated class pairs, and that nucleus-to-cytoplasm ratio and granularity-texture descriptors were the least discriminative features for this subtype. Under severe class imbalance, models learn decision boundaries that preferentially reassign ambiguous cells to more frequently represented adjacent classes, directly suppressing true positive detection for the rarest and most morphologically ambiguous category.</p>
<p>The subgroup analysis stratified by feature-extraction paradigm revealed a preliminary but intriguing heterogeneity pattern. Hand-crafted pipelines demonstrated the lowest between-model variability, with I-squared values of zero percent for erythroblast, monoblast, and myeloblast sensitivity, a finding the authors attribute to the deterministic nature of explicitly engineered descriptors and the smaller, carefully curated cell subsets these methods operated on. End-to-end deep learning pipelines showed the highest heterogeneity, with I-squared reaching 99.0 percent for myeloblast sensitivity, consistent with greater architectural diversity, augmentation variability, and sensitivity to operating-point selection. Hybrid pipelines occupied an intermediate position, plausibly reflecting the regularizing influence of structured autoencoder representations. Notably, end-to-end models did not uniformly outperform simpler approaches, suggesting that architectural complexity alone does not guarantee reproducible diagnostic accuracy when training data are limited or class-imbalanced.</p>
<p>A critical methodological finding emerged from the quality assessment using the QUADAS-AI framework, an AI-adapted extension of the established QUADAS-2 tool. All six studies were rated high risk in the patient selection domain because they used curated, single-centre, case-control image cohorts rather than consecutive clinical smear series. All were rated high or unclear risk in the index test domain because training and testing used the same retrospective dataset with image-level rather than patient-level splitting, making data leakage likely and inflating apparent performance. The reference standard was the sole domain rated low risk, as cell labels were assigned by expert hematopathologists using established morphological criteria. The pooled estimates therefore represent best-case performance under controlled, curated conditions and should be interpreted as upper-bound approximations of real-world diagnostic accuracy.</p>
<p>Perhaps the most fundamental constraint on the entire evidence base is that all six studies evaluated their models on the same resource: the LMU AML Cytomorphology dataset, comprising over 18,000 expertly annotated single-cell images from peripheral blood smears of 100 AML patients and 100 non-malignant controls. This shared dataset means the underlying patient population, image acquisition protocol, staining method, and labeling process are identical across the evidence base, introducing expected between-model error correlation and rendering the pooled confidence intervals narrower than would be obtained from genuinely independent cohorts. The authors characterize their synthesis as a model-level benchmark within a fixed data environment rather than a conventional multi-study diagnostic accuracy meta-analysis, a distinction that carries profound implications for clinical translation.</p>
<p>A pre-specified sensitivity analysis tested whether the algebraic reconstruction of contingency counts, required for four of the six studies that did not publish complete confusion matrices, had materially biased the results. Restricting the analysis to the two studies with directly reported or figure-transcribed matrices produced absolute differences in pooled sensitivity ranging from just 0.1 percent for erythroblasts to 6.7 percent for promyelocytes, with specificity differences between 0.4 and 1.4 percent. No systematic directional bias was identified across subtypes, and confidence intervals fully overlapped in all cases, supporting the robustness of the primary findings. The restricted two-study analysis provides the more conservative estimate range, while the six-study pooled values represent an upper bound under the shared dataset structure.</p>
<p>The authors conclude that current AI classifiers are best positioned as adjunctive decision-support tools for pre-screening and highlighting morphologically suspicious cells, rather than as stand-alone diagnostic systems, particularly for promyelocytes where sensitivity remains suboptimal and missed classification carries direct consequences for AML management and treatment stratification. They also caution that high specificity in a one-versus-all framework should not be mistaken for high positive predictive value in routine clinical smears, where promyelocyte prevalence is typically low and even a 99.7 percent specificity may generate a substantial false discovery burden. Five priorities emerge for future work: independent multi-center datasets with prospective collection and patient-level splits, generative augmentation strategies that preserve morphologically relevant feature diversity for rare subtypes, domain adaptation across staining protocols and imaging hardware, training objectives that penalize false negatives in clinically important minority classes, and standardized reporting of full confusion matrices and decision-curve analysis. Until those standards are met, the promise of AI-assisted leukemia diagnosis remains tantalizingly close but not yet clinically deployable on its own.</p>
<p><strong>Subject of Research:</strong> Meta-analysis of artificial intelligence methods for classifying immature white blood cell subtypes in acute myeloid leukemia</p>
<p><strong>Article Title:</strong> Systematic review and meta-analysis of AI methods for immature WBC subtype classification</p>
<p><strong>Article References:</strong> Elhassan, T., Rahim, M. S. M., Elhaj, F. A., Haroon, A., &amp; Aljurf, M. (2026). Systematic review and meta-analysis of AI methods for immature WBC subtype classification. <em>Machine Learning with Applications, 26</em>, Article 101017. <a href="https://doi.org/10.1016/j.mlwa.2026.101017" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101017</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101017" rel="noopener noreferrer">10.1016/j.mlwa.2026.101017</a></p>
<p><strong>Keywords:</strong> acute myeloid leukemia, white blood cells, deep learning, meta-analysis, cytomorphology, promyelocytes, convolutional neural networks, diagnostic accuracy, QUADAS-AI, blood smear, machine learning, class imbalance</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222762</post-id>	</item>
	</channel>
</rss>
