<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>machine learning prognosis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/machine-learning-prognosis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 18:00:04 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>machine learning prognosis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Machine Learning Model Predicts Three-Year Survival in Rare Blood Cancer</title>
		<link>https://scienmag.com/machine-learning-model-predicts-three-year-survival-in-rare-blood-cancer/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 18:00:04 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[B-cell lymphoma]]></category>
		<category><![CDATA[CatBoost]]></category>
		<category><![CDATA[Chinese hematology research]]></category>
		<category><![CDATA[clinical scoring systems]]></category>
		<category><![CDATA[data-driven prognostic models]]></category>
		<category><![CDATA[explainable AI in healthcare]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[hematology]]></category>
		<category><![CDATA[IgM antibody production]]></category>
		<category><![CDATA[IPSSWM]]></category>
		<category><![CDATA[LIME]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning prognosis]]></category>
		<category><![CDATA[multi-center clinical study]]></category>
		<category><![CDATA[overall survival]]></category>
		<category><![CDATA[patient survival variability]]></category>
		<category><![CDATA[predictive medicine]]></category>
		<category><![CDATA[prognosis]]></category>
		<category><![CDATA[rare blood cancer]]></category>
		<category><![CDATA[SHAP]]></category>
		<category><![CDATA[three-year survival prediction]]></category>
		<category><![CDATA[Waldenström Macroglobulinemia]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217790</guid>

					<description><![CDATA[Chinese researchers built an interpretable CatBoost-based machine learning model that predicts three-year overall survival in Waldenström Macroglobulinemia patients with an AUC above 0.82 on unseen data.]]></description>
										<content:encoded><![CDATA[<p>Waldenström Macroglobulinemia is one of the rarest and most enigmatic cancers of the blood, a slow-growing B-cell lymphoma that produces abnormal amounts of IgM antibody and follows a course that can differ dramatically from one patient to the next. Some patients live for decades with minimal intervention, while others deteriorate rapidly despite aggressive therapy. For decades, clinicians have relied on a handful of laboratory values and clinical scoring systems to guess where an individual patient falls on that spectrum. Now, a team of Chinese hematologists and data scientists has shown that a carefully engineered machine learning model can predict which patients will survive three years after diagnosis with a level of accuracy that substantially outperforms traditional prognostic tools, and, crucially, can explain exactly why it reaches its conclusions.</p>
<p>The study, published in Annals of Hematology, brought together 179 patients with Waldenström Macroglobulinemia treated at four medical centers across China: Fujian Medical University Union Hospital in Fuzhou, West China Hospital of Sichuan University in Chengdu, the Affiliated Hospital of Southwest Medical University in Luzhou, and Chengdu Seventh People&#8217;s Hospital. The researchers divided the cohort into a training set of 134 patients, which the algorithms would learn from, and a held-out test set of 45 patients, which would serve as an honest examination the models had never seen. Stratified randomization was used to split the data, ensuring that the proportions of survivors and non-survivors remained consistent across both groups and that the test set faithfully mirrored the population the model would eventually face in the clinic.</p>
<p>The team systematically evaluated five distinct machine learning algorithms, each representing a different mathematical philosophy for extracting patterns from clinical data. Four of them belong to the family of ensemble tree-based methods: CatBoost, XGBoost, LightGBM, and Random Forest. These algorithms build large collections of decision trees, each one trained to correct the errors of its predecessors or to vote independently on the outcome, and their aggregated judgment typically exceeds the accuracy of any single tree. The fifth contender, Logistic Regression, is a classical statistical technique that fits a weighted linear combination of predictors to the log-odds of survival. It remains a benchmark in medical prediction precisely because of its transparency, but it cannot capture the nonlinear interactions and threshold effects that often govern biology, such as the possibility that an elevated IgM level matters only in combination with a particular hemoglobin threshold.</p>
<p>Before any algorithm was trained, the researchers confronted one of the most persistent dangers in clinical machine learning: the temptation to feed the model every available variable and hope it sorts things out. With a rare disease like Waldenström Macroglobulinemia, where sample sizes are inherently limited, including hundreds of features invites overfitting, the phenomenon in which a model memorizes the idiosyncrasies of its training data rather than learning generalizable biology. To guard against this, the team employed Recursive Feature Elimination with cross-validation, or RFECV. The method works by training the model on the full feature set, discarding the least informative variables, and repeating the process iteratively, with cross-validation at each step confirming that predictive performance is preserved. In this way the pipeline acts like a sculptor, chiseling away redundant and noisy variables until only the features with genuine prognostic weight remain.</p>
<p>The competition produced a clear winner. CatBoost, a gradient boosting algorithm developed by Yandex that uses ordered boosting to reduce a subtle form of target leakage and handles categorical variables natively, retained twenty pivotal features through the RFECV screening process and formed the optimized feature subset. Its discriminative performance, measured by the area under the receiver operating characteristic curve, or AUC, reached 0.9322 on the training set, with a 95 percent confidence interval of 0.901 to 0.963, and 0.8235 on the unseen test set, with a confidence interval of 0.761 to 0.886. An AUC of 0.5 corresponds to random guessing, while 1.0 represents perfect discrimination, so a test-set value above 0.82 in a cohort of only 45 patients represents strong, clinically meaningful separation between those likely to survive three years and those at high risk of earlier death.</p>
<p>What elevates this work above many similar machine learning studies in oncology is the authors&#8217; insistence on interpretability. Black-box models have faced well-earned skepticism from clinicians, who are rightly reluctant to act on a prediction they cannot inspect. The researchers therefore applied two complementary explanation frameworks: Shapley Additive Explanations, known as SHAP, and Local Interpretable Model-agnostic Explanations, or LIME. SHAP, rooted in cooperative game theory, distributes the credit for each individual prediction among the input features in a mathematically consistent way, quantifying exactly how much each variable pushed a given patient&#8217;s predicted risk up or down. LIME takes the opposite approach, building a simple, locally faithful approximation of the complex model around a single patient to reveal which factors dominated that specific case. Together, the two tools allow a hematologist to audit the model&#8217;s logic patient by patient rather than trusting it blindly.</p>
<p>The SHAP-based feature importance analysis identified three determinants as the most critical drivers of the three-year mortality prediction: treatment status, the International Prognostic Scoring System for Waldenström Macroglobulinemia risk stratification, and hepatomegaly, the enlargement of the liver that occurs when malignant lymphoplasmacytic cells infiltrate the organ. Each of these carries a clear clinical logic. Whether and how a patient has been treated directly shapes disease control; the IPSSWM score, which incorporates age, hemoglobin, platelet count, beta-2 microglobulin, and monoclonal IgM concentration, is the field&#8217;s established risk framework; and hepatomegaly signals a greater tumor burden and organ involvement. The fact that the algorithm converged on variables that clinicians already recognize as meaningful, while weighing them in data-driven proportions, demonstrates what the authors describe as clinically actionable biological interpretability, a model whose internal reasoning aligns with, and refines, medical understanding rather than contradicting it.</p>
<p>The practical implications reach well beyond the statistics. For a disease that is currently incurable and managed with a sequence of therapies including rituximab-based immunochemotherapy, BTK inhibitors such as ibrutinib and zanubrutinib, and BCL-2 antagonists, knowing early which patients face the highest three-year mortality risk could fundamentally change clinical decision-making. High-risk patients might be steered toward more intensive frontline regimens, enrolled in clinical trials of novel agents, or monitored with greater frequency, while lower-risk patients could be spared overtreatment and its associated toxicities. The framework enables the kind of personalized risk stratification that the era of precision medicine has promised, delivered through a tool that runs on routine clinical variables rather than expensive genomic profiling, making it feasible even in resource-limited settings where Waldenström Macroglobulinemia expertise is scarce.</p>
<p>Several caveats deserve honest acknowledgment. The cohort of 179 patients, though substantial for such a rare malignancy, is modest by machine learning standards, and the confidence interval around the test-set AUC reflects that uncertainty. The retrospective, multi-center Chinese cohort means the model must be externally validated in independent populations, ideally across different ethnicities and health care systems, before widespread deployment. Treatment status itself is a variable entangled with disease severity, since the sickest patients often receive different therapies, and disentangling cause from correlation remains a challenge for any observational model. The authors also note that the article was shared early to provide faster access to peer-reviewed, accepted research, with a final Version of Record to follow, and the work was supported by the Fujian Provincial Natural Science Foundation of China, a Fujian provincial health technology project, and the National Natural Science Foundation of China.</p>
<p>Even with those limitations, the study offers a compelling template for how artificial intelligence should enter rare-disease oncology: not as an inscrutable oracle, but as a transparent, auditable partner that ranks the variables clinicians already care about, quantifies its own uncertainty, and justifies every prediction it makes. As similar interpretable frameworks are validated across larger and more diverse cohorts, the line between statistical prediction and personalized medicine will continue to blur, and for patients facing a rare, incurable lymphoma, that convergence may arrive not a moment too soon.</p>
<p><strong>Subject of Research:</strong> Machine learning prediction of three-year overall survival in Waldenström Macroglobulinemia</p>
<p><strong>Article Title:</strong> Predicting the 3-year overall survival in patients with Waldenström Macroglobulinemia using machine learning algorithms</p>
<p><strong>Article References:</strong> Huang, X., Zhang, C., Zhu, Y., Wang, X., Zhu, J., Zheng, Z., Zhan, R., &amp; Wang, S. (2026). Predicting the 3-year overall survival in patients with Waldenström Macroglobulinemia using machine learning algorithms. <em>Annals of Hematology</em>. <a href="https://doi.org/10.1007/s00277-026-07296-3" rel="noopener noreferrer">https://doi.org/10.1007/s00277-026-07296-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00277-026-07296-3" rel="noopener noreferrer">10.1007/s00277-026-07296-3</a></p>
<p><strong>Keywords:</strong> Waldenström Macroglobulinemia, machine learning, CatBoost, overall survival, prognosis, SHAP, LIME, IPSSWM, feature selection, hematology, predictive medicine, B-cell lymphoma</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217790</post-id>	</item>
	</channel>
</rss>
