<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>stability of AI models in fertility treatment &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/stability-of-ai-models-in-fertility-treatment/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 23:44:01 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>stability of AI models in fertility treatment &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI That Sees the Whole Picture: New Model Brings Stability to Embryo Selection in IVF</title>
		<link>https://scienmag.com/ai-that-sees-the-whole-picture-new-model-brings-stability-to-embryo-selection-in-ivf/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 23:44:01 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[advancements in fertility clinic AI systems]]></category>
		<category><![CDATA[AI model consistency in embryo viability prediction]]></category>
		<category><![CDATA[AI-assisted embryo selection]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[attention mechanisms]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[comparative embryo assessment methods]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[embryo cohort analysis]]></category>
		<category><![CDATA[embryo cohort comparison]]></category>
		<category><![CDATA[embryo ranking and scoring algorithms]]></category>
		<category><![CDATA[embryo selection]]></category>
		<category><![CDATA[improving clinical trust in AI tools]]></category>
		<category><![CDATA[In vitro fertilization]]></category>
		<category><![CDATA[in vitro fertilization technology]]></category>
		<category><![CDATA[integrating human expertise with AI in IVF]]></category>
		<category><![CDATA[Kendall's W]]></category>
		<category><![CDATA[model stability]]></category>
		<category><![CDATA[multi-instance learning]]></category>
		<category><![CDATA[multi-instance learning in IVF]]></category>
		<category><![CDATA[reproductive medicine]]></category>
		<category><![CDATA[stability of AI models in fertility treatment]]></category>
		<category><![CDATA[time-lapse imaging]]></category>
		<category><![CDATA[underspecification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=232470</guid>

					<description><![CDATA[A new multi-instance learning model called Cohort-AI evaluates entire IVF embryo cohorts rather than individual embryos, dramatically improving the consistency and reliability of AI-assisted embryo rankings across two fertility centers.]]></description>
										<content:encoded><![CDATA[<p>In vitro fertilization is, at its heart, a comparative exercise. When an embryologist sits down to decide which embryo to transfer, they rarely judge a single embryo in isolation. Instead, they weigh every available embryo against its siblings from the same cycle, asking which one stands the best chance of becoming a baby. Yet most artificial intelligence tools now entering fertility clinics do something fundamentally different: they score each embryo independently, one image at a time, with no knowledge of the other embryos in the cohort. A new study published in Heliyon argues that this mismatch between how AI models and how human experts actually work may be a major reason why AI-assisted embryo selection has struggled to earn clinical trust—and it proposes a fix that dramatically improves the consistency of the rankings these systems produce.</p>
<p>The research, led by Prudhvi Thirumalaraju, Manoj Kumar Kanakasabapathy, and Hadi Shafiee of Brigham and Women&#8217;s Hospital and Harvard Medical School, together with collaborators at Weill Cornell Medicine, introduces a model called Cohort-AI. Rather than evaluating embryos one by one, Cohort-AI uses a machine learning framework known as multi-instance learning to assess an entire patient&#8217;s set of embryos simultaneously, predicting the cumulative live-birth outcome of the cohort and ranking each embryo by its relative contribution to that outcome. The approach mirrors the way embryologists reason: an embryo that looks average on its own might be the best option in a weak cohort, while a good-looking embryo might rank lower in a set full of excellent candidates.</p>
<p>The motivation for the work comes from an uncomfortable finding in recent literature. A previous evaluation of eight commercial AI algorithms for embryo ranking found that agreement between the algorithms and human embryologists was lower than expected, and in some cases the models disagreed with each other. Most strikingly, two of the commercial models produced rankings that were statistically indistinguishable from random chance. Even when the models were fine-tuned on center-specific data, the discrepancies persisted. The problem, the researchers argue, is a phenomenon known in machine learning as underspecification: when many different internal solutions can achieve similar overall accuracy, small variations in training—such as the random seed used to initialize the network&#8217;s weights—can produce large, unpredictable changes in individual decisions, even while headline performance metrics look stable.</p>
<p>To test whether a cohort-aware design could tame this instability, the team trained fifty replicate versions of both a conventional single-instance convolutional neural network and the new Cohort-AI model, differing only in their random initialization. All training used more than 10,000 Day-5 embryo images from 1,258 patients at Massachusetts General Hospital, captured with time-lapse incubators. A separate dataset of 648 embryos from 53 patients at Weill Cornell was held out entirely as an independent external test set, with no tuning or adaptation of any kind. The researchers then measured how consistently the fifty replicate models of each type ranked the same embryos for the same patients.</p>
<p>The results were stark. On the Massachusetts General Hospital test set of 92 patient cohorts, the replicate Cohort-AI models achieved an average Kendall&#8217;s W—a statistical measure of rank-order agreement ranging from zero, meaning no agreement, to one, meaning perfect agreement—of approximately 0.94. The single-instance models managed only about 0.36. On the external Cornell data, the pattern held: roughly 0.91 for Cohort-AI versus 0.34 for the conventional approach, with both differences highly statistically significant. In practical terms, the conventional models, trained identically except for a random starting point, frequently disagreed with themselves about which embryo was best, while the cohort-aware models produced nearly identical rankings across all fifty replicates.</p>
<p>Consistency translated directly into fewer dangerous mistakes. The team defined a critical error as a case in which a model ranked a degenerate, grade 1 embryo as the top choice even though a viable blastocyst of grade 3 or better was present in the same cohort. On the internal test set, single-instance models made such errors at an average rate of about 12.4 percent, with individual models ranging from roughly 4 to 22 percent. Cohort-AI cut that rate to about 2.3 percent. On the external Cornell data the gap widened further: about 17.3 percent for the conventional models versus 1.24 percent for Cohort-AI. Notably, the variance in error rates across replicate models also collapsed for the cohort-aware approach, and formal testing showed that this variance did not significantly increase when the model was applied to the unfamiliar Cornell data—evidence that the stability survived a genuine distribution shift between two different clinics and microscope models.</p>
<p>The study also probed what the model had actually learned. Using attention mechanisms—the component of the network that assigns weights to different inputs when forming a decision—the researchers found that Cohort-AI consistently concentrated its attention on embryos of higher morphological grade, as judged by the modified Gardner grading system used in clinics, and assigned the lowest attention to empty wells in the culture dish. Feature-space visualizations showed that replicate Cohort-AI models clustered tightly together in how they represented embryos from the same patient, whereas the conventional models diverged widely. The authors are careful to note that attention weights are treated as a descriptive weighting mechanism, not as causal proof of embryo importance, but the alignment with embryologist grading suggests the model&#8217;s comparative reasoning resembles human judgment rather than exploiting spurious shortcuts.</p>
<p>Perhaps most importantly for clinical credibility, the stability did not come at the cost of performance. In retrospective comparisons against historical clinical outcomes, Cohort-AI&#8217;s top-ranked embryos matched the embryos clinicians actually transferred 68.7 percent of the time at Massachusetts General Hospital, compared with 44.7 percent for the single-instance models, and the live-birth rate associated with those top selections was 44.8 percent versus 41.0 percent—both above the center&#8217;s baseline of 35.1 percent for the evaluated patients. At Cornell, Cohort-AI again showed higher transfer and live-birth rates with at least twice the consistency. When the analysis was restricted to cycles known to contain at least one successful live birth, the median Cohort-AI model produced 29 live births from about 35 transfers on the internal data, while the median conventional model produced only 16 from about 23 transfers. The lowest-performing Cohort-AI replicates performed comparably to the highest-performing conventional models, meaning clinicians would no longer need to gamble on which replicate they happened to deploy.</p>
<p>A single-patient case study crystallized why these differences matter. For one patient with three high-quality blastocysts, one moderate blastocyst, and two degenerate embryos, consensus among five embryologists was unambiguous. Yet 33 of the 50 conventional models ranked a degenerate embryo above a high-quality one, 31 ranked a high-quality embryo last, and seven—including the three models with the highest validation accuracy—placed a degenerate embryo first. Pairwise rank correlations among the conventional models averaged just 0.27, with some pairs producing exactly opposite orderings. All fifty Cohort-AI models, by contrast, ranked the high-quality embryos at the top and the degenerate ones at the bottom, with an average pairwise correlation of 0.99. The two embryos Cohort-AI consistently ranked highest were among those that had resulted in live births. The lesson, the authors argue, is that strong validation accuracy is an insufficient proxy for clinical reliability in ranking tasks.</p>
<p>The implications extend beyond embryology. The authors contend that embryo selection is inherently a comparative problem, not a binary classification, and that validation frameworks designed for diagnostic AI—where each sample has a clear ground truth—are poorly suited to it. They call for standardized benchmarks that measure not only accuracy but also ranking reproducibility, cross-site robustness, and behavior across software updates, noting that AI adoption in IVF has surged from roughly a quarter of clinics in 2022 to more than half in 2025 even as many clinicians report low confidence in interpreting these systems. The team is candid about limitations: the study is retrospective, the associations with live birth are correlational rather than causal, both centers used the same family of time-lapse incubators, and no blinded comparison against embryologist rankings was performed. Prospective, multi-center trials remain the necessary next step. But the central message is clear and potentially field-changing: for AI to become dependable infrastructure in IVF, rankings must remain predictable as data and software evolve—and building the comparison into the model itself, rather than leaving it out, appears to be a powerful way to get there.</p>
<p><strong>Subject of Research:</strong> A multi-instance learning AI model for stable and reliable AI-assisted embryo selection in IVF</p>
<p><strong>Article Title:</strong> Cohort-AI: A multi-instance learning approach for improved stability and reliability in AI-assisted embryo selection</p>
<p><strong>Article References:</strong> Thirumalaraju, P., Kanakasabapathy, M. K., Kandula, H., Kandula, T., Katkuri, A. V. R., Cipriano, C., Malmsten, J. E., Zaninovic, N., Bormann, C. L., &amp; Shafiee, H. (2026). Cohort-AI: A multi-instance learning approach for improved stability and reliability in AI-assisted embryo selection. <em>Heliyon, 12</em>(15), Article e45457. <a href="https://doi.org/10.1016/j.heliyon.2026.e45457" rel="noopener noreferrer">https://doi.org/10.1016/j.heliyon.2026.e45457</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.heliyon.2026.e45457" rel="noopener noreferrer">10.1016/j.heliyon.2026.e45457</a></p>
<p><strong>Keywords:</strong> in vitro fertilization, embryo selection, artificial intelligence, multi-instance learning, deep learning, reproductive medicine, Kendall&#x27;s W, model stability, underspecification, attention mechanisms, clinical decision support, time-lapse imaging</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">232470</post-id>	</item>
	</channel>
</rss>
