<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>benchmarking toxicity prediction accuracy &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/benchmarking-toxicity-prediction-accuracy/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 00:50:47 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>benchmarking toxicity prediction accuracy &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage</title>
		<link>https://scienmag.com/ai-toxicity-predictions-overstated-new-benchmark-exposes-hidden-data-leakage/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 00:50:47 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[addressing data leakage in bioinformatics]]></category>
		<category><![CDATA[AI toxicity prediction bias]]></category>
		<category><![CDATA[applicability domain]]></category>
		<category><![CDATA[benchmarking toxicity prediction accuracy]]></category>
		<category><![CDATA[chemical structure similarity in machine learning]]></category>
		<category><![CDATA[ClinTox]]></category>
		<category><![CDATA[data leakage]]></category>
		<category><![CDATA[data leakage in chemical datasets]]></category>
		<category><![CDATA[evaluation of toxicity prediction algorithms]]></category>
		<category><![CDATA[generalization issues in chemical machine learning]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[impact of data leakage on model performance]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[molecular analogs and model overfitting]]></category>
		<category><![CDATA[predictive toxicology]]></category>
		<category><![CDATA[probability calibration]]></category>
		<category><![CDATA[reliable toxicity prediction in drug discovery]]></category>
		<category><![CDATA[scaffold splitting]]></category>
		<category><![CDATA[SIDER]]></category>
		<category><![CDATA[structural similarity in drug toxicity datasets]]></category>
		<category><![CDATA[Tox21]]></category>
		<category><![CDATA[ToxBench]]></category>
		<category><![CDATA[ToxBench benchmark for toxicity models]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=236278</guid>

					<description><![CDATA[A new leakage-audited benchmark shows that random data splitting inflates toxicity model performance by up to 0.079 AUROC points while calibration and applicability domain tools reveal when predictions can be trusted.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models that promise to predict whether a new drug candidate will be toxic may be flattering themselves. A team at Gachon University in South Korea has built a new benchmark, called ToxBench, that systematically audits how toxicity-prediction algorithms are tested, and the results suggest that many reported performance figures are inflated by a subtle but pervasive flaw: data leakage. The study, published in BMC Bioinformatics, quantifies exactly how much performance drops when structurally similar molecules are kept out of the test set, and it goes further by examining whether the probabilities models produce can actually be trusted.</p>
<p>The core problem is deceptively simple. When researchers evaluate a toxicity model, they typically shuffle their dataset randomly and hold out a fraction for testing. But chemical datasets are full of close molecular relatives, analogs that differ only by a small substituent yet share nearly identical structures. Under random splitting, a near-twin of a test compound often sits in the training data, and the model effectively memorizes the answer rather than learning generalizable chemistry. ToxBench was designed to measure this inflation directly rather than assume it away.</p>
<p>The benchmark assembles three widely used toxicology datasets: Tox21, covering 7,538 compounds across 12 toxicity endpoints; ClinTox, with 1,379 compounds and 2 tasks focused on clinical toxicity; and SIDER, containing 1,350 compounds and 27 side-effect tasks. All three were processed through a transparent standardization pipeline in which the removal of compounds with conflicting labels is explicitly reported, an unusual level of bookkeeping that makes the benchmark reproducible from end to end.</p>
<p>Four model classes were put through their paces: Random Forest, XGBoost, a multilayer perceptron, and a Graph Neural Network that reads molecules as molecular graphs. Each model was trained under both random and Bemis-Murcko scaffold-based splits, the latter grouping compounds by their shared molecular scaffold so that entire structural families fall on one side of the divide. Every configuration was repeated across five independent random seeds, yielding 120 distinct experimental conditions, a deliberate hedge against the luck-of-the-draw variability that plagues small molecular datasets.</p>
<p>The headline finding is stark. Moving from random to scaffold splitting reduced the area under the receiver operating characteristic curve, or AUROC, by 0.057 to 0.079 points across all four model classes on Tox21, with a mean drop of 0.070. On SIDER, three of the four models lost 0.031 to 0.035 AUROC points. In a field where claimed improvements between competing algorithms are often smaller than that margin, the implication is uncomfortable: a substantial share of the progress reported in predictive toxicology may reflect leakage rather than genuine learning.</p>
<p>Crucially, the authors refused to treat scaffold splitting as a magic fix. Instead of assuming that grouping by scaffold eliminates leakage, they audited the residual structural similarity between training and test sets directly. On Tox21, the fraction of test compounds with a nearest-neighbor Tanimoto similarity above 0.6 to the training set falls from 43.9 percent under random splitting to 12.1 percent under scaffold splitting. That is a dramatic reduction, but it means roughly one test compound in eight still has a close chemical cousin in the training data, and the benchmark quantifies this residue rather than ignoring it.</p>
<p>Not every dataset behaved as expected. ClinTox, the smallest of the three, showed a reversed performance ordering between split types, and its seed-to-seed variance of plus or minus 0.085 to 0.160 AUROC points confirmed that results on this dataset are dominated by sampling noise rather than by the split strategy itself. With extreme class imbalance and only 1,379 compounds, ClinTox serves as a cautionary tale: benchmarks built on small, skewed datasets can produce rankings that are essentially statistical artifacts.</p>
<p>The study also tackled a question that receives far less attention than raw accuracy: whether a model&#8217;s predicted probabilities mean what they say. When a toxicity model reports an 80 percent chance of hepatotoxicity, regulators and medicinal chemists need that number to be calibrated, meaning that among compounds assigned that score, roughly 80 percent really are toxic. Post-hoc calibration cut the expected calibration error by 67 to 68 percent on ClinTox, showing that reliable probability estimates are achievable with simple corrections. On Tox21 and SIDER, raw Random Forest predictions were already well calibrated, with expected calibration errors of just 0.018 and 0.057 respectively, and post-hoc methods added nothing. Scaffold splitting, however, consistently worsened calibration across all three datasets, meaning the harder, more honest evaluation regime also produces less trustworthy probabilities.</p>
<p>The benchmark&#8217;s third pillar addresses a practical question for anyone deploying these models: when should they be trusted at all? Applicability domain analysis revealed a consistent positive relationship between the Tanimoto similarity of a test compound to its nearest training neighbor and predictive reliability, with low-similarity compounds showing notably reduced AUROC under scaffold splitting. In other words, models are most trustworthy for compounds that resemble their training data, and performance degrades predictably as inputs drift into unfamiliar chemical space. Ensemble-based uncertainty estimates proved useful here too: filtering out predictions the model itself was not confident about improved AUROC by up to 0.025 points on SIDER, offering a simple operational rule for flagging unreliable outputs before they reach a decision-maker.</p>
<p>Together, these analyses turn ToxBench into more than a leaderboard. The authors argue that scaffold splitting, multi-seed evaluation, and applicability domain filtering should become standard practice in toxicity prediction benchmarking, and their leakage audit provides a template for how to verify that a split is genuinely doing its job. For a field whose predictions increasingly inform drug safety decisions, the message is clear: reported accuracy is only as credible as the split that produced it, and knowing when a model is guessing outside its comfort zone may matter as much as how often it is right.</p>
<p><strong>Subject of Research:</strong> A leakage-audited multi-task benchmark evaluating machine learning models for predictive toxicology with calibration, uncertainty, and applicability domain analysis</p>
<p><strong>Article Title:</strong> ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis</p>
<p><strong>Article References:</strong> Kartic, Seo, Y., Yi, S., &amp; Park, T.-S. (2026). ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06621-x" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06621-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06621-x" rel="noopener noreferrer">10.1186/s12859-026-06621-x</a></p>
<p><strong>Keywords:</strong> predictive toxicology, machine learning, data leakage, scaffold splitting, ToxBench, probability calibration, uncertainty quantification, applicability domain, Tox21, ClinTox, SIDER, graph neural networks</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">236278</post-id>	</item>
	</channel>
</rss>
