Monday, October 5, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage

October 5, 2026
in Biology
Drew Townsend
By Drew Townsend Scienmag Editorial Profile - Cell Biology
Reading Time: 4 mins read
0
AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage

AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Machine learning models that promise to predict whether a new drug candidate will be toxic may be flattering themselves. A team at Gachon University in South Korea has built a new benchmark, called ToxBench, that systematically audits how toxicity-prediction algorithms are tested, and the results suggest that many reported performance figures are inflated by a subtle but pervasive flaw: data leakage. The study, published in BMC Bioinformatics, quantifies exactly how much performance drops when structurally similar molecules are kept out of the test set, and it goes further by examining whether the probabilities models produce can actually be trusted.

The core problem is deceptively simple. When researchers evaluate a toxicity model, they typically shuffle their dataset randomly and hold out a fraction for testing. But chemical datasets are full of close molecular relatives, analogs that differ only by a small substituent yet share nearly identical structures. Under random splitting, a near-twin of a test compound often sits in the training data, and the model effectively memorizes the answer rather than learning generalizable chemistry. ToxBench was designed to measure this inflation directly rather than assume it away.

The benchmark assembles three widely used toxicology datasets: Tox21, covering 7,538 compounds across 12 toxicity endpoints; ClinTox, with 1,379 compounds and 2 tasks focused on clinical toxicity; and SIDER, containing 1,350 compounds and 27 side-effect tasks. All three were processed through a transparent standardization pipeline in which the removal of compounds with conflicting labels is explicitly reported, an unusual level of bookkeeping that makes the benchmark reproducible from end to end.

Four model classes were put through their paces: Random Forest, XGBoost, a multilayer perceptron, and a Graph Neural Network that reads molecules as molecular graphs. Each model was trained under both random and Bemis-Murcko scaffold-based splits, the latter grouping compounds by their shared molecular scaffold so that entire structural families fall on one side of the divide. Every configuration was repeated across five independent random seeds, yielding 120 distinct experimental conditions, a deliberate hedge against the luck-of-the-draw variability that plagues small molecular datasets.

The headline finding is stark. Moving from random to scaffold splitting reduced the area under the receiver operating characteristic curve, or AUROC, by 0.057 to 0.079 points across all four model classes on Tox21, with a mean drop of 0.070. On SIDER, three of the four models lost 0.031 to 0.035 AUROC points. In a field where claimed improvements between competing algorithms are often smaller than that margin, the implication is uncomfortable: a substantial share of the progress reported in predictive toxicology may reflect leakage rather than genuine learning.

Crucially, the authors refused to treat scaffold splitting as a magic fix. Instead of assuming that grouping by scaffold eliminates leakage, they audited the residual structural similarity between training and test sets directly. On Tox21, the fraction of test compounds with a nearest-neighbor Tanimoto similarity above 0.6 to the training set falls from 43.9 percent under random splitting to 12.1 percent under scaffold splitting. That is a dramatic reduction, but it means roughly one test compound in eight still has a close chemical cousin in the training data, and the benchmark quantifies this residue rather than ignoring it.

Not every dataset behaved as expected. ClinTox, the smallest of the three, showed a reversed performance ordering between split types, and its seed-to-seed variance of plus or minus 0.085 to 0.160 AUROC points confirmed that results on this dataset are dominated by sampling noise rather than by the split strategy itself. With extreme class imbalance and only 1,379 compounds, ClinTox serves as a cautionary tale: benchmarks built on small, skewed datasets can produce rankings that are essentially statistical artifacts.

The study also tackled a question that receives far less attention than raw accuracy: whether a model’s predicted probabilities mean what they say. When a toxicity model reports an 80 percent chance of hepatotoxicity, regulators and medicinal chemists need that number to be calibrated, meaning that among compounds assigned that score, roughly 80 percent really are toxic. Post-hoc calibration cut the expected calibration error by 67 to 68 percent on ClinTox, showing that reliable probability estimates are achievable with simple corrections. On Tox21 and SIDER, raw Random Forest predictions were already well calibrated, with expected calibration errors of just 0.018 and 0.057 respectively, and post-hoc methods added nothing. Scaffold splitting, however, consistently worsened calibration across all three datasets, meaning the harder, more honest evaluation regime also produces less trustworthy probabilities.

The benchmark’s third pillar addresses a practical question for anyone deploying these models: when should they be trusted at all? Applicability domain analysis revealed a consistent positive relationship between the Tanimoto similarity of a test compound to its nearest training neighbor and predictive reliability, with low-similarity compounds showing notably reduced AUROC under scaffold splitting. In other words, models are most trustworthy for compounds that resemble their training data, and performance degrades predictably as inputs drift into unfamiliar chemical space. Ensemble-based uncertainty estimates proved useful here too: filtering out predictions the model itself was not confident about improved AUROC by up to 0.025 points on SIDER, offering a simple operational rule for flagging unreliable outputs before they reach a decision-maker.

Together, these analyses turn ToxBench into more than a leaderboard. The authors argue that scaffold splitting, multi-seed evaluation, and applicability domain filtering should become standard practice in toxicity prediction benchmarking, and their leakage audit provides a template for how to verify that a split is genuinely doing its job. For a field whose predictions increasingly inform drug safety decisions, the message is clear: reported accuracy is only as credible as the split that produced it, and knowing when a model is guessing outside its comfort zone may matter as much as how often it is right.

Subject of Research: A leakage-audited multi-task benchmark evaluating machine learning models for predictive toxicology with calibration, uncertainty, and applicability domain analysis

Article Title: ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis

Article References: Kartic, Seo, Y., Yi, S., & Park, T.-S. (2026). ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06621-x

Image Credits: AI Generated

DOI: 10.1186/s12859-026-06621-x

Keywords: predictive toxicology, machine learning, data leakage, scaffold splitting, ToxBench, probability calibration, uncertainty quantification, applicability domain, Tox21, ClinTox, SIDER, graph neural networks

Cite Scienmag News

Drew Townsend. (October 5, 2026). AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage. Scienmag. https://scienmag.com/ai-toxicity-predictions-overstated-new-benchmark-exposes-hidden-data-leakage/

Drew Townsend. "AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage." Scienmag, 5 October 2026, https://scienmag.com/ai-toxicity-predictions-overstated-new-benchmark-exposes-hidden-data-leakage/. Accessed 5 October 2026.

Drew Townsend. "AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage." Scienmag. October 5, 2026. https://scienmag.com/ai-toxicity-predictions-overstated-new-benchmark-exposes-hidden-data-leakage/

Tags: addressing data leakage in bioinformaticsAI toxicity prediction biasapplicability domainbenchmarking toxicity prediction accuracychemical structure similarity in machine learningClinToxdata leakagedata leakage in chemical datasetsevaluation of toxicity prediction algorithmsgeneralization issues in chemical machine learningGraph Neural Networksimpact of data leakage on model performanceMachine learningmolecular analogs and model overfittingpredictive toxicologyprobability calibrationreliable toxicity prediction in drug discoveryscaffold splittingSIDERstructural similarity in drug toxicity datasetsTox21ToxBenchToxBench benchmark for toxicity modelsuncertainty quantification
Share26Tweet16
Previous Post

Platypus-Inspired Algorithm Brings Animal Sensing to Optimization

Next Post

Job Satisfaction Emerges as Key Bridge Between University Culture and Faculty Loyalty

Related Posts

Computers Take the Guesswork Out of Antisense Drug Design
Biology

Computers Take the Guesswork Out of Antisense Drug Design

October 5, 2026
A Hidden DNA Insertion Paints Quinoa Leaves in Red and Green
Biology

A Hidden DNA Insertion Paints Quinoa Leaves in Red and Green

October 4, 2026
Mapping Climate Refuges to Give Panama’s Harlequin Toads a Second Chance
Biology

Mapping Climate Refuges to Give Panama’s Harlequin Toads a Second Chance

October 4, 2026
Iron’s Master Switch: How Cells and the Body Keep the Perfect Balance
Biology

Iron’s Master Switch: How Cells and the Body Keep the Perfect Balance

October 4, 2026
A Methylation Signature in Brain Fluid Predicts Who Survives a Deadly Stroke
Biology

A Methylation Signature in Brain Fluid Predicts Who Survives a Deadly Stroke

October 4, 2026
Scientists Build a 44-SNP Genetic Barcode to Catch Mislabelled Samples Across Labs
Biology

Scientists Build a 44-SNP Genetic Barcode to Catch Mislabelled Samples Across Labs

October 4, 2026
Next Post
Job Satisfaction Emerges as Key Bridge Between University Culture and Faculty Loyalty

Job Satisfaction Emerges as Key Bridge Between University Culture and Faculty Loyalty

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Job Satisfaction Emerges as Key Bridge Between University Culture and Faculty Loyalty
  • AI Toxicity Predictions Overstated: New Benchmark Exposes Hidden Data Leakage
  • Platypus-Inspired Algorithm Brings Animal Sensing to Optimization
  • Rebuilding the Nose from Bone: New 3D Method Aims to Sharpen Forensic Facial Reconstruction

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading