Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing

October 1, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing

Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing

Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Machine-learning pipelines that predict diabetes, hypertension, chronic kidney disease and heart disease from routine clinical records often begin with a seductive idea: delete the outliers, balance the classes, and the classifier will see the signal more clearly. A new study in Results in Engineering puts that idea under the microscope and finds it wanting. In a leakage-free benchmark spanning ten public chronic-disease cohorts, twelve outlier detectors, and 6,500 per-fold evaluation records, researchers Rifkat R. Davronov and Fatima T. Adilova report that not one cleaning method significantly outperformed simply doing nothing at all. Worse, the cleaning step preferentially deletes exactly the patients a screening model exists to find.

The paper’s most striking number concerns how these pipelines have been measured, not how they perform. When feature selection, scaling, outlier removal and the popular SMOTE-ENN resampling technique are all fitted on the full dataset before cross-validation — a practice the authors trace through much of the published literature — reported F1 scores are inflated by an average of 28.9 percentage points, with a range from 0.4 to 61.2 points depending on the cohort. The dominant culprit is resampling before splitting: SMOTE synthesises new minority-class points by interpolating between existing ones, so a test-fold patient can become a near-duplicate of training data, while the accompanying Edited Nearest Neighbours rule deletes borderline majority records using their labels. The evaluation, in effect, grades the model on questions it has already seen.

The inflation is largest precisely where the underlying task is hardest. On a small hypertension cohort of 175 patients, where blood pressure is essentially unpredictable from anthropometric measurements alone, the flawed protocol reports an F1 of 93.75 percent; the corrected protocol yields 35.5 percent, with the area under the ROC curve hovering near chance. On an easy, well-posed chronic kidney disease diagnosis task, the difference between protocols is a mere 0.4 points. Leakage, the authors conclude, does not add a constant bonus — it manufactures performance in proportion to how little genuine signal exists, which is why mid-to-high-nineties accuracy figures on small clinical cohorts deserve deep suspicion.

Under the corrected protocol, with every fitted step confined inside the training fold, the benchmark’s verdict on outlier removal is unambiguous. Twelve detectors — including tuned DBSCAN, Local Outlier Factor, Isolation Forest, One-Class SVM, robust covariance estimation, HDBSCAN, and the recent distribution-based methods ECOD and COPOD — were compared against a no-cleaning control over 50 folds per cell, using the Nadeau–Bengio corrected resampled t-test with Holm adjustment to account for overlapping training sets. Across all 130 method-by-cohort cells, no detector significantly beat the control. The single statistically significant result in either direction was a loss: HDBSCAN on the diabetic retinopathy cohort, dropping F1 by 9.37 points.

The authors did not merely tune their own detectors carefully and declare victory. To test the objection that poor configuration explains the null result, they granted an oracle: for each cohort, the best of 44 parameter settings of their proposed method was selected with hindsight on the very folds being reported. Even this impossible advantage yielded a mean gain of only 1.24 points — less than the 3.20 points that picking the best of 44 genuinely equivalent settings would produce from fold-to-fold noise alone. On two cohorts, the oracle’s chosen configuration removed nothing whatsoever. Across all 440 evaluated configurations, F1 correlated negatively with the fraction of data removed on seven of nine informative cohorts, as strongly as r = −0.99. Within this family of methods, removing more is monotonically worse.

If cleaning were merely useless, it would be a wasted step. The study’s forensic analysis shows it is actively hazardous. Every detector examined removes minority-class records preferentially — the proposed YadroSeg method deleted 32.8 percent of minority-class patients against 18.5 percent of majority-class records, a pattern significant on 8 of 10 cohorts, while Isolation Forest showed the most extreme ratio at 4.4 to 1. Triaging the deleted rows against physiological ranges, only about 15 percent contained impossible values that could be called genuine data errors; roughly 63 percent were rare but clinically plausible patients. The mechanism is not a defect in any particular algorithm: removal correlates strongly with local sparsity in feature space (r between 0.43 and 0.69) and barely at all with label impurity. Density-based detectors find sparse regions, and in clinical data sparse regions are populated by uncommon patients, not corrupted records.

The study is also a remarkable act of scientific self-correction. An earlier version of this work reported that its graph-based detector, YadroSeg, achieved the best F1 on three of five datasets, reaching 98.87 percent — a figure produced by the very leakage the revised paper dissects. The same configuration scores 100.00 under the flawed protocol and 58.51 under the corrected one. The authors pre-registered a series of tests to understand why their density-variation analysis, adapted from image segmentation, failed to transfer to tabular data, and report that two of their own registered explanations were refuted by the experiments designed to check them. The final finding is structural: driving the coring threshold until no core vertex survives leaves the set of removed records bit-identical on all ten cohorts, because the density sequence over a constructed clinical feature graph lacks the small set of dominant boundaries that thresholding presupposes.

What survives the transfer is the density ordering itself. When exposed as a continuous score rather than collapsed into a threshold, the peeling sequence recovered injected corruption better than the published algorithm on ten of ten cohorts, ranking second of nine scorers behind LOF. Measuring this honestly required the authors to correct their own evaluation a second time: their initial comparison scored detectors on the arbitrary masks they emitted, at flag rates differing by more than an order of magnitude. At a matched flag budget, every detector tested ranked corrupted records two to three times above chance — though none exceeded a detection F1 of 0.275, confirming that none of them is a reliable corruption detector, particularly for flipped labels, which leave a record exactly where it was in feature space.

The corrected implementation also brings a scalability lesson. The original formulation required a dense distance matrix that scaled as N to the power 2.67 and became infeasible beyond 32,000 records; the revised version, using radius queries on a spatial index, measures at N to the power 1.48 — competitive with DBSCAN and LOF, though an order of magnitude behind Isolation Forest’s effectively flat profile. The authors withdraw their earlier claim of O(N log N) scaling as unsupported by measurement, and confirm the exponents on two real clinical registries of over 100,000 and 250,000 records. A large-cohort confirmation arm on those registries, reaching nearly nine times the sample size of the largest primary cohort, reproduced the central null: cleaning never helped, and where differences reached significance they ran against it.

The practical recommendations are pointed. Every fitted preprocessing step should live inside the cross-validation loop, ideally enforced structurally through pipeline tools rather than manual sequencing. Any study adopting outlier removal should include a no-cleaning control and report the class composition of what it deletes. Calibration should be reported alongside discrimination, since resampling is known to damage probability calibration — a property the study observed across all configurations. And for reviewers, the single most informative check on a paper using the cleaning-plus-resampling-plus-ensemble template is whether the resampler is fitted inside or outside the validation loop. On the balance of evidence, the authors recommend against routine density-based outlier removal in clinical tabular prediction: a measurable risk of discarding rare patients, with no measurable benefit in return.

Subject of Research: Evaluation leakage and outlier removal in machine-learning prediction of chronic diseases from clinical tabular data

Article Title: Graph-based density-variation outlier removal for chronic disease prediction: A leakage-free benchmark and component-wise ablation

Article References: Davronov, R. R., & Adilova, F. T. (2026). Graph-based density-variation outlier removal for chronic disease prediction: A leakage-free benchmark and component-wise ablation. Results in Engineering, 32, Article 112842. https://doi.org/10.1016/j.rineng.2026.112842

Image Credits: AI Generated

DOI: 10.1016/j.rineng.2026.112842

Keywords: machine learning, outlier detection, data leakage, chronic disease prediction, SMOTE, class imbalance, DBSCAN, clinical data, cross-validation, YadroSeg, reproducibility, density-based clustering

Cite Scienmag News

Denise Maddox. (October 1, 2026). Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing. Scienmag. https://scienmag.com/widely-used-data-cleaning-step-in-chronic-disease-ai-fails-rigorous-testing/

Denise Maddox. "Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing." Scienmag, 1 October 2026, https://scienmag.com/widely-used-data-cleaning-step-in-chronic-disease-ai-fails-rigorous-testing/. Accessed 1 October 2026.

Denise Maddox. "Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing." Scienmag. October 1, 2026. https://scienmag.com/widely-used-data-cleaning-step-in-chronic-disease-ai-fails-rigorous-testing/

Tags: best practices for data preparation in healthcare AIchronic disease data cleaningchronic disease predictionclass imbalanceclinical datacross-validationdata leakagedata leakage in clinical machine learningDBSCANdensity-based clusteringeffects of resampling techniques like SMOTE-ENN on model accuracyevaluation of outlier removal methods in health dataimpact of data preprocessing on disease prediction modelsinfluence of data leakage on reported machine learning performancelimitations of commonMachine learningmachine learning bias in healthcareoutlier detectionoutlier detection in medical datasetspitfalls of standard data cleaning steps in chronic disease predictionreproducibilityrigorous testing of clinical data preprocessing pipelinesSMOTEYadroSeg
Share26Tweet16
Previous Post

Harare’s Gridlock Is an Economic Squeeze, and a Planning Matrix May Untangle It

Next Post

Exercise Before Cancer Surgery Boosts Fitness, Function and Recovery, Major Review Finds

Related Posts

AI in Government: Landmark Review Maps Three Decades of Policy Research
Technology and Engineering

AI in Government: Landmark Review Maps Three Decades of Policy Research

October 1, 2026
Why Thick Steel Plates Turn Brittle in the Cold: A New Map of Hidden Weak Zones
Technology and Engineering

Why Thick Steel Plates Turn Brittle in the Cold: A New Map of Hidden Weak Zones

October 1, 2026
New Benchmark Captures the Hidden Difficulty of Scheduling Psychology Clinic Interns
Technology and Engineering

New Benchmark Captures the Hidden Difficulty of Scheduling Psychology Clinic Interns

October 1, 2026
New Open-Source Tool Turns Cluster-Label Agreement Into a Tunable Dial for Benchmarking
Technology and Engineering

New Open-Source Tool Turns Cluster-Label Agreement Into a Tunable Dial for Benchmarking

October 1, 2026
When Machine Learning Labels Lie: Auditing Crypto Risk Models Exposes Hidden Weaknesses
Technology and Engineering

When Machine Learning Labels Lie: Auditing Crypto Risk Models Exposes Hidden Weaknesses

October 1, 2026
Gentle Corrugation and Narrowband Emitters Push Microcavity OLEDs Toward 69% Quantum Efficiency
Technology and Engineering

Gentle Corrugation and Narrowband Emitters Push Microcavity OLEDs Toward 69% Quantum Efficiency

October 1, 2026
Next Post
Exercise Before Cancer Surgery Boosts Fitness, Function and Recovery, Major Review Finds

Exercise Before Cancer Surgery Boosts Fitness, Function and Recovery, Major Review Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Exercise Before Cancer Surgery Boosts Fitness, Function and Recovery, Major Review Finds
  • Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing
  • Harare’s Gridlock Is an Economic Squeeze, and a Planning Matrix May Untangle It
  • New Maize Hybrids Beat Fall Armyworm and Outyield Top Commercial Checks in Nigeria

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading