Machine-learning pipelines that predict diabetes, hypertension, chronic kidney disease and heart disease from routine clinical records often begin with a seductive idea: delete the outliers, balance the classes, and the classifier will see the signal more clearly. A new study in Results in Engineering puts that idea under the microscope and finds it wanting. In a leakage-free benchmark spanning ten public chronic-disease cohorts, twelve outlier detectors, and 6,500 per-fold evaluation records, researchers Rifkat R. Davronov and Fatima T. Adilova report that not one cleaning method significantly outperformed simply doing nothing at all. Worse, the cleaning step preferentially deletes exactly the patients a screening model exists to find.
The paper’s most striking number concerns how these pipelines have been measured, not how they perform. When feature selection, scaling, outlier removal and the popular SMOTE-ENN resampling technique are all fitted on the full dataset before cross-validation — a practice the authors trace through much of the published literature — reported F1 scores are inflated by an average of 28.9 percentage points, with a range from 0.4 to 61.2 points depending on the cohort. The dominant culprit is resampling before splitting: SMOTE synthesises new minority-class points by interpolating between existing ones, so a test-fold patient can become a near-duplicate of training data, while the accompanying Edited Nearest Neighbours rule deletes borderline majority records using their labels. The evaluation, in effect, grades the model on questions it has already seen.
The inflation is largest precisely where the underlying task is hardest. On a small hypertension cohort of 175 patients, where blood pressure is essentially unpredictable from anthropometric measurements alone, the flawed protocol reports an F1 of 93.75 percent; the corrected protocol yields 35.5 percent, with the area under the ROC curve hovering near chance. On an easy, well-posed chronic kidney disease diagnosis task, the difference between protocols is a mere 0.4 points. Leakage, the authors conclude, does not add a constant bonus — it manufactures performance in proportion to how little genuine signal exists, which is why mid-to-high-nineties accuracy figures on small clinical cohorts deserve deep suspicion.
Under the corrected protocol, with every fitted step confined inside the training fold, the benchmark’s verdict on outlier removal is unambiguous. Twelve detectors — including tuned DBSCAN, Local Outlier Factor, Isolation Forest, One-Class SVM, robust covariance estimation, HDBSCAN, and the recent distribution-based methods ECOD and COPOD — were compared against a no-cleaning control over 50 folds per cell, using the Nadeau–Bengio corrected resampled t-test with Holm adjustment to account for overlapping training sets. Across all 130 method-by-cohort cells, no detector significantly beat the control. The single statistically significant result in either direction was a loss: HDBSCAN on the diabetic retinopathy cohort, dropping F1 by 9.37 points.
The authors did not merely tune their own detectors carefully and declare victory. To test the objection that poor configuration explains the null result, they granted an oracle: for each cohort, the best of 44 parameter settings of their proposed method was selected with hindsight on the very folds being reported. Even this impossible advantage yielded a mean gain of only 1.24 points — less than the 3.20 points that picking the best of 44 genuinely equivalent settings would produce from fold-to-fold noise alone. On two cohorts, the oracle’s chosen configuration removed nothing whatsoever. Across all 440 evaluated configurations, F1 correlated negatively with the fraction of data removed on seven of nine informative cohorts, as strongly as r = −0.99. Within this family of methods, removing more is monotonically worse.
If cleaning were merely useless, it would be a wasted step. The study’s forensic analysis shows it is actively hazardous. Every detector examined removes minority-class records preferentially — the proposed YadroSeg method deleted 32.8 percent of minority-class patients against 18.5 percent of majority-class records, a pattern significant on 8 of 10 cohorts, while Isolation Forest showed the most extreme ratio at 4.4 to 1. Triaging the deleted rows against physiological ranges, only about 15 percent contained impossible values that could be called genuine data errors; roughly 63 percent were rare but clinically plausible patients. The mechanism is not a defect in any particular algorithm: removal correlates strongly with local sparsity in feature space (r between 0.43 and 0.69) and barely at all with label impurity. Density-based detectors find sparse regions, and in clinical data sparse regions are populated by uncommon patients, not corrupted records.
The study is also a remarkable act of scientific self-correction. An earlier version of this work reported that its graph-based detector, YadroSeg, achieved the best F1 on three of five datasets, reaching 98.87 percent — a figure produced by the very leakage the revised paper dissects. The same configuration scores 100.00 under the flawed protocol and 58.51 under the corrected one. The authors pre-registered a series of tests to understand why their density-variation analysis, adapted from image segmentation, failed to transfer to tabular data, and report that two of their own registered explanations were refuted by the experiments designed to check them. The final finding is structural: driving the coring threshold until no core vertex survives leaves the set of removed records bit-identical on all ten cohorts, because the density sequence over a constructed clinical feature graph lacks the small set of dominant boundaries that thresholding presupposes.
What survives the transfer is the density ordering itself. When exposed as a continuous score rather than collapsed into a threshold, the peeling sequence recovered injected corruption better than the published algorithm on ten of ten cohorts, ranking second of nine scorers behind LOF. Measuring this honestly required the authors to correct their own evaluation a second time: their initial comparison scored detectors on the arbitrary masks they emitted, at flag rates differing by more than an order of magnitude. At a matched flag budget, every detector tested ranked corrupted records two to three times above chance — though none exceeded a detection F1 of 0.275, confirming that none of them is a reliable corruption detector, particularly for flipped labels, which leave a record exactly where it was in feature space.
The corrected implementation also brings a scalability lesson. The original formulation required a dense distance matrix that scaled as N to the power 2.67 and became infeasible beyond 32,000 records; the revised version, using radius queries on a spatial index, measures at N to the power 1.48 — competitive with DBSCAN and LOF, though an order of magnitude behind Isolation Forest’s effectively flat profile. The authors withdraw their earlier claim of O(N log N) scaling as unsupported by measurement, and confirm the exponents on two real clinical registries of over 100,000 and 250,000 records. A large-cohort confirmation arm on those registries, reaching nearly nine times the sample size of the largest primary cohort, reproduced the central null: cleaning never helped, and where differences reached significance they ran against it.
The practical recommendations are pointed. Every fitted preprocessing step should live inside the cross-validation loop, ideally enforced structurally through pipeline tools rather than manual sequencing. Any study adopting outlier removal should include a no-cleaning control and report the class composition of what it deletes. Calibration should be reported alongside discrimination, since resampling is known to damage probability calibration — a property the study observed across all configurations. And for reviewers, the single most informative check on a paper using the cleaning-plus-resampling-plus-ensemble template is whether the resampler is fitted inside or outside the validation loop. On the balance of evidence, the authors recommend against routine density-based outlier removal in clinical tabular prediction: a measurable risk of discarding rare patients, with no measurable benefit in return.
Subject of Research: Evaluation leakage and outlier removal in machine-learning prediction of chronic diseases from clinical tabular data
Article Title: Graph-based density-variation outlier removal for chronic disease prediction: A leakage-free benchmark and component-wise ablation
Article References: Davronov, R. R., & Adilova, F. T. (2026). Graph-based density-variation outlier removal for chronic disease prediction: A leakage-free benchmark and component-wise ablation. Results in Engineering, 32, Article 112842. https://doi.org/10.1016/j.rineng.2026.112842
Image Credits: AI Generated
DOI: 10.1016/j.rineng.2026.112842
Keywords: machine learning, outlier detection, data leakage, chronic disease prediction, SMOTE, class imbalance, DBSCAN, clinical data, cross-validation, YadroSeg, reproducibility, density-based clustering
Cite Scienmag News
Denise Maddox. (October 1, 2026). Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing. Scienmag. https://scienmag.com/widely-used-data-cleaning-step-in-chronic-disease-ai-fails-rigorous-testing/
Denise Maddox. "Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing." Scienmag, 1 October 2026, https://scienmag.com/widely-used-data-cleaning-step-in-chronic-disease-ai-fails-rigorous-testing/. Accessed 1 October 2026.
Denise Maddox. "Widely Used Data-Cleaning Step in Chronic Disease AI Fails Rigorous Testing." Scienmag. October 1, 2026. https://scienmag.com/widely-used-data-cleaning-step-in-chronic-disease-ai-fails-rigorous-testing/

