For the past two years, a cloud has hung over one of computational biology’s most ambitious goals: using deep learning to predict how a cell’s entire gene expression program responds when a gene is switched off, dialed up, or silenced entirely. A string of high-profile benchmarking studies concluded that sophisticated neural networks, including transformer-based foundation models trained on millions of single-cell profiles, failed to beat something almost embarrassingly simple — the mean baseline, which merely averages the perturbed expression profiles seen during training. If a model that ignores the identity of the perturbed gene performs as well as a state-of-the-art neural network, the entire enterprise of in silico genetic screens for drug discovery looks shaky. Now a team at Shift Bioscience, working with Bo Wang of the University of Toronto, argues that the models were never as bad as the benchmarks suggested. The problem, they report in Nature Biotechnology, lies in the rulers used to measure them.
The researchers’ central claim is that the metrics most commonly used to score perturbation prediction models — mean squared error (MSE), mean absolute error (MAE), and the control-referenced Pearson correlation, Pearson(Δ_ctrl) — are frequently miscalibrated. A calibrated metric should reward any predictor that captures genuine perturbation-specific signal and penalize predictors that do not. To test whether a metric meets that standard, the team borrowed a concept from experimental biology: controls. The mean baseline serves as an intuitive negative control, since it contains no perturbation-specific information. But benchmarks had lacked a positive control — a predictor that, by construction, contains real signal — leaving an ambiguity whenever models scored poorly. Low scores could mean the model failed, or simply that the metric was too insensitive to notice success.
The team’s earlier work had proposed a technical-duplicate baseline as a positive control: split the cells from each perturbation into two halves, use one half to predict the other, and see how well that works. Because both halves share the same perturbation-specific signal, the duplicate should outperform the mean baseline if a metric is working properly. On the Norman19 dataset, with roughly 102 differentially expressed genes (DEGs) per perturbation, it did. But on the Replogle22 K562 genome-wide Perturb-seq dataset, where perturbations are far weaker — only about 3.25 DEGs per perturbation on average — the technical duplicate actually underperformed the mean baseline on MSE in roughly 95 percent of individual perturbations. The culprit, the researchers found, is an artifact they call signal dilution: when perturbations touch only a handful of genes out of thousands, error metrics are dominated by the vast transcriptomic background, and a better estimate of unperturbed expression — which the mean baseline provides thanks to greater statistical power — looks more accurate than a prediction carrying real but sparse biological signal.
To build a positive control that works even for weak perturbations, the team introduced the interpolated duplicate. This baseline blends the technical duplicate and the mean baseline gene by gene, using an interpolation weight α derived from the adjusted P value of a differential expression analysis. Genes showing strong statistical evidence of being perturbed are weighted toward the technical duplicate; genes that appear unaffected are weighted toward the mean baseline. The result is a predictor that consistently outperforms the mean baseline across datasets, providing the reliable positive control the field had been missing. With negative and positive controls in hand, the researchers could finally ask the question that had gone unanswered: which benchmarking metrics can actually tell the two apart?
Their answer takes the form of a new meta-metric called the dynamic range fraction, or DRF. For any candidate metric, DRF measures how much of the theoretical gap between a perfect prediction and the negative control is actually covered by the empirical gap between the positive control and the negative control. If the positive control comfortably beats the negative control, the metric is well calibrated and DRF approaches one. If DRF hovers near zero, the metric is essentially blind to perturbation-specific signal, no matter how good the model is. Applied per perturbation, DRF turns metric evaluation from a matter of taste into a measurable quantity — and the results were striking.
Across 14 datasets and 18 metrics spanning direct reconstruction error, reference-based delta metrics, and retrieval-based metrics, the workhorse metrics of the field fared poorly. MSE and Pearson(Δ_ctrl) showed low DRF values in most perturbations of the Replogle22 K562 dataset, meaning they could barely distinguish a signal-bearing predictor from an uninformative one. The evidence was not merely statistical. When the team examined SP2, a transcription factor whose inhibition in Replogle22 K562 yields a DRF for MSE near zero, gene set enrichment analysis of the differential expression ranks still recovered SP2’s own binding targets among the top ENCODE/ChEA gene sets — clear biological signal that MSE simply could not see. In contrast, weighted and rank-based metrics, including weighted MSE (WMSE), weighted R²Δ, and the normalized inverse rank (NIR), consistently showed higher calibration. These metrics share a design principle: they upweight the small fraction of genes that actually respond to a perturbation, rather than letting the silent transcriptome drown the signal out.
With well-calibrated metrics in place, the team re-benchmarked nine models on the task of predicting responses to perturbations held out during training: scGPT, GEARS, PRESAGE, scLambda, CellFlow, and four foundation-model embedding probes built on Geneformer, ESM2, scGPT, and GenePT representations. The reversal was dramatic. Under poorly calibrated metrics such as MSE and Pearson(Δ_ctrl), earlier models like scGPT and GEARS showed no advantage over baselines, consistent with prior benchmarking reports. Under well-calibrated metrics, however, even these earlier models mostly outperformed the uninformative baselines — their perturbation-specific signal had been present all along, merely obscured by the choice of ruler. Newer architectures went further: PRESAGE, an attention-based model that encodes biological prior knowledge, and scLambda, a variational autoencoder with language-model-derived embeddings, often beat the baselines even under poorly calibrated metrics. On every metric tested, at least one deep learning model outperformed all baselines, and the findings replicated in the independent Nadig25 HepG2 dataset.
The combination-prediction task, where models must predict the transcriptomic effect of perturbing two genes at once, told a more nuanced story. On Norman19, the field’s favorite benchmark, the additive baseline — which simply sums the effects of the two single perturbations — proved nearly unbeatable, and the team explains why: about 96 percent of effects in that dataset are additive, and the training set covers only 31 of 4,950 possible gene pairs, or 0.63 percent of the combinatorial space, leaving almost nothing to learn about nonadditive interactions. The additive baseline captured roughly 88 percent of the ideal performance gap on the best-calibrated metrics. On Wessels23, a dataset with ten times the combinatorial coverage, the picture changed: multiple models surpassed the additive baseline under WMSE and weighted R²Δ, and PRESAGE beat it on 15 of 18 metrics. The team recommends that future studies stop using Norman19 as the sole test of combinatorial modeling, since it largely measures arithmetic rather than learned genetic interactions.
Importantly, the researchers checked that their conclusions reflect biology rather than metric trivia. Models that scored well under well-calibrated metrics also performed better on two downstream tasks: pathway recovery, measured by the correlation between gene set enrichment results computed on predicted versus true perturbation effects, and neighborhood-structure recovery, measured by the overlap between nearest-neighbor graphs built from predicted and observed perturbation shifts. Pathway recovery rankings correlated strongly with well-calibrated metric rankings (average Spearman ρ of 0.76 in Replogle22 K562 and 0.84 in Wessels23), and neighborhood recovery tracked retrieval-based metrics most closely. Sensitivity analyses varying cell counts, sequencing depth, and the number of highly variable genes confirmed that well-calibrated metrics retain their advantage even when data quality degrades.
The study does not declare victory for deep learning outright. The authors caution against relying on any single metric, note that the unseen-context task — predicting responses in unobserved cell types or donors — remains unaddressed, and acknowledge that any fixed set of models represents only a snapshot of a fast-moving field. But the reframing is significant. Prior benchmarks played a constructive role in exposing modeling limitations, and architectural advances have begun to address them; an independent preprint by Cole and colleagues reached complementary conclusions about the value of prior-knowledge embeddings. What this work establishes is that the debate over whether neural networks can predict genetic perturbations was, in part, a debate about measurement. With calibration-aware evaluation built on proper positive and negative controls, the models look considerably more capable than the field believed — and the path toward reliable in silico perturbation screens looks considerably shorter.
Subject of Research: Metric calibration for benchmarking deep learning models of genetic perturbation responses
Article Title: Deep learning perturbation models can outperform baselines on calibrated metrics
Article References: Miller, H. E., Mejia, G. M., Leblanc, F. J. A., Swain, B., Wang, B., & de Lima Camillo, L. P. (2026). Deep learning perturbation models can outperform baselines on calibrated metrics. Nature Biotechnology. https://doi.org/10.1038/s41587-026-03307-w
Image Credits: AI Generated
DOI: 10.1038/s41587-026-03307-w
Keywords: deep learning, genetic perturbation, Perturb-seq, benchmarking, metric calibration, single-cell RNA-seq, dynamic range fraction, weighted MSE, foundation models, drug discovery, computational biology, Nature Biotechnology
Cite Scienmag News
Juliet Wilcox. (October 8, 2026). Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction. Scienmag. https://scienmag.com/flawed-benchmarks-hid-ai-success-in-gene-perturbation-prediction/
Juliet Wilcox. "Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction." Scienmag, 8 October 2026, https://scienmag.com/flawed-benchmarks-hid-ai-success-in-gene-perturbation-prediction/. Accessed 8 October 2026.
Juliet Wilcox. "Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction." Scienmag. October 8, 2026. https://scienmag.com/flawed-benchmarks-hid-ai-success-in-gene-perturbation-prediction/

