Long COVID, the persistent and often debilitating condition that follows acute SARS-CoV-2 infection, remains notoriously difficult to anticipate at the bedside. Clinicians have long hoped that blood proteins measured during the acute phase of infection might serve as early warning signals, flagging which patients will go on to develop lasting symptoms. A new study published in Scientific Reports puts one such candidate strategy through a rigorous stress test, and the results are a sobering reminder of how hard it is to move a predictive model from one patient cohort, or one laboratory platform, to another.
The study, conducted by Chuan-Xin Duan of the National Cancer Center in Beijing and Lang Yang of China-Japan Friendship Hospital, focused on two proteins: fibroblast growth factor 2 (FGF2) and peroxiredoxin 3 (PRDX3). Both have plausible biological connections to the processes thought to underlie Long COVID, including tissue repair, inflammation, and oxidative stress. The researchers asked a deceptively simple question: does adding measurements of these two proteins, taken during acute infection, meaningfully improve the prediction of Long COVID beyond what can already be achieved with a patient’s age and sex alone?
What distinguishes this analysis is the discipline of its design. Rather than searching through hundreds or thousands of proteins for the best-performing combination, the authors fixed their protein pair before ever looking at outcome data. This choice was made possible by a practical constraint: FGF2 and PRDX3 were the only proteins with one-to-one assay mappings between the different proteomic platforms used by the cohorts under study. In other words, these were the only markers whose measurements could be compared in a like-for-like manner across datasets generated with different technologies. By locking in the pair in advance, the researchers avoided the well-known trap of overfitting, in which models appear accurate on the data used to build them but fail on new patients.
The analytical pipeline drew on two independent cohorts. The first, Su/INCOV, served as the reference or training dataset. It contained 204 participants with complete predictor information, but only 125 of them had linked follow-up data documenting whether they developed Long COVID. The strict analysis ultimately included 122 participants, of whom 75 experienced the outcome. The second cohort, from Zurich, served as the validation set, comprising 113 participants with 40 events. Crucially, the researchers did not refit their model to the Zurich data. Instead, they applied the coefficients estimated in the Su cohort directly to the Zurich participants, standardizing protein values within each cohort beforehand. This procedure, known as coefficient transfer, is a stringent test of whether a model truly generalizes or merely memorizes the quirks of its original dataset.
To evaluate performance, the researchers relied on the Brier score, a standard metric in prediction research that combines discrimination and calibration into a single number. Lower Brier scores indicate better prediction. The key comparison was between a baseline model using only age and sex, designated C0, and an enhanced model that also incorporated FGF2 and PRDX3, designated C2. The difference between these two scores, after statistical correction for optimism, would reveal whether the proteins added genuine predictive value.
The answer, in both cohorts, was no. In the Su/INCOV reference cohort, the optimism-corrected difference between the protein model and the baseline model was +0.00342, with a 95 percent interval spanning from −0.00608 to +0.02834. Because this interval crosses zero, the difference is statistically inconclusive; the protein model might be slightly better or slightly worse, but there is no reliable evidence of improvement. The Zurich results told a similar story. The transferred model produced a difference of +0.01093 (95 percent interval, −0.01721 to +0.03882), and a strict-model transfer gave +0.00902 (−0.01809 to +0.03696). Again, the intervals included zero, and the positive point estimates suggested that adding the proteins may even have made prediction marginally worse.
Calibration emerged as a particular problem. Both models overpredicted Long COVID risk when applied to the Zurich cohort, meaning they systematically estimated higher probabilities than what actually occurred. In prediction research, calibration failures of this kind are often more damaging than modest losses of discrimination, because they distort the risk estimates that would inform clinical decisions. A patient told they have a 60 percent chance of developing Long COVID, when the true figure is closer to 35 percent, faces unnecessary anxiety and potentially inappropriate interventions. When the researchers, at reviewers’ request, updated the model’s intercept to correct for this miscalibration, the difference between the protein model and the baseline remained inconclusive, reinforcing the overall finding.
The authors are candid about the limits of their pipeline. Because protein standardization depends on the characteristics of a target cohort, the approach as implemented is not ready to predict risk for an individual patient. Standardization, which rescales protein values to a common reference distribution, is a necessary step when combining measurements from different proteomic platforms. But it introduces a dependency on the very cohort to which the model is being applied, which complicates the clean transfer of coefficients from one dataset to another. This technical subtlety is often overlooked in biomarker studies that celebrate apparent predictive success within a single dataset, only to see that success evaporate in external validation.
The study also highlights the challenges of working with modest sample sizes. With 75 events in the reference cohort and 40 in the validation cohort, the statistical intervals around the performance differences are wide. A truly helpful biomarker effect of small magnitude could hide within those intervals, just as a harmful one could. Larger, prospectively designed cohorts with complete follow-up would be needed to settle the question definitively. The researchers note that their secondary analysis used only publicly accessible, deidentified data from the source studies, and that while the original result-blind protocol and statistical analysis plan are preserved in a versioned reproducibility archive, the analysis was not registered in a public study registry.
For the broader field of Long COVID research, the message is one of methodological caution rather than outright pessimism. FGF2 and PRDX3, at least as measured in this pipeline, did not reliably improve prediction beyond age and sex, two of the most basic demographic predictors available. But the study’s framework, with its pre-specified biomarker pair, strict coefficient transfer, and honest reporting of inconclusive results, offers a template for how candidate biomarkers should be evaluated before being promoted as clinical tools. As proteomic technologies mature and larger cohorts become available, the search for acute-phase predictors of Long COVID will continue. This study demonstrates that the bar for claiming success must include not just accuracy within a dataset, but demonstrable transportability across cohorts and platforms, with calibration that holds when the model meets patients it has never seen.
Subject of Research: Evaluation of acute-phase blood proteins FGF2 and PRDX3 for predicting Long COVID across cohorts and proteomic platforms
Article Title: Evaluation of acute FGF2 and PRDX3 for long COVID prediction across cohorts and proteomic platforms
Article References: Duan, C.-X., & Yang, L. (2026). Evaluation of acute FGF2 and PRDX3 for long COVID prediction across cohorts and proteomic platforms. Scientific Reports. https://doi.org/10.1038/s41598-026-75618-6
Image Credits: AI Generated
DOI: 10.1038/s41598-026-75618-6
Keywords: Long COVID, FGF2, PRDX3, biomarkers, proteomics, prediction models, coefficient transfer, calibration, Brier score, transportability, reproducibility, SARS-CoV-2
Cite Scienmag News
Denise Maddox. (October 10, 2026). Blood Proteins FGF2 and PRDX3 Fail to Improve Long COVID Prediction Across Cohorts. Scienmag. https://scienmag.com/blood-proteins-fgf2-and-prdx3-fail-to-improve-long-covid-prediction-across-cohorts/
Denise Maddox. "Blood Proteins FGF2 and PRDX3 Fail to Improve Long COVID Prediction Across Cohorts." Scienmag, 10 October 2026, https://scienmag.com/blood-proteins-fgf2-and-prdx3-fail-to-improve-long-covid-prediction-across-cohorts/. Accessed 10 October 2026.
Denise Maddox. "Blood Proteins FGF2 and PRDX3 Fail to Improve Long COVID Prediction Across Cohorts." Scienmag. October 10, 2026. https://scienmag.com/blood-proteins-fgf2-and-prdx3-fail-to-improve-long-covid-prediction-across-cohorts/

