When a patient is diagnosed with a soft tissue sarcoma of the arm or leg, one of the first questions they ask is deceptively simple: what are my chances? Answering that question accurately shapes everything that follows, from the aggressiveness of surgery to the decision to give radiation or chemotherapy, and from follow-up scheduling to the honest conversations doctors must have with families. For years, oncologists have relied on statistical tools called nomograms to turn tumor characteristics into personalized survival estimates. Two of the best known are the Memorial Sloan Kettering Cancer Center nomogram, developed at one of the world’s most famous cancer hospitals, and Sarculator, a more recent web-based calculator built on large European sarcoma databases. A new study from Türkiye now puts these two tools side by side in a group of patients with extremity sarcomas, and the results, while preliminary, offer a rare glimpse of how such prediction models behave outside the populations where they were born.
The research, published in BMC Cancer by a team from Dokuz Eylül University in Izmir, took the form of a single-center retrospective comparison. The investigators combed their institutional records for patients with localized, resectable soft tissue sarcoma of the extremities who underwent curative-intent surgery between 2008 and 2024. They settled on 41 patients, a modest number by any standard, and followed them for a median of 103 months, more than eight and a half years. During that time, 12 patients died, every one of them from sarcoma itself rather than from other causes. The five-year overall survival rate in the cohort was 78.0 percent, falling to 61.1 percent at ten years, figures that sit comfortably within the range reported for extremity sarcomas in larger international series.
To judge the two nomograms, the researchers relied on two complementary statistical concepts that are central to modern prognostic modeling. The first is discrimination, measured here with Harrell’s C-index, which asks how well a model separates patients who die earlier from those who live longer. A C-index of 0.5 means the model performs no better than a coin flip, while a value of 1.0 represents perfect ordering of every patient’s outcome. The second is calibration, which asks a subtler question: when the model says a patient has, say, a 70 percent chance of surviving five years, do roughly 70 percent of similar patients actually make it that far? A tool can rank patients correctly yet still be badly miscalibrated, promising more or less survival than reality delivers, and both dimensions matter when the numbers end up in a clinic note or a consent discussion.
Because the two tools were designed to answer questions at different time points, the comparison required careful methodological choices. The MSKCC nomogram generates predictions at four, eight, and twelve years after surgery, while Sarculator produces five- and ten-year estimates. The team used each horizon-specific prediction as a prognostic marker and computed C-index values over the full follow-up period, attaching bias-corrected and accelerated bootstrap 95 percent confidence intervals to every estimate. They prespecified a single paired comparison at the closest available horizons, Sarculator’s five-year prediction against the MSKCC four-year prediction, and were explicit that this contrast does not test identical prediction targets. With only 12 deaths in the entire cohort, they treated all discrimination and calibration findings as exploratory and hypothesis-generating, emphasizing uncertainty over p-values.
Against that cautious backdrop, the headline finding was clear enough. The MSKCC nomogram produced a higher C-index point estimate, hovering around 0.81 across its prediction horizons, compared with roughly 0.71 for Sarculator. In the prespecified paired comparison, the difference in C-index was 0.096, with a 95 percent confidence interval stretching from 0.004 to 0.225. The statistical tests told slightly different stories: a Wald test suggested the difference was significant at p equals 0.042, while a paired bootstrap test gave a p-value of 0.072, which fails the conventional significance threshold. The authors were refreshingly candid about this discordance, concluding that the wide interval, the mismatched horizons, and the small number of events together preclude any claim that one tool is definitively superior to the other.
Calibration told a similar story of cautious asymmetry. When the researchers plotted predicted survival against observed survival, the MSKCC nomogram’s estimates tracked the actual experience of the cohort most closely at the four- and eight-year marks. Sarculator, by contrast, tended to overestimate survival, and the discrepancy was most pronounced at the ten-year horizon, where the tool’s optimism diverged furthest from what the Izmir patients actually experienced. Overestimating survival is not a trivial flaw. A model that paints too rosy a picture could, in principle, lead clinicians to under-treat a high-risk patient or to schedule follow-up less intensively than the underlying biology warrants, although the authors stress that their small cohort prevents any precise quantification of how large this miscalibration really is.
Why should a comparison of 41 patients matter to anyone beyond the walls of a single Turkish sarcoma clinic? The answer lies in a persistent gap in the validation literature. Both nomograms were developed and refined in large, predominantly Western cohorts, and head-to-head comparisons restricted exclusively to extremity-localized disease, and conducted in non-Western populations, have been scarce. Soft tissue sarcomas are a heterogeneous family of more than 70 histological subtypes, and the composition of cases, the referral patterns, and even the prevalence of undifferentiated pleomorphic sarcoma versus other subtypes can vary between regions. A tool that performs beautifully in New York or Milan may drift when applied in Izmir, and the only way to find out is to test it, even imperfectly, in new settings.
The study’s technical approach also offers a small masterclass in how prognostic tools should be evaluated. Rather than simply reporting a single C-index and moving on, the investigators used horizon-specific predictions as the markers being tested, acknowledged that comparing a five-year forecast with a four-year forecast is an apples-to-oranges exercise, and leaned on bootstrap resampling to quantify uncertainty honestly. The bias-corrected and accelerated bootstrap method they employed adjusts for skewness and bias in the resampled distribution, producing confidence intervals that are more reliable in small samples than naive approximations. Their decision to prespecify a single primary comparison also guards against the inflation of false positives that comes from testing many contrasts and cherry-picking the significant ones, a discipline that small retrospective studies often lack.
The limitations, which the authors enumerate with unusual transparency, are nonetheless substantial. Forty-one patients and twelve events provide very little statistical fuel; a rough rule of thumb in prognostic modeling suggests that reliable estimation requires far more events per variable examined. The retrospective single-center design means the cohort reflects one institution’s referral patterns, surgical practices, and pathology reporting over a 16-year span, during which imaging, staging, and adjuvant therapy norms all evolved. Calibration could only be assessed graphically rather than with formal statistical tests, and the different prediction horizons of the two tools mean the head-to-head comparison is inherently approximate. The authors themselves conclude that no inference of comparative superiority can be drawn, and they frame the work explicitly as hypothesis-generating, calling for confirmation in larger, externally validated cohorts.
Even so, the study lands at a meaningful moment. Prediction tools in oncology are proliferating rapidly, powered by ever-larger registries and increasingly sophisticated machine learning methods, yet the quiet work of external validation, especially in populations far from the development data, often lags behind. Nomograms like MSKCC’s and calculators like Sarculator are already embedded in clinical practice, quoted in tumor boards and counseling sessions around the world. This Turkish analysis does not dethrone either tool, but it does something arguably more valuable: it demonstrates a rigorous, honest template for testing them, and it raises a concrete, testable hypothesis that the older MSKCC model may discriminate better while the newer Sarculator may run optimistic at long horizons. For the thousands of patients who face an extremity sarcoma diagnosis each year, the eventual answer to that hypothesis, tested in large and diverse cohorts, could mean survival estimates they can genuinely trust.
Subject of Research: Comparison of the MSKCC and Sarculator prognostic nomograms for predicting survival in extremity soft tissue sarcoma
Article Title: Discrimination and calibration of the MSKCC and sarculator nomograms in extremity soft tissue sarcoma: a single-center retrospective comparison
Article References: İzgör, R. B., Özalp, F. R., Aktepe, O. H., Karaoğlu, A., & Atağ, E. (2026). Discrimination and calibration of the MSKCC and sarculator nomograms in extremity soft tissue sarcoma: a single-center retrospective comparison. BMC Cancer. https://doi.org/10.1186/s12885-026-17087-8
Image Credits: AI Generated
DOI: 10.1186/s12885-026-17087-8
Keywords: soft tissue sarcoma, extremity sarcoma, MSKCC nomogram, Sarculator, prognostic nomogram, overall survival, C-index, calibration, external validation, retrospective study, BMC Cancer, Dokuz Eylül University
Cite Scienmag News
Nathaniel Bowman. (October 5, 2026). Head-to-Head Test of Two Sarcoma Survival Calculators Favors the Older Model. Scienmag. https://scienmag.com/head-to-head-test-of-two-sarcoma-survival-calculators-favors-the-older-model/
Nathaniel Bowman. "Head-to-Head Test of Two Sarcoma Survival Calculators Favors the Older Model." Scienmag, 5 October 2026, https://scienmag.com/head-to-head-test-of-two-sarcoma-survival-calculators-favors-the-older-model/. Accessed 5 October 2026.
Nathaniel Bowman. "Head-to-Head Test of Two Sarcoma Survival Calculators Favors the Older Model." Scienmag. October 5, 2026. https://scienmag.com/head-to-head-test-of-two-sarcoma-survival-calculators-favors-the-older-model/








