Every year, thousands of people pedal or walk to exhaustion on a laboratory cycle ergometer or treadmill while a metabolic cart measures the air they breathe, breath by breath. Hidden in those gas exchange curves is one of the most informative numbers in exercise physiology: the first ventilatory threshold, or VT1, the intensity at which the body shifts from moderate to heavy exercise as lactate begins to accumulate faster than it can be cleared. A new study published in Sports Medicine – Open by researchers at KU Leuven and University Hospitals Leuven in Belgium reports that a freely available computer algorithm can now pinpoint this threshold almost as reliably as a pair of experienced human assessors, a finding that could reshape how exercise tests are analysed in research settings and, eventually, in the clinic.
VT1 matters because it is a non-invasive, effort-independent window into aerobic fitness. Unlike peak oxygen uptake, which depends on whether a person truly pushes to their limit, VT1 emerges from underlying physiology: as exercise intensity rises, lactate is buffered by bicarbonate, producing extra carbon dioxide. This causes carbon dioxide output to climb disproportionately relative to oxygen uptake, and ventilation follows. The classic V-slope method, introduced by Beaver, Wasserman and Whipp in 1985, detects the breakpoint in the relationship between carbon dioxide output and oxygen uptake. Exercise prescriptions built around VT1 have been shown to deliver greater improvements in peak oxygen uptake and metabolic health than prescriptions based on fixed percentages of maximum capacity, making accurate threshold detection far more than an academic exercise.
The problem has always been that finding that breakpoint is largely an art. Trained physicians inspect the plotted curve and mark where the slope changes, often cross-checking with complementary techniques such as the ventilatory equivalents or end-tidal pressure methods. Even among experienced assessors, disagreement occurs in roughly 39 percent of cases, with variability of about 11 percent in healthy individuals and up to 18 percent among less experienced raters. Visual assessment is also slow, a serious bottleneck when large datasets from thousands of tests must be analysed. Automated detection promises objectivity, reproducibility and speed, but earlier attempts have struggled: a software tool tested in more than 500 asymptomatic volunteers produced mean differences of 22 percent with limits of agreement exceeding 800 mL/min, while a convolutional neural network approach reported average differences of 11 percent.
The Belgian team, led by Margot Vermeiren, set out to validate a revised V-slope algorithm that builds on the original Beaver approach but replaces its iterative breakpoint search with a fully data-driven mathematical procedure. The algorithm fits both a linear regression and a quadratic curve to the carbon dioxide-oxygen relationship across the incremental exercise phase, then locates the point of maximum perpendicular distance between the two fits. That point divides the data into lower and upper segments, each fitted with its own linear regression. The breakpoint is accepted as VT1 only if the difference in slope between the two segments exceeds 0.1, a criterion inherited from the original method to filter out spurious breakpoints caused by noise. If the criterion is not met, the algorithm is allowed to examine neighbouring values before giving up.
To test the algorithm, the researchers drew on the iCOMPEER registry, a retrospective collection of cardiopulmonary exercise tests performed on a cycle ergometer at University Hospitals Leuven between 2010 and 2020. From 3,466 individuals, strict exclusions removed anyone with cardiovascular, pulmonary or autoimmune disease, diabetes, malignancy or bariatric surgery, along with users of certain medications, people aged 80 or older, and those with a body mass index below 18. Of 813 eligible subjects, 271 had raw time-series data available, and after removing tests with indeterminate or failed assessments the final cohort comprised 262 adults, with a mean age of 46 years, 61 percent women, predominantly overweight, and a good average exercise capacity at 93 percent of predicted maximum. Both the manual and automated assessments were applied to identically pre-processed signals, with outliers corrected using a custom Python pipeline.
The reference standard was deliberately rigorous. Two experienced observers independently inspected every curve and marked the visual V-slope breakpoint, and their average value became the manual VT1. Tests in which no discernible breakpoint existed were classified as indeterminate and excluded when both observers agreed. The automated method, implemented through the publicly available get_GET package in Python, then processed the same data. The team compared the two approaches using a battery of complementary statistics: Bland-Altman analysis for bias and limits of agreement, intraclass correlation coefficients for agreement, Deming regression to account for error in both methods, and formal equivalence testing with pre-defined margins.
The headline result is striking. Median oxygen uptake at VT1 was 1,151 mL/min with the algorithm versus 1,122 mL/min by eye, a mean difference of just 29 mL/min, or 2.2 percent, with limits of agreement of minus 11.7 to plus 16.0 percent after adjusting for the fact that measurement error grows with the magnitude of the measurement. The intraclass correlation coefficient reached 0.97, and Deming regression yielded a slope of 0.95 with an intercept of 26.1 mL/min, indicating a small proportional bias whereby automated values ran slightly higher at greater fitness levels. Equivalence testing with a margin of 100 mL/min was significant, confirming that the methods agree within a clinically meaningful band. Crucially, the inter-method agreement of 2.2 percent closely approached the 0 to 11 percent variability reported between experienced cardiologists, and the algorithm’s ICC of 0.97 actually exceeded the 0.93 agreement between the two human raters themselves.
The picture is less tidy at the level of the individual. The absolute limits of agreement, spanning roughly minus 140 to plus 199 mL/min, exceeded the pre-specified 100 mL/min boundary, and 21 percent of participants fell outside that margin. Agreement was better in women than in men, likely reflecting finer physiological resolution from the smaller work-rate increments used in their protocols. The algorithm also failed to identify a valid threshold in about 2 percent of tests, most of which were also indeterminate by eye, suggesting that these failures reflect genuinely ambiguous curve shapes rather than a specific weakness of the software. The researchers catalogued several sources of discordance: near-linear curves without a clear breakpoint, coarse data resolution near the threshold, gradual transitions where the eye picks the first deviation while the algorithm anchors to maximum curvature, artefactual data points that distort regression fitting, and early thresholds that leave an excess of high-intensity data points pulling the mathematical breakpoint rightward.
Heart rate at VT1, often used to translate the threshold into training zones, told a similar story. Agreement was good, with an ICC of 0.91, a mean bias of 4 beats per minute, and formal equivalence within a 7 bpm margin, though the sex pattern reversed, with marginally better agreement in men, possibly because women’s smaller stroke volumes demand steeper heart rate responses that amplify small threshold discrepancies.
The authors conclude that the revised V-slope algorithm is a reliable, standardized and time-efficient tool for large-scale research applications in healthy adults, where modest individual disagreement can be tolerated and expert availability is limited. But they are careful about the clinic: because individual estimates may differ meaningfully from expert assessment in a non-negligible share of cases, automated VT1 values should be reviewed by an experienced assessor before guiding exercise prescription for a specific patient. The study population was also exclusively healthy, so validation in people with heart failure, lung disease or metabolic disorders, where ventilatory responses are more complex, remains an essential next step. Still, with the algorithm and data-cleaning code openly available on GitHub, the work marks a significant step toward making one of exercise science’s most valuable measurements faster, more consistent, and accessible at scale.
Subject of Research: Validation of an automated V-slope algorithm for detecting the first ventilatory threshold during cardiopulmonary exercise testing in healthy adults
Article Title: Validation of a Revised V-Slope Algorithm for Automated First Ventilatory Threshold Detection in Healthy Adults
Article References: Vermeiren, M., Goetschalckx, K., Van Opstal, H., Claes, J., Van Langenhoven, L., Cauwenberghs, N., Willems, R., Ntalianis, E., Kuznetsova, T., Cornelissen, V., & Sinnaeve, P. (2026). Validation of a Revised V-Slope Algorithm for Automated First Ventilatory Threshold Detection in Healthy Adults. Sports Medicine – Open, 12(1), Article 153. https://doi.org/10.1186/s40798-026-01124-8
Image Credits: AI Generated
DOI: 10.1186/s40798-026-01124-8
Keywords: cardiopulmonary exercise testing, ventilatory threshold, V-slope method, aerobic fitness, exercise prescription, algorithm validation, Bland-Altman analysis, oxygen uptake, rehabilitation, sports medicine, Validation, Revised
Cite Scienmag News
Ophelia Keating. (October 11, 2026). Algorithm Matches Human Experts in Reading Key Fitness Threshold From Exercise Tests. Scienmag. https://scienmag.com/algorithm-matches-human-experts-in-reading-key-fitness-threshold-from-exercise-tests/
Ophelia Keating. "Algorithm Matches Human Experts in Reading Key Fitness Threshold From Exercise Tests." Scienmag, 11 October 2026, https://scienmag.com/algorithm-matches-human-experts-in-reading-key-fitness-threshold-from-exercise-tests/. Accessed 11 October 2026.
Ophelia Keating. "Algorithm Matches Human Experts in Reading Key Fitness Threshold From Exercise Tests." Scienmag. October 11, 2026. https://scienmag.com/algorithm-matches-human-experts-in-reading-key-fitness-threshold-from-exercise-tests/

