For millions of people living with chronic obstructive pulmonary disease, or COPD, pulmonary rehabilitation is one of the most effective treatments available. The structured programs of exercise training, breathing techniques, and education can improve walking capacity, reduce breathlessness, and lift quality of life. Yet the approach has a stubborn problem: adherence. Repetitive exercise routines can feel monotonous, and many patients drop out before they gain the full benefit. In recent years, researchers have turned to virtual reality and exergaming, the use of video game-style technology to make exercise more engaging, as a possible complement to conventional rehabilitation. A new umbrella review published in BMC Complementary Medicine and Therapies now takes stock of that rapidly growing evidence base, and its findings are a masterclass in scientific caution.
The review, led by Junting Sai and colleagues at the First Affiliated Hospital of Henan University of Chinese Medicine in Zhengzhou, is what researchers call an umbrella review: a systematic examination of systematic reviews. Rather than analyzing individual clinical trials directly, the team gathered and appraised the highest tier of evidence synthesis, the meta-analyses that pool results from many trials. The researchers searched nine databases and identified twelve meta-analyses published between 2020 and 2025 that examined virtual reality or exergaming-complemented pulmonary rehabilitation in people with COPD. The work was prospectively registered with PROSPERO, the international registry of systematic review protocols, under registration number CRD420251143714.
The methodological architecture of the study is worth understanding, because it explains why the conclusions are so measured. The team assessed the quality of each included meta-analysis using AMSTAR-2, a sixteen-item instrument designed to evaluate how rigorously a systematic review was conducted, covering aspects such as protocol registration, duplicate study selection, risk-of-bias assessment, and appropriate statistical synthesis. They measured overlap among primary studies using the Corrected Covered Area, or CCA, a metric that quantifies how much the same trials are being counted repeatedly across different reviews. And they graded confidence in the findings using GRADE, the Grading of Recommendations Assessment, Development and Evaluation framework, which rates certainty from high down to very low based on study limitations, inconsistency, indirectness, and imprecision.
The results of the quality appraisal were sobering. Of the twelve meta-analyses, five were rated as low quality and seven as critically low quality under AMSTAR-2. Not a single review reached moderate or high quality. This matters because meta-analyses inherit the flaws of the trials they pool; if the reviews themselves were conducted sloppily, with inadequate searches, missing protocol registration, or poor handling of bias, then their pooled estimates may be misleading no matter how impressive the statistics look. The authors deliberately chose not to run a second-level pooled analysis of the review-level estimates, reasoning that because the reviews share the same underlying primary studies, their results are statistically dependent and should not be treated as independent observations. Instead, they presented the individual meta-analytic estimates in a structured descriptive format.
The overlap analysis revealed just how dependent those estimates are. Outcome-specific CCA values ranged from 20.00 percent to 50.00 percent, which the authors characterize as very high primary-study overlap across all assessed outcomes. In practical terms, this means that the twelve reviews were not drawing on twelve independent bodies of evidence; many of the same clinical trials were being re-analyzed and re-pooled again and again. When the same patients and the same interventions are counted multiple times across reviews, the apparent volume of evidence inflates, and the impression of convergent findings can be an illusion created by repetition rather than genuine replication.
The review also exposed a technological gap between the hype and the reality. Most of the primary studies underlying the meta-analyses evaluated screen-based or motion-sensing exergaming interventions, systems in which patients exercise while following on-screen prompts or controlling games with body movements captured by sensors. Far fewer studies tested immersive virtual reality, in which patients wear head-mounted displays or enter fully projected environments that surround them with simulated worlds. This distinction is not trivial. Immersive VR is precisely the technology that has captured public imagination and driven much of the enthusiasm for gamified rehabilitation, yet the current evidence base mostly reflects simpler, non-immersive systems. The authors explicitly warn that their findings should not be interpreted as firm evidence for immersive VR interventions specifically.
So what did the pooled evidence actually show? Review-level estimates generally favored adding virtual reality or exergaming to pulmonary rehabilitation for several rehabilitation-related outcomes, including measures of exercise capacity, lung function parameters such as forced expiratory volume in one second, health-related quality of life, psychological status, and peripheral oxygen saturation. However, the picture was far from uniformly positive. Findings on dyspnea, the distressing breathlessness that defines COPD, were inconsistent across reviews. More tellingly, most estimates of improvement in the six-minute walk distance, the standard field test of functional exercise capacity, did not clearly exceed the commonly cited minimal clinically important difference, a threshold generally placed at 25 to 30 meters. A statistically significant gain that falls short of this threshold may be detectable on a spreadsheet but imperceptible in a patient’s daily life.
This is where the GRADE assessment delivers the review’s most important message: certainty of evidence was rated very low for all assessed outcomes. In GRADE terminology, very low certainty means that the true effect may be substantially different from the estimate of effect, and that further research is very likely to change the conclusions. The authors are careful to state that, because of the combination of low methodological quality, very high overlap, and very low certainty, the statistically significant findings should not be interpreted as definitive evidence of clinically important benefit. Instead, they frame the entire evidence base as exploratory rather than confirmatory, a hypothesis-generating foundation rather than a basis for clinical practice.
Why does this matter beyond the academic literature? COPD is a leading cause of death and disability worldwide, and pulmonary rehabilitation, despite its proven value, remains underused and underfunded in many health systems. If gamified technology genuinely boosts motivation, engagement, and adherence, it could extend the reach of rehabilitation into patients’ homes and make long-term exercise sustainable for people who struggle with conventional programs. The physiological rationale is plausible: interactive feedback, goal-setting, and immersive distraction can reduce perceived exertion and make training sessions more tolerable. But enthusiasm must be calibrated against evidence, and this review demonstrates that the current evidence, while abundant in quantity, is thin in quality and heavily recycled.
The authors’ prescription for the field is clear. Firm clinical recommendations will require better-designed primary trials and rigorously conducted systematic reviews that clearly distinguish immersive virtual reality from screen-based exergaming, rather than lumping heterogeneous technologies together under a single label. Future trials should be adequately powered, use standardized outcome measures, and report effects against minimal clinically important difference thresholds so that readers can judge clinical relevance rather than mere statistical significance. For now, patients and clinicians curious about gamified rehabilitation can take a measured message from this work: the approach may help with selected rehabilitation outcomes, and it appears generally feasible, but the technology’s true clinical value remains an open question that only better science can answer.
Subject of Research: Virtual reality and exergaming-complemented pulmonary rehabilitation for chronic obstructive pulmonary disease
Article Title: Virtual reality/exergaming-complemented pulmonary rehabilitation in chronic obstructive pulmonary disease: an umbrella review of systematic reviews and meta-analyses
Article References: Virtual reality/exergaming-complemented pulmonary rehabilitation in chronic obstructive pulmonary disease: an umbrella review of systematic reviews and meta-analyses. (n.d.). https://doi.org/10.1186/s12906-026-05619-5
Image Credits: AI Generated
DOI: 10.1186/s12906-026-05619-5
Keywords: COPD, pulmonary rehabilitation, virtual reality, exergaming, umbrella review, meta-analysis, AMSTAR-2, GRADE, 6-minute walk distance, dyspnea, immersive VR, evidence quality
Cite Scienmag News
Barbara Leach. (October 1, 2026). Virtual Reality Workouts for Lung Disease Show Promise, but the Evidence Is Shaky. Scienmag. https://scienmag.com/virtual-reality-workouts-for-lung-disease-show-promise-but-the-evidence-is-shaky/
Barbara Leach. "Virtual Reality Workouts for Lung Disease Show Promise, but the Evidence Is Shaky." Scienmag, 1 October 2026, https://scienmag.com/virtual-reality-workouts-for-lung-disease-show-promise-but-the-evidence-is-shaky/. Accessed 1 October 2026.
Barbara Leach. "Virtual Reality Workouts for Lung Disease Show Promise, but the Evidence Is Shaky." Scienmag. October 1, 2026. https://scienmag.com/virtual-reality-workouts-for-lung-disease-show-promise-but-the-evidence-is-shaky/

