Home sleep apnea testing has become one of the most talked-about developments in sleep medicine, promising to move diagnosis out of the laboratory and into the bedroom. For adults, portable monitoring is now routine, but extending the approach to young children has proven far more difficult. A new letter published in the Journal of Clinical Sleep Medicine raises pointed questions about how the accuracy of these home-based devices should be measured in preschool-aged patients, and its arguments carry weight for anyone following the shift toward at-home diagnostics.
The letter, written by Anmol Devi of Ghulam Muhammad Mahar Medical College in Sukkur, Pakistan, examines a recent study by Murray and colleagues that evaluated a photoplethysmography-based home sleep apnea test, known as PPG-HSAT, against polysomnography in children aged two to six years. Polysomnography, or PSG, remains the gold standard for diagnosing obstructive sleep apnea in children, but it is expensive, resource-intensive, and often uncomfortable for young patients. A reliable home alternative would be transformative, allowing clinicians to screen and diagnose far more children in their own sleep environments. The appeal is obvious, which is precisely why Devi argues that the methodological details behind accuracy claims deserve close scrutiny.
The first concern centers on what happens when the home test simply fails to produce usable data. In the original study, 43 children attempted concurrent testing with both the home device and laboratory polysomnography. Four of them, roughly nine percent, had no interpretable PPG-HSAT data at all. Among the 39 children with some captured data, a further 11, or 28 percent, had less than two hours of recorded sleep. Despite these substantial losses, the study calculated sensitivity, specificity, positive predictive value, and negative predictive value only among the children whose home recordings were successfully captured. The headline figures were striking: a sensitivity of 94.1 percent, meaning the device caught nearly all true cases, but a specificity of just 22.7 percent, meaning it flagged many children as positive who did not have the condition by the reference standard.
Devi’s argument is that these numbers may not mean what they appear to mean. In diagnostic accuracy research, excluding missing or inconclusive index-test results is a well-recognized source of bias. When only technically successful recordings are analyzed, the study population becomes a selected subset, potentially healthier in terms of device tolerance or simply easier to monitor, and therefore less representative of the intended clinical population. A device that fails in a quarter of attempted recordings is not the same device, in clinical terms, as one evaluated only where it worked. The reported sensitivity and specificity may reflect performance among technically successful recordings rather than the overall diagnostic effectiveness of the test as it would be deployed in practice.
The solution Devi proposes draws on established methodological literature. An intention-to-diagnose analysis, analogous in spirit to intention-to-treat analysis in therapeutic trials, would classify technically unsuccessful studies as indeterminate or failed tests rather than discarding them. This approach, supported by a scoping review of methods for handling missing values and inconclusive results in diagnostic studies published in Statistical Methods in Medical Research, yields estimates that better reflect real-world performance. It answers the question clinicians actually face: if I prescribe this home test for a toddler suspected of having sleep apnea, how likely am I to get an accurate answer, or any answer at all? A test with excellent accuracy conditional on success but a high failure rate may be far less useful than its published numbers suggest.
The second methodological concern involves the reference standard itself. Polysomnography is only as unbiased as the people interpreting it, and Devi notes that in the original study, PSG technologists and physicians were not blinded to study participation. Moreover, the PSG interpretation incorporated audiovisual findings, including snoring, mouth breathing, and airway-protective maneuvers. These observations are clinically valuable, but they also carry subjective judgment, and their integration into the diagnostic classification opens the door to what the literature calls diagnostic review bias, which arises when interpreters of the reference standard are aware of the index-test result or the investigational context.
The letter acknowledges that the original authors reported the SleepImage data, the commercial PPG-based system under evaluation, were reviewed independently at a later time. That is an important safeguard for the index test. What remains unclear, Devi writes, is whether the PSG interpretation was fully independent of the investigational testing context. If the sleep technologists scoring the reference standard knew that a child was part of a study evaluating a home apnea test, even subtle expectations could color their reading of borderline respiratory events. This may be particularly consequential in pediatric obstructive sleep apnea, where classification often hinges on borderline cases and where clinical and audiovisual findings, such as a parent’s report of snoring or visible mouth breathing, can influence whether a child crosses the diagnostic threshold.
Diagnostic review bias is not a hypothetical worry. Reviews of diagnostic accuracy methodology, including work published in Radiology Research and Practice on recognizing and addressing sources of bias, have repeatedly shown that unblinded reference-standard interpretation can inflate apparent agreement between tests and distort estimates of sensitivity and specificity. In fields like radiology, where imaging studies are routinely compared against clinical reference standards, blinding procedures are now considered a core quality marker of accuracy research. Sleep medicine, Devi implies, should hold itself to the same standard, especially as commercial home-testing platforms seek regulatory approval and clinical adoption on the basis of validation studies.
Importantly, the letter does not claim that the original study’s findings are invalid. Devi explicitly states that these issues do not invalidate the study, but that they may affect the reported estimates of diagnostic performance. The distinction matters. The Murray study addresses an important clinical question, and the underlying data remain valuable. The concern is about interpretation: how far the reported sensitivity of 94.1 percent and specificity of 22.7 percent can be generalized to the population of children who would actually be offered the test, including those whose recordings fail or fall short of adequate sleep duration.
The broader stakes extend well beyond a single study. Pediatric obstructive sleep apnea affects a meaningful share of preschool children and is associated with behavioral problems, poor growth, cardiovascular strain, and impaired quality of life. Untreated, it can shape a child’s development during critical years. Yet access to polysomnography is limited in many regions, and waiting lists for pediatric sleep laboratories can stretch for months. If home sleep apnea tests can be validated rigorously for young children, the public health payoff could be enormous: earlier diagnosis, earlier treatment with adenotonsillectomy or other interventions, and reduced burden on families. That promise is exactly why methodological rigor in validation studies is not an academic nicety but a prerequisite for safe clinical adoption.
Devi’s recommendations are concrete. Future evaluations of pediatric home sleep apnea tests should account for technically unsuccessful recordings, reporting them transparently and incorporating them into intention-to-diagnose analyses, and should ensure independent, ideally blinded, interpretation of the reference standard. Such measures would provide more robust and clinically meaningful estimates of accuracy, allowing clinicians and regulators to judge these devices on the terms that matter: how they perform in the full population of children who would use them, not just the subset in whom they happen to work. As home testing moves from novelty toward standard of care, the letter serves as a timely reminder that in diagnostic medicine, the fine print of methodology often determines whether a promising technology earns its place at the bedside, or in this case, the child’s own bedroom.
Subject of Research: Methodological biases in validating home sleep apnea tests against polysomnography in preschool children
Article Title: Methodological considerations in assessing diagnostic accuracy of home sleep apnea tests in children
Article References: Devi, A. (2026). Methodological considerations in assessing diagnostic accuracy of home sleep apnea tests in children. Journal of Clinical Sleep Medicine, 22(1), Article 159. https://doi.org/10.1007/s44470-026-00185-6
Image Credits: AI Generated
DOI: 10.1007/s44470-026-00185-6
Keywords: home sleep apnea test, pediatric sleep apnea, polysomnography, photoplethysmography, diagnostic accuracy, intention-to-diagnose analysis, diagnostic review bias, PPG-HSAT, children aged 2 to 6, sleep medicine, reference standard blinding, Journal of Clinical Sleep Medicine
Cite Scienmag News
Ophelia Keating. (September 13, 2026). Hidden Pitfalls in Home Sleep Apnea Tests for Young Children. Scienmag. https://scienmag.com/hidden-pitfalls-in-home-sleep-apnea-tests-for-young-children/
Ophelia Keating. "Hidden Pitfalls in Home Sleep Apnea Tests for Young Children." Scienmag, 13 September 2026, https://scienmag.com/hidden-pitfalls-in-home-sleep-apnea-tests-for-young-children/. Accessed 13 September 2026.
Ophelia Keating. "Hidden Pitfalls in Home Sleep Apnea Tests for Young Children." Scienmag. September 13, 2026. https://scienmag.com/hidden-pitfalls-in-home-sleep-apnea-tests-for-young-children/

