For decades, the diagnosis of obstructive sleep apnea has required a night wired to electrodes in a sleep laboratory, or at best a home testing kit that many patients never return. Now a consumer smartwatch has formally entered that territory. A new validation study published in the Journal of Clinical Sleep Medicine examined the Samsung Galaxy Watch, which recently received clearance from the US Food and Drug Administration to detect moderate to severe obstructive sleep apnea over a two-night monitoring period. The device relies on photoplethysmography, heart rate tracking, and proprietary sleep pattern information to generate an estimated apnea-hypopnea index, a number that approximates how frequently a sleeper’s breathing is disrupted. A companion commentary by sleep physicians Rami N. Khayat and Matthew S. Floyd of Pennsylvania State University argues that the study offers a reasonable roadmap for evaluating wearable diagnostics, while cautioning that the technology is not yet ready to stand alone in clinical practice.
The significance of the moment is hard to overstate. Sleep is a quantifiable and clinically rich physiological state, making it a natural target for continuous health monitoring ever since wrist-worn actigraphy was first introduced by Kripke and colleagues in 1978. Yet while wearable hardware has evolved at a breakneck pace, the standards for validating these technologies have lagged far behind. Experts have increasingly called for formal guidelines governing how consumer devices that deliver diagnostic assessments and medical advice should be tested. Adoption, meanwhile, has already crossed a tipping point: more than 40 percent of US adults now use wearable devices. As photoplethysmography and oximetry sensors mature, these gadgets are no longer merely counting steps but making claims once reserved for medical equipment.
The study by Alavi and colleagues put the Galaxy Watch through an unusually rigorous test. The investigators enrolled 147 participants and gathered data on two separate nights, comparing the watch’s estimated apnea-hypopnea index against full in-lab polysomnography, the gold standard for sleep apnea diagnosis. The two monitored nights were separated by four to ten nights of at-home recording, a design that aligned directly with the FDA-cleared two-night decision framework while also capturing the natural night-to-night variability of the disorder. The researchers benchmarked the watch not only against conventional respiratory event frequency metrics but also against hypoxic burden, an emerging measure of the cumulative oxygen deprivation a patient experiences during sleep.
The headline numbers are striking. Against the polysomnography-derived apnea-hypopnea index scored at a 4 percent oxygen desaturation threshold, the watch’s estimate achieved an area under the receiver operating characteristic curve of 0.94, with a correlation of 0.87, its strongest diagnostic performance. Even more impressive, among participants classified into the high-risk hypoxic burden stratum, the algorithm achieved 100 percent sensitivity and 100 percent specificity. This suggests that optical photoplethysmography and wrist oximetry are especially discriminative when the cumulative hypoxemic stimulus is large. The performance profile aligns with a hypoxia-based screening approach and hints at reduced discriminatory power for milder disease, where arousals rather than oxygen drops dominate the pathology.
Equally notable was the study’s attention to a well-known weakness of optical sensing: skin pigmentation. Pulse oximeters measure light absorption through tissue, and darker skin tones have historically degraded their accuracy. The investigators conducted structured subgroup analyses across sex, body mass index, age, ethnicity, and objectively measured skin pigmentation. They found consistent sensitivity above 91 percent and consistent specificity across darker skin tones, addressing a vulnerability that has plagued optical pulse oximetry for generations and that regulators have scrutinized intensely in recent years.
But the commentary authors identify serious caveats that temper the enthusiasm. The study population was enriched for individuals with either a prior diagnosis of moderate-to-severe obstructive sleep apnea or a high pre-test probability, defined by a STOP-Bang score of 3 or greater. Of the 147 participants who completed both laboratory nights, 97 had moderate-to-severe sleep apnea on both nights, while only 29 had an apnea-hypopnea index below 15 on both nights. This spectrum bias artificially inflates positive predictive values, which exceeded 95 percent in the study. In a low-prevalence general screening population, the false positive rate would be substantially higher, a statistical reality that could overwhelm sleep clinics with unnecessary referrals.
The raw performance figures at the manufacturer’s default threshold tell a more complicated story. At an estimated apnea-hypopnea index cutoff of 15 events per hour, the device showed high sensitivity of 94.1 percent but a notably modest specificity of only 66.7 percent against the 4 percent oxygen desaturation index, and 65.8 percent against the standard apnea-hypopnea index. That figure stands in contrast to the 87.7 percent specificity reported in the FDA De Novo authorization summary. When the investigators shifted the cutoff to a cohort-optimized threshold of 25.95 events per hour, specificity rose to 94.9 percent but sensitivity fell to 82.4 percent. In other words, the device can be tuned to catch most cases or to avoid most false alarms, but not both simultaneously, and clinicians deploying it must understand exactly where that trade-off sits.
Real-world data quality posed another problem. Among the 147 participants who completed both in-lab polysomnography sessions, 57 lacked valid watch data on one or both laboratory nights because of poor signal quality, motion artifact, or inadequate wrist coupling, an invalidation rate of 38.8 percent during attended studies. The researchers supplemented missing data with home recordings for research purposes, but unwitnessed home recordings remain susceptible to unmeasured confounders such as positional therapy, alcohol consumption, sleep fragmentation, or undiagnosed central apnea events. A wearable that fails to produce usable data in nearly four out of ten supervised sessions highlights the ongoing engineering challenge of sensor displacement and optical signal loss in wrist-worn consumer hardware.
The study also excluded patients with significant cardiac disease, chronic pulmonary disease, movement disorders, and comorbid sleep disorders. These are precisely the comorbid conditions that are common among patients referred for sleep evaluation, meaning the algorithm’s performance in multimorbid populations remains unproven. The US Preventive Services Task Force, moreover, does not currently support screening for obstructive sleep apnea in the absence of symptoms, citing insufficient evidence of screening accuracy and inconsistent evidence on health outcomes. Any wearable-based screening program would therefore operate in a gray zone of preventive medicine, generating findings that guidelines do not yet know how to handle.
The practical implications for clinicians are concrete. Used as a sole screening tool at the default 15 events per hour threshold, the watch would generate a significant number of false positives, burdening downstream sleep clinic resources and confirmatory diagnostic pathways. Confirmatory testing with polysomnography or home sleep apnea testing would still be required before initiating therapy. Yet the commentary authors see genuine value in the device’s ability to identify high-risk individuals with significant hypoxic burden, positioning it as an accessible triage entry point rather than a diagnostic endpoint. In their view, the work by Alavi and colleagues provides a reasonable roadmap for validating emerging wearable devices and algorithms, and industry’s willingness to seek FDA approval for clinical claims is itself an encouraging development. Deployment, they conclude, would still require careful appraisal of positive and negative predictive values within each specific population and clinical setting. The smartwatch has crossed a threshold, but it has not yet arrived at its destination.
Subject of Research: Validation of a consumer smartwatch algorithm for detecting moderate-to-severe obstructive sleep apnea
Article Title: A smart watch crosses the threshold into clinical practice—or not quite yet?
Article References: Khayat, R. N., & Floyd, M. S. (2026). A smart watch crosses the threshold into clinical practice—or not quite yet?. Journal of Clinical Sleep Medicine, 22(1), Article 181. https://doi.org/10.1007/s44470-026-00195-4
Image Credits: AI Generated
DOI: 10.1007/s44470-026-00195-4
Keywords: wearable technology, obstructive sleep apnea, smartwatch, photoplethysmography, sleep medicine, FDA clearance, hypoxic burden, polysomnography, diagnostic validation, screening, oximetry, skin pigmentation
Cite Scienmag News
Ophelia Keating. (October 6, 2026). Smartwatch Sleep Apnea Detection Shows Promise but Faces Clinical Hurdles. Scienmag. https://scienmag.com/smartwatch-sleep-apnea-detection-shows-promise-but-faces-clinical-hurdles/
Ophelia Keating. "Smartwatch Sleep Apnea Detection Shows Promise but Faces Clinical Hurdles." Scienmag, 6 October 2026, https://scienmag.com/smartwatch-sleep-apnea-detection-shows-promise-but-faces-clinical-hurdles/. Accessed 6 October 2026.
Ophelia Keating. "Smartwatch Sleep Apnea Detection Shows Promise but Faces Clinical Hurdles." Scienmag. October 6, 2026. https://scienmag.com/smartwatch-sleep-apnea-detection-shows-promise-but-faces-clinical-hurdles/

