A scientific dispute over whether narcolepsy shortens lives has escalated into a detailed methodological exchange in the Journal of Clinical Sleep Medicine. A team of sleep medicine researchers led by Amir Sharafkhaneh of Baylor College of Medicine and the Michael E. DeBakey VA Medical Center has published a formal reply to critics who questioned the reliability of their large retrospective study on mortality in narcolepsy. The original investigation, a 25-year propensity-matched cohort study drawing on national Veterans Health Administration data, reported that patients with clinically identified narcolepsy experienced elevated mortality and higher healthcare utilization compared with matched sleep clinic patients. The critics, Mehri and Finsterer, argued that such conclusions should rest on more reliable study designs, cataloguing a series of limitations they attributed to retrospective cohort methodology. The authors’ response, published as an open-access author’s reply, defends the study’s architecture while candidly acknowledging the constraints inherent in working with electronic health record data at national scale.
At the heart of the exchange is a question that has long troubled sleep researchers: how do you study mortality in a disease as rare as narcolepsy without waiting decades for a prospective cohort to mature? Narcolepsy affects roughly one in two thousand people, and its two main forms, narcolepsy type 1 with cataplexy and narcolepsy type 2 without cataplexy, are distinguished by features that require specialized testing to confirm. The authors argue that for rare conditions, electronic health record based cohort studies remain indispensable precisely because the prospective designs their critics favor are seldom feasible at the scale or duration needed to detect mortality signals. Their study leveraged the Veterans Health Administration’s longitudinal records, which follow patients across decades of care, offering a follow-up window of 25 years that no single-center prospective study could realistically match.
One of the sharpest criticisms concerned bias. Mehri and Finsterer invoked limitations commonly associated with retrospective designs, including recall bias, selection bias, and confounding. The authors’ reply draws a technical distinction here that is worth unpacking. Recall bias, the distortion that arises when patients or investigators remember exposures imperfectly, is a meaningful threat in studies that rely on interviews or chart reviews conducted after the fact. It is not, however, a meaningful threat in administrative electronic health record cohorts, where exposures and outcomes are recorded prospectively at the point of care, independent of anyone’s memory. A diagnosis code entered during a clinic visit and a death recorded in a vital status file are not subject to the fallibilities of recollection. Selection bias and confounding, by contrast, the authors concede, remain legitimate concerns in any observational dataset, and they describe the specific countermeasures they deployed.
Those countermeasures centered on propensity score matching, a statistical technique widely used in non-randomized research to approximate the balance achieved in randomized trials. In the narcolepsy study, patients with narcolepsy were matched to comparator patients from general sleep clinics on age, sex, race and ethnicity, and index diagnosis year, with additional adjustment for body mass index and the Charlson Comorbidity Index, a validated composite measure of baseline disease burden. The authors also align themselves with earlier correspondents, Talari and Goyal, who noted that retrospective designs cannot establish causality. Their own manuscript, they emphasize, explicitly frames the findings as associations rather than causal effects. This is a crucial framing for readers: the study does not claim that narcolepsy itself causes death, but that in real-world clinical practice, patients diagnosed and managed for narcolepsy died at higher rates than comparable sleep clinic patients, even when those comparators carried a substantially heavier comorbidity burden at baseline.
The reliability of the mortality outcome itself became another point of contention, and here the authors marshal quantitative evidence. Vital status data in the Veterans Health Administration have been independently validated against the National Death Index, the United States’ authoritative repository of death records. That validation work, published by Sohn and colleagues, found a sensitivity of 98.3 percent and exact-date agreement of 97.6 percent, figures that support the accuracy and completeness of the VHA’s mortality ascertainment. In practical terms, this means the study’s primary endpoint, all-cause death, was measured against a benchmark that few national datasets can match. The authors argue that these psychometric properties of the data source underpin the reliability of their central finding, whatever reservations may attach to the diagnostic classification of the exposure.
That diagnostic classification was the critics’ most technically substantive objection. Definitive diagnosis of narcolepsy type 1 under the International Classification of Sleep Disorders, third edition, requires objective confirmation: polysomnography followed by a multiple sleep latency test demonstrating shortened sleep latency and early REM onset, or cerebrospinal fluid measurement revealing low hypocretin-1, the neuropeptide whose loss underlies the disorder’s pathophysiology. The critics asked how many participants had cataplexy and whether all had undergone these tests. The authors’ answer is a candid acknowledgment of the limits of administrative data: polysomnography and multiple sleep latency test results are not captured in a standardized format across the national VHA dataset, and cerebrospinal fluid hypocretin testing is rarely performed in routine Veterans Affairs practice. These data simply cannot be ascertained reliably from administrative records, a constraint that shaped the study’s entire design philosophy.
Rather than attempting an unattainable gold-standard confirmation, the researchers employed a computable phenotyping algorithm, requiring at least two narcolepsy diagnosis codes separated by 30 to 390 days. This temporal separation requirement is designed to improve diagnostic specificity by filtering out random or provisional miscoding, since a genuine narcolepsy diagnosis tends to persist across multiple encounters while a coding error is less likely to be repeated weeks or months later. The authors note that this approach aligns with prior large electronic health record based studies of narcolepsy. They also push back on the implicit assumption that the diagnostic gold standard is itself flawless, citing work by Trotti and colleagues showing that approximately half of patients with noncataplectic central hypersomnia received a different diagnosis when the multiple sleep latency test was repeated. Even in tertiary referral settings, the boundary between narcolepsy type 2 and idiopathic hypersomnia is unstable, which complicates any demand for perfect phenotypic precision.
The critics also raised the specter of confounding by conditions and medications that mimic or modify hypersomnolence: endocrine disorders such as hypothyroidism, and drug classes including antiepileptics, analgesics, antihistamines, and immunosuppressants. The authors offer a reasoned rebuttal on both fronts. Hypothyroidism and similar conditions, they argue, would more plausibly be misclassified within the general sleep clinic comparator group than within the narcolepsy cohort, where the two-code requirement reduces the likelihood of diagnostic confusion. As for the medication classes, none has a well-established direct effect on the mortality pathways most likely operative in narcolepsy, namely cardiovascular events and accident-related deaths, and adding them would have introduced collinearity without addressing the principal causal hypotheses. The study instead focused on comorbidity groups with established relevance to mortality risk in narcolepsy: cardiovascular, metabolic, pulmonary, neurologic, and psychiatric conditions. The authors caution that expanding the model to every conceivable confounder would reduce interpretability, destabilize the statistical models, and exceed the constraints of available structured data.
The exchange is a window into a broader tension in modern epidemiology. Big data cohorts built from electronic health records offer scale, duration, and real-world generalizability that carefully phenotyped prospective studies cannot match, but they sacrifice diagnostic granularity in return. The authors of the reply do not dispute that prospectively defined cohorts are essential for answering mechanistic questions and refining diagnostic precision; they simply note that such cohorts are logistically difficult to obtain for rare disorders and cannot provide the scale or quarter-century of follow-up available in the VHA population. Their study, they reiterate, was never designed to establish causality or to revalidate diagnostic criteria, but to assess whether narcolepsy, as diagnosed and managed in ordinary clinical practice, is associated with increased mortality and healthcare utilization. On that question, they argue, the answer stands: the elevated mortality persisted even against a comparator group with a higher baseline comorbidity burden, reinforcing the clinical relevance of the finding and the case that real-world evidence, carefully constructed, deserves a seat at the table in rare-disease research.
Subject of Research: Mortality risk and study methodology in narcolepsy research using large electronic health record cohorts
Article Title: Reply to “The question of whether narcolepsy increases mortality should be answered based on reliable studies”
Article References: Sharafkhaneh, A., BaHammam, A. S., Dauvilliers, Y., Azarian, M., Thorpy, M., Han, F., Senel, G. B., Aksu, M., & Razjouyan, J. (2026). Reply to “The question of whether narcolepsy increases mortality should be answered based on reliable studies”. Journal of Clinical Sleep Medicine, 22(1), Article 110. https://doi.org/10.1007/s44470-026-00121-8
Image Credits: AI Generated
DOI: 10.1007/s44470-026-00121-8
Keywords: narcolepsy, mortality, electronic health records, retrospective cohort study, propensity score matching, Veterans Health Administration, sleep medicine, computable phenotyping, polysomnography, multiple sleep latency test, confounding, hypocretin
Cite Scienmag News
Phoebe Ingram. (October 6, 2026). Narcolepsy Mortality Debate: Researchers Defend 25-Year VA Cohort Study. Scienmag. https://scienmag.com/narcolepsy-mortality-debate-researchers-defend-25-year-va-cohort-study/
Phoebe Ingram. "Narcolepsy Mortality Debate: Researchers Defend 25-Year VA Cohort Study." Scienmag, 6 October 2026, https://scienmag.com/narcolepsy-mortality-debate-researchers-defend-25-year-va-cohort-study/. Accessed 6 October 2026.
Phoebe Ingram. "Narcolepsy Mortality Debate: Researchers Defend 25-Year VA Cohort Study." Scienmag. October 6, 2026. https://scienmag.com/narcolepsy-mortality-debate-researchers-defend-25-year-va-cohort-study/

