<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>statistical methodology &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/statistical-methodology/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 18:09:53 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>statistical methodology &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Circular Diagnosis: Why a Landmark Vasculitis Comparison May Prove Less Than It Claims</title>
		<link>https://scienmag.com/circular-diagnosis-why-a-landmark-vasculitis-comparison-may-prove-less-than-it-claims/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 18:09:53 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[ACR/EULAR 2022]]></category>
		<category><![CDATA[circular reasoning]]></category>
		<category><![CDATA[classification criteria]]></category>
		<category><![CDATA[clinical differences in vasculitis]]></category>
		<category><![CDATA[disease grouping flaws]]></category>
		<category><![CDATA[disease-specific diagnostic strategies]]></category>
		<category><![CDATA[epidemiology]]></category>
		<category><![CDATA[giant cell arteritis]]></category>
		<category><![CDATA[giant cell arteritis comparison]]></category>
		<category><![CDATA[inflammatory disease diagnosis]]></category>
		<category><![CDATA[large-vessel vasculitis]]></category>
		<category><![CDATA[longitudinal vasculitis research]]></category>
		<category><![CDATA[medical research methodology]]></category>
		<category><![CDATA[multiple testing]]></category>
		<category><![CDATA[rare vasculitis study critique]]></category>
		<category><![CDATA[referral bias]]></category>
		<category><![CDATA[rheumatology]]></category>
		<category><![CDATA[statistical analysis in vasculitis]]></category>
		<category><![CDATA[statistical methodology]]></category>
		<category><![CDATA[Takayasu arteritis]]></category>
		<category><![CDATA[Takayasu arteritis diagnosis]]></category>
		<category><![CDATA[vascular imaging in vasculitis]]></category>
		<category><![CDATA[vasculitis]]></category>
		<category><![CDATA[Vasculitis classification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217882</guid>

					<description><![CDATA[A new commentary argues that a major Italian study comparing Takayasu arteritis and giant cell arteritis reached its conclusions circularly, because the features used to classify patients were the same ones reported as differences between the diseases.]]></description>
										<content:encoded><![CDATA[<p>A statistical dispute with far-reaching implications for how medicine classifies rare inflammatory diseases has erupted over one of the largest head-to-head comparisons of Takayasu arteritis and giant cell arteritis ever assembled. The original study, conducted across three Italian centers and published in Immunity, Inflammation and Disease, compared 59 patients with Takayasu arteritis against 37 with giant cell arteritis, following them for a median of five years. Its authors concluded that the two conditions are clinically distinct entities that demand disease-specific diagnostic and therapeutic strategies. Now, in a formal comment on the paper, a team of clinicians argues that the study&#8217;s central conclusion rests on a subtle but fundamental logical flaw: the very features used to sort patients into the two groups were then reported back as evidence that the groups differ.</p>
<p>The critique, authored by Shubhendu Mohanty, Adarsh Jyoti Lakra, Hima Bindu Mantravadi, Prerna Uniyal and Dhanya Dedeepya, does not dispute the value of the underlying cohort. Assembling nearly a hundred well-characterized patients with two rare vasculitides, imaged to a common protocol and tracked over years, is genuinely difficult work, and the descriptive account of vascular distribution and treatment patterns stands as a useful contribution to the literature. The objection is narrower and more technical: it concerns what kind of question this particular study design is capable of answering, and whether the answer the authors reached was baked into the study before the first patient was enrolled.</p>
<p>The heart of the problem lies in the 2022 classification criteria from the American College of Rheumatology and the European Alliance of Associations for Rheumatology, which were used to allocate patients to the two diagnostic groups. Classification criteria are, by design, tools built to discriminate between named diseases. They are constructed from clinical and imaging features known to separate the conditions, and they are validated for that discriminating purpose. When a study then compares the groups on those same features and reports the differences as findings, the analysis becomes circular. The classifier and the outcome are the same variable, and the result is a foregone conclusion dressed in the language of discovery.</p>
<p>Age offers the clearest illustration of this circularity. The Takayasu criteria require disease onset at or below 60 years of age, while the giant cell arteritis criteria require patients to be 50 or older. The Italian study reported a median age of 33 in the Takayasu group against 76 in the giant cell arteritis group, with a p-value below 0.001. But as the commentators point out, this striking difference simply restates the entry rule. No patient over 60 could appear in the Takayasu group, and few under 50 could appear in the giant cell arteritis group. Reporting the age gap as a distinguishing feature of the two diseases adds no new biological information; it merely echoes the arithmetic of the classification scheme.</p>
<p>Vascular distribution presents a more consequential version of the same problem. The study&#8217;s methods state explicitly that the differentiation of cranial giant cell arteritis from its large-vessel form, and from Takayasu arteritis, was based on angiographic assessment. Yet the results then compare angiographic vascular distribution between the groups and report involvement of the axillary arteries, the aortic arch, the mesenteric vessels and the renal arteries as findings that distinguish the diseases. Polymyalgia rheumatica, which was present in 37.8 percent of the giant cell arteritis group and in none of the Takayasu group, falls into the same category, because it is a recognized component of the giant cell arteritis phenotype that informs the diagnosis in the first place. In each case, the feature that appears to separate the groups is the feature that was used to separate them.</p>
<p>Crucially, the commentators are not arguing that Takayasu arteritis and giant cell arteritis are the same disease. Their claim is epistemological rather than clinical: a cohort assembled through classification criteria cannot adjudicate whether the two entities are distinct, because the criteria were engineered to make them distinct. The features that could genuinely settle the question are precisely the ones the criteria do not use, and the Italian study actually measured several of them. Erythrocyte sedimentation rate and C-reactive protein, the two classic inflammatory markers, showed no significant difference between the groups, with p-values of 0.722 and 0.448 respectively. Neurologic and pulmonary manifestations likewise did not differ, and the long-term remission rate was reported as not significantly different in the abstract. An analysis restricted to these criteria-independent variables, the commentators suggest, would constitute a real test of the distinctness question, and would arguably make a more interesting paper.</p>
<p>The critique raises a second, independent statistical concern: the sheer number of comparisons. Across Tables 1, 2 and 3, the original study performed roughly thirty, eleven and ten statistical tests respectively, all evaluated at a significance threshold of 0.05 with no adjustment for multiple testing. At that volume of testing, chance alone guarantees that two or three findings will cross the significance threshold even if no true differences exist. Several of the differences that survived into the paper&#8217;s conclusion were marginal: dermatologic involvement at p equals 0.04, chronic liver disease at p equals 0.02, axillary artery involvement at p equals 0.02, and low-dose glucocorticoid use at p equals 0.04. The commentators note that Fisher&#8217;s exact test, which the authors appropriately chose for the sparse data, validates each individual test but does nothing to address the multiplicity problem. Their proposed remedy is straightforward: designate a small number of prespecified comparisons as primary, and label everything else exploratory.</p>
<p>A third concern involves confounding by referral pattern, a problem the original authors themselves acknowledged in their limitations section. One of the three participating centers is a regional tertiary referral center for Takayasu arteritis, which the authors offered as an explanation for the imbalance in group sizes. But the consequences of that arrangement extend well beyond group size. A referral center concentrates severe and refractory disease, so the Takayasu group was enriched for exactly the patients most likely to receive aggressive treatment. The study reported high-dose glucocorticoid use in 59.3 percent of Takayasu patients against 18.9 percent of giant cell arteritis patients, combination therapy with glucocorticoids plus immunosuppressants in 86.4 percent against 51.4 percent, and interventional vascular procedures in 25.4 percent against 5.4 percent. These figures, the commentators argue, reflect differences between two differently recruited populations at least as much as differences between two diseases, and the conclusion that the conditions respond differently to therapy does not follow from them.</p>
<p>The comment also catalogs a series of numerical inconsistencies that the authors are asked to reconcile at proof stage. The abstract reports the p-value for time from symptom onset to diagnosis as 0.025, while Table 1 and the results text give 0.029. The caption to Table 3 describes a cohort of 84 patients, although the study comprises 96. The subclavian artery row of Table 2 lists the giant cell arteritis proportion as 35140, presumably a typographical error for 35.14 percent. And the proportion of patients followed for at least 60 months appears as 72 percent in the abstract and discussion but 75 percent in the methods. None of these discrepancies is individually decisive, but together they underscore the importance of careful proofreading in studies that will inform clinical classification.</p>
<p>The broader lesson extends well beyond this single cohort. Classification criteria are indispensable tools for enrolling homogeneous patient groups into research studies and for standardizing clinical communication, but they are diagnostic heuristics, not ground truth. When researchers compare groups defined by such criteria on the variables that constitute them, they risk converting a definitional artifact into a purported biological finding. The commentators close with a constructive proposal: the Italian cohort remains a valuable resource, and a comparison restricted to criteria-independent features, governed by a prespecified analysis plan, would allow the data to speak to the question the study set out to answer. Until such an analysis is performed, the claim that Takayasu arteritis and giant cell arteritis require fundamentally different approaches remains plausible, but unproven by this evidence, a distinction that matters for the clinicians treating these rare and potentially devastating diseases of the large arteries.</p>
<p><strong>Subject of Research:</strong> Methodological critique of a cross-sectional comparison of Takayasu arteritis and giant cell arteritis</p>
<p><strong>Article Title:</strong> Comment on “Takayasu Arteritis and Giant Cell Arteritis: Results From a Cross‐Sectional Study of 96 Italian Patients”</p>
<p><strong>Article References:</strong> Mohanty, S., Lakra, A. J., Mantravadi, H. B., Uniyal, P., &amp; Dedeepya, D. (2026). Comment on “Takayasu Arteritis and Giant Cell Arteritis: Results From a Cross‐Sectional Study of 96 Italian Patients”. <em>Immunity, Inflammation and Disease, 14</em>(9), Article e70513. <a href="https://doi.org/10.1002/iid3.70513" rel="noopener noreferrer">https://doi.org/10.1002/iid3.70513</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1002/iid3.70513" rel="noopener noreferrer">10.1002/iid3.70513</a></p>
<p><strong>Keywords:</strong> Takayasu arteritis, giant cell arteritis, vasculitis, classification criteria, ACR/EULAR 2022, circular reasoning, multiple testing, referral bias, rheumatology, statistical methodology, large-vessel vasculitis, epidemiology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217882</post-id>	</item>
		<item>
		<title>CBL Deletion in Stage II Melanoma: Why Statistics Matter Before Calling It a Prognostic Biomarker</title>
		<link>https://scienmag.com/cbl-deletion-in-stage-ii-melanoma-why-statistics-matter-before-calling-it-a-prognostic-biomarker/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 15:37:28 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[11q23.1-3 deletion]]></category>
		<category><![CDATA[AJCC staging]]></category>
		<category><![CDATA[CBL]]></category>
		<category><![CDATA[CBL gene deletion in melanoma]]></category>
		<category><![CDATA[chromosome 11q deletions]]></category>
		<category><![CDATA[clinical implications of genetic markers]]></category>
		<category><![CDATA[genomic landscape of melanoma]]></category>
		<category><![CDATA[importance of statistical rigor in genomic studies]]></category>
		<category><![CDATA[Melanoma genetic profiling]]></category>
		<category><![CDATA[melanoma genomics]]></category>
		<category><![CDATA[melanoma tumor genetics]]></category>
		<category><![CDATA[methodology in cancer biomarker research]]></category>
		<category><![CDATA[molecular subgroups in melanoma]]></category>
		<category><![CDATA[multiple testing adjustment]]></category>
		<category><![CDATA[prognostic biomarker]]></category>
		<category><![CDATA[prognostic biomarkers in melanoma]]></category>
		<category><![CDATA[RAS-mutated melanoma]]></category>
		<category><![CDATA[relapse-free survival]]></category>
		<category><![CDATA[REMARK guidelines]]></category>
		<category><![CDATA[stage II melanoma]]></category>
		<category><![CDATA[stage II melanoma treatment]]></category>
		<category><![CDATA[statistical methodology]]></category>
		<category><![CDATA[statistical validation of cancer biomarkers]]></category>
		<category><![CDATA[subgroup analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206483</guid>

					<description><![CDATA[A new correspondence in the British Journal of Cancer cautions that unadjusted statistical significance and sub-stage confounding must be resolved before the CBL deletion can be accepted as a prognostic biomarker in stage II melanoma.]]></description>
										<content:encoded><![CDATA[<p>A deletion on chromosome 11q has been proposed as one of the most intriguing prognostic markers to emerge from genomic profiling of early-stage melanoma, but a new correspondence in the British Journal of Cancer argues that the statistical foundations beneath this claim deserve far closer scrutiny before the marker is translated into clinical practice. In a letter to the editor published on 21 September 2026, researchers Jingyi Han, Xinger Gao and Wenjun Jiang of the Department of Clinical Laboratory at the First Affiliated Hospital of Dalian Medical University respond to a comprehensive genetic landscape study of stage II melanoma, praising its scale while raising pointed methodological concerns about how a deletion within the CBL gene region was transformed from a genome-wide observation into a candidate biomarker for specific molecular subgroups of patients.</p>
<p>The original study, led by Lindner and colleagues, analysed tumours from 193 treatment-naïve patients with stage II melanoma, a disease stage in which the primary tumour is thick but has not yet spread to distant sites, and in which clinicians urgently need better tools to decide who requires intensive surveillance or adjuvant therapy. Among the many alterations catalogued in that work, one finding stood out: a deletion spanning the chromosomal region 11q23.1-3, which contains the CBL gene, was associated with relapse-free survival in the overall cohort when tested in a rigorous multivariate analysis. The Dalian-based team is quick to acknowledge the strength of that result, describing it as a compelling foundation for exploring CBL loss as a potential prognostic biomarker. Their concerns arise not from the headline finding, but from what happens when that finding is pushed into finer molecular subgroups.</p>
<p>CBL is no incidental passenger. The gene encodes an E3 ubiquitin ligase, a molecular machine that tags proteins for degradation and thereby acts as a critical regulator of signalling pathways driven by receptor tyrosine kinases. Because melanomas are classically stratified by driver mutations in BRAF, RAS and NF1, with a residual group classified as triple wild-type, any genomic alteration that appears to carry prognostic weight within one of these subtypes immediately attracts attention. The Lindner study reported that the 11q23.1-3 deletion showed a prognostic trend within the RAS-mutated subgroup of patients, and it is precisely this subgroup-specific claim that Han, Gao and Jiang dissect in their correspondence.</p>
<p>The heart of their critique concerns the difference between an unadjusted p-value and an adjusted one. In the RAS-mutated subgroup, the original report highlighted an unadjusted p-value of 0.044 for relapse-free survival, a figure that sits just below the conventional 0.05 threshold and therefore appears, at first glance, to signal genuine statistical significance. But when the analysis was corrected for multiple testing, the adjusted p-value rose to 0.178, well above the threshold that most researchers would accept as evidence of a reliable effect. The distinction is far from pedantic. When investigators test many genomic subgroups simultaneously, as happens when BRAF, RAS, NF1 and triple wild-type tumours are each interrogated for prognostic associations, the probability of stumbling across at least one apparently significant result by pure chance rises steeply. This phenomenon, known as a Type I error, is the false positive that multiple-testing adjustments are designed to suppress.</p>
<p>Han and colleagues argue that in exploratory subgroup analyses spanning multiple genomic subtypes, adjusting for multiple comparisons is generally recommended to prevent exactly these spurious discoveries. They point to the influential 2007 New England Journal of Medicine commentary by Wang, Lagakos, Ware, Hunter and Drazen on the reporting of subgroup analyses in clinical trials, a paper that has shaped how statisticians and clinicians interpret claims carved out of broader datasets. That commentary warned that subgroup findings are frequently overinterpreted, particularly when unadjusted significance levels are emphasized over corrected ones. By foregrounding the unadjusted p-value of 0.044 while the adjusted figure of 0.178 tells a more cautious story, the original presentation, the correspondents suggest, risks conveying a degree of predictive confidence in the RAS-mutated subgroup that the data do not yet support.</p>
<p>There is also the matter of sub-stage confounding, a second analytical nuance the letter raises. Stage II melanoma is not a single homogeneous category. Under the American Joint Committee on Cancer eighth edition staging system, refined in the landmark 2017 update by Gershenwald and colleagues, stage II encompasses patients with tumours of markedly different thicknesses and ulceration statuses, and these features themselves carry powerful prognostic information. When a cohort is subdivided first by molecular subtype and then examined for survival associations, imbalances in tumour thickness, ulceration or other clinicopathological variables between patients with and without the CBL region deletion can masquerade as genuine biological effects. Disentangling whether the deletion independently forecasts relapse, or merely travels alongside known risk factors that happen to cluster within the subgroup, demands careful covariate adjustment and transparent reporting of how residual confounding was handled.</p>
<p>The correspondents anchor their argument in established reporting standards, invoking the REMARK guidelines, the Reporting Recommendations for Tumor Marker Prognostic Studies published by McShane and colleagues in 2005 in the Journal of the National Cancer Institute. REMARK was developed precisely because biomarker prognostic studies have historically been plagued by small samples, selective reporting and optimistic interpretation, leading to markers that fail repeatedly upon validation. The guidelines call for complete documentation of statistical methods, prespecified hypotheses, transparent handling of multiple testing and honest characterisation of exploratory versus confirmatory findings. Emphasising adjusted p-values, Han, Gao and Jiang contend, aligns with these norms and helps readers accurately gauge the robustness of the CBL alteration within specific molecular subsets rather than being swept up in an apparently significant number.</p>
<p>None of this diminishes the value of the underlying discovery. The genomic classification of cutaneous melanoma established by The Cancer Genome Atlas Network in 2015 demonstrated that melanoma biology divides cleanly into the BRAF-mutant, RAS-mutant, NF1-mutant and triple wild-type categories, and subsequent efforts to layer prognostic information onto that framework have been a major research priority. A driver gene and biomarker candidate emerging from a 193-patient cohort of therapy-naïve stage II patients is genuinely noteworthy, particularly for a disease stage in which sentinel lymph node status and tumour thickness remain the dominant but imperfect guides to management. The Dalian team frames its letter as constructive engagement, crediting the original authors with a robust cohort and a rigorous multivariate analysis in the overall population, while urging that subgroup-level claims be contextualised with the statistical caution they require.</p>
<p>The broader lesson radiates well beyond melanoma genomics. Modern high-throughput studies routinely generate dozens or hundreds of candidate associations, and the path from an exploratory signal to a clinically actionable biomarker runs through validation in independent cohorts, replication under pre-specified analytical plans and harmonisation with existing staging and risk models. A deletion at 11q23.1-3 affecting CBL may yet prove to be a genuine driver event with prognostic power, and the original study&#8217;s evidence in the overall cohort suggests the hypothesis is worth pursuing vigorously. But as Han, Gao and Jiang make clear, the credibility of that pursuit depends on how the statistics are handled at each step, and on whether the field resists the temptation to treat a subgroup p-value of 0.044 as a verdict rather than a prompt for further, more stringently powered investigation.</p>
<p>For patients with stage II melanoma, the stakes are concrete: biomarkers of this kind could ultimately refine who is monitored most intensively, who is considered for adjuvant intervention and who can be reassured. Ensuring that such tools rest on statistically sound foundations is therefore not an academic quibble but a patient-safety issue. The correspondence, received on 5 June 2026, revised on 14 June and accepted on 3 September before publication on 21 September, stands as a reminder that in precision oncology, the rigour of the analysis is inseparable from the value of the discovery, and that the most important filters between a genomic observation and a clinical biomarker are multiple-testing correction, confounder control and disciplined adherence to reporting guidelines such as REMARK.</p>
<p><strong>Subject of Research:</strong> Methodological evaluation of the 11q23.1-3 CBL deletion as a prognostic biomarker in stage II melanoma</p>
<p><strong>Article Title:</strong> Methodological considerations in defining CBL as a prognostic biomarker in stage II melanoma</p>
<p><strong>Article References:</strong> Methodological considerations in defining CBL as a prognostic biomarker in stage II melanoma. (n.d.). <a href="https://doi.org/10.1038/s41416-026-03628-2" rel="noopener noreferrer">https://doi.org/10.1038/s41416-026-03628-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41416-026-03628-2" rel="noopener noreferrer">10.1038/s41416-026-03628-2</a></p>
<p><strong>Keywords:</strong> stage II melanoma, CBL, 11q23.1-3 deletion, prognostic biomarker, RAS-mutated melanoma, multiple testing adjustment, REMARK guidelines, subgroup analysis, relapse-free survival, AJCC staging, melanoma genomics, statistical methodology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206483</post-id>	</item>
		<item>
		<title>Hospital IT Vendors May Not Drive Digital Maturity, New Statistical Scrutiny Warns</title>
		<link>https://scienmag.com/hospital-it-vendors-may-not-drive-digital-maturity-new-statistical-scrutiny-warns/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:31:47 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[analytics and telehealth integration in hospitals]]></category>
		<category><![CDATA[cluster-robust inference]]></category>
		<category><![CDATA[digital health maturity assessment]]></category>
		<category><![CDATA[digital maturity]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[health informatics]]></category>
		<category><![CDATA[health information system vendor characteristics]]></category>
		<category><![CDATA[health IT vendors]]></category>
		<category><![CDATA[healthcare digital transformation challenges]]></category>
		<category><![CDATA[healthcare research statistical scrutiny]]></category>
		<category><![CDATA[hospital digital capabilities evaluation]]></category>
		<category><![CDATA[hospital digitalization]]></category>
		<category><![CDATA[hospital information system selection factors]]></category>
		<category><![CDATA[hospital information systems]]></category>
		<category><![CDATA[hospital IT vendor influence on hospital digital maturity]]></category>
		<category><![CDATA[impact of commercial vendors on healthcare IT]]></category>
		<category><![CDATA[Journal of Medical Systems]]></category>
		<category><![CDATA[limitations of vendor impact studies in healthcare]]></category>
		<category><![CDATA[market share]]></category>
		<category><![CDATA[methodological issues in digital health studies]]></category>
		<category><![CDATA[provider-level analysis]]></category>
		<category><![CDATA[statistical analysis in healthcare research]]></category>
		<category><![CDATA[statistical methodology]]></category>
		<category><![CDATA[vendor selection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196355</guid>

					<description><![CDATA[A new commentary in the Journal of Medical Systems argues that statistical flaws including provider-level clustering, self-inclusion, and market-size dependence may undermine claims linking health information system vendors to hospital digital maturity.]]></description>
										<content:encoded><![CDATA[<p>A short but pointed methodological commentary published in the Journal of Medical Systems is challenging how health services researchers interpret one of the more persistent questions in digital health: whether the characteristics of a hospital&#8217;s health information system vendor actually shape how digitally mature that hospital becomes. In a correspondence piece, Hao Lyu, Yaowen Hu, and Shucai Fan of Zhejiang Provincial People&#8217;s Hospital argue that a recent study linking hospital information system choice to digital maturity rests on statistical foundations that may be far shakier than its conclusions suggest, and that several subtle inferential problems could be steering the field toward confident answers the data cannot yet support.</p>
<p>The debate centers on a study by Backes and colleagues, published earlier in the same journal, which asked whether the choice of hospital information system influences digital maturity scores. Digital maturity, in this context, is typically measured through structured national assessment frameworks that grade hospitals on the sophistication of their clinical, administrative, and technical digital capabilities, from electronic documentation and data exchange to advanced analytics and telehealth integration. Because hospitals overwhelmingly rely on commercial vendors for their core information systems, the intuition that vendor characteristics, such as market presence, product breadth, or implementation experience, might correlate with maturity outcomes is compelling. Policymakers increasingly want to know whether choosing the right vendor can accelerate digital transformation, and vendor selection has become a strategic decision with multi-million-dollar consequences.</p>
<p>Lyu and colleagues do not dispute that the question matters. What they dispute is whether the analytical approach used to answer it can support the conclusions drawn. Their commentary organizes its critique around three technical pillars: provider-level inference, self-inclusion, and market-size dependence. Each of these describes a distinct way that the statistical machinery of the original analysis could produce misleading estimates, and together they form a checklist that the authors believe should be applied to any study attempting to connect vendor characteristics to hospital outcomes.</p>
<p>The first pillar concerns what the authors call provider-level inference. When researchers examine whether vendor characteristics are associated with hospital digital maturity, the vendor characteristics themselves vary at the level of the provider, not the hospital. Multiple hospitals in a dataset may share the same vendor, meaning their values on vendor-level explanatory variables are identical. Ignoring this clustering and treating each hospital as an independent observation inflates the effective sample size for vendor-level effects and can dramatically overstate statistical confidence. The correspondence points to the established literature on cluster-robust inference, including widely cited methodological guidance by MacKinnon, Nielsen, and Webb, which emphasizes that standard errors must account for the structure of the data. Recent work by Huang on the failure of cluster-robust methods in small samples adds a further caution: when the number of clusters, in this case the number of distinct vendors, is limited, even cluster-robust corrections can perform poorly, producing confidence intervals that are too narrow and p-values that are too optimistic.</p>
<p>The implications are substantial. If a study includes thousands of hospitals served by only a handful of major vendors, the effective information about vendor-level effects is bounded by the number of vendors, not the number of hospitals. Any claim that a particular vendor attribute, such as market share or product portfolio breadth, is significantly associated with maturity outcomes must survive inference procedures that respect this clustering structure. The commentary argues that without such corrections, the reported associations may reflect statistical artifacts rather than genuine market dynamics, and the field risks building policy recommendations on findings that would not replicate.</p>
<p>The second pillar, self-inclusion, addresses a subtler but equally consequential problem. In many studies of this type, the vendors being evaluated as potential drivers of digital maturity are themselves embedded in the market being studied, and in some analytical framings, entities can effectively appear on both sides of the regression equation. When a provider characteristic is derived from data that includes the very hospitals whose outcomes it is meant to predict, the explanatory variable and the outcome variable become mechanically entangled. This can induce spurious correlation: the predictor partially contains information about the outcome by construction, rather than by any real-world causal pathway. Lyu and colleagues argue that the original analysis did not adequately separate the measurement of provider characteristics from the hospital populations used to compute them, leaving open the possibility that at least part of the observed association is an artifact of this circularity rather than evidence that vendor choice shapes maturity.</p>
<p>The third pillar, market-size dependence, concerns how vendor characteristics are defined and scaled. Characteristics such as vendor market share are inherently relative quantities that depend on the size and composition of the market being measured. A vendor serving a large fraction of hospitals in one region may serve a tiny fraction in another, and the same vendor may occupy different market positions in different hospital segments. If the analysis pools heterogeneous markets or computes vendor characteristics over an ill-defined population, the resulting measures can conflate vendor quality or strategy with simple market structure. An association between market share and digital maturity might then reflect regional differences in healthcare infrastructure, funding, or policy environments, rather than any property of the vendors themselves. The commentary suggests that without careful attention to how the relevant market is delimited and how provider characteristics are normalized, the estimated relationships remain open to confounding by market size.</p>
<p>The correspondence also situates the debate within a broader international context. Assessing hospital digital maturity has become a priority across health systems, and a recent viewpoint in the Journal of Medical Internet Res compared national assessment approaches in five countries, revealing just how differently countries operationalize the concept. Meanwhile, surveys of digital health companies&#8217; experiences with electronic health record interfaces, including work published in the Journal of the American Medical Informatics Association, highlight how deeply vendor capabilities and interoperability practices shape what hospitals can actually achieve with their systems. Taken together, this literature underscores that vendor-hospital relationships are real and consequential, which makes it all the more important, the authors contend, that the statistical evidence linking them be rigorous.</p>
<p>For hospital leaders and procurement officials, the practical message is one of caution. If the association between vendor characteristics and digital maturity is weaker or less certain than early studies suggest, then decisions driven by the assumption that a particular vendor guarantees maturity gains may be misplaced. Investments in organizational readiness, staff training, workflow redesign, and governance may matter as much as or more than vendor selection, a conclusion consistent with decades of health informatics research showing that technology adoption succeeds or fails on sociotechnical grounds. The commentary does not claim that vendors are irrelevant; rather, it insists that the field currently lacks the inferential rigor needed to quantify exactly how much vendor characteristics contribute to maturity outcomes.</p>
<p>The authors of the correspondence, who report no funding and no competing interests, frame their intervention as constructive: a set of analytical safeguards, provider-level clustering with appropriate robust or hierarchical standard errors, careful exclusion of self-referential constructs, and explicit attention to market definitions, that future studies should adopt before drawing policy-relevant conclusions. As health systems worldwide pour resources into digital transformation and vendors compete to position their platforms as engines of maturity, the message from Hangzhou is clear: before declaring that system choice drives digital maturity, researchers must first ensure their statistics can legitimately make that claim. Until then, the true drivers of hospital digitalization remain an open and urgently important question.</p>
<p><strong>Subject of Research:</strong> Methodological critique of statistical associations between health information system provider characteristics and hospital digital maturity</p>
<p><strong>Article Title:</strong> Clarifying Associations Between HIS Provider Characteristics and Hospital Digital Maturity: Provider-Level Inference, Self-Inclusion, and Market-Size Dependence</p>
<p><strong>Article References:</strong> Lyu, H., Hu, Y., &amp; Fan, S. (2026). Clarifying Associations Between HIS Provider Characteristics and Hospital Digital Maturity: Provider-Level Inference, Self-Inclusion, and Market-Size Dependence. <em>Journal of Medical Systems, 50</em>(1), Article 129. <a href="https://doi.org/10.1007/s10916-026-02456-4" rel="noopener noreferrer">https://doi.org/10.1007/s10916-026-02456-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10916-026-02456-4" rel="noopener noreferrer">10.1007/s10916-026-02456-4</a></p>
<p><strong>Keywords:</strong> hospital information systems, digital maturity, health IT vendors, cluster-robust inference, provider-level analysis, market share, electronic health records, health informatics, statistical methodology, hospital digitalization, vendor selection, Journal of Medical Systems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196355</post-id>	</item>
	</channel>
</rss>
