A short letter published in the Annals of Biomedical Engineering has ignited a debate that reaches far beyond one laboratory study, touching on one of the most common and consequential errors in biomedical research: the misidentification of the experimental unit. Shun Yang and Hao Zhang, orthopedic researchers at the Central Hospital of Dalian University of Technology, have formally asked the authors of a prominent zirconia–silver implant coating study to clarify how they counted their data and how they interpreted a null statistical result. Their critique, published as a letter to the editor on 30 September 2026, argues that without this clarification, readers cannot tell whether the study truly demonstrated that a promising antibacterial coating preserves bone integration, or merely failed to detect a difference that might exist.
The study under scrutiny, led by Stefania Brogini and colleagues and published earlier in the same journal, evaluated a sol–gel zirconia–silver coating applied to titanium implants. The coating is designed to address a persistent clinical problem: implant-associated infections. Silver ions are potent antibacterial agents, and embedding silver within a ceramic zirconia matrix offers a way to deliver that antimicrobial activity at the implant surface without releasing enough metal to harm surrounding tissue. The original research reported an integrated assessment covering biocompatibility, short-term antibacterial efficacy, and osseointegration, the process by which living bone grows into direct contact with an implant surface. The authors concluded that osseointegration was preserved despite the addition of silver, a finding with obvious appeal for anyone hoping to design implants that resist infection without compromising fixation.
Yang and Zhang’s concern begins with a discrepancy in the numbers. According to the methods section of the original paper, the animal experiment involved 12 control and 12 coated implant sites distributed across 17 rats. Yet Table 2 of that paper reports 14 observations for the control group and 19 for the coated group when measuring bone-to-implant contact, the standard histological metric of osseointegration. The letter writers ask a deceptively simple question: what do those numbers represent? Are they individual implants, histological sections, microscopic images, or regions of interest within images? Each answer carries different statistical consequences, and the distinction matters enormously for how much confidence the results deserve.
The issue at stake is what statisticians call the experimental unit, the smallest entity to which a treatment is independently applied and which can serve as the basis for a valid comparison. In an implant study, the experimental unit is typically the individual animal, or at minimum the individual implant, because two implants placed in the same rat share the same circulation, immune system, healing capacity, and local bone environment. Measurements taken from multiple sections or images of the same implant are not independent observations; they are repeated measures that inflate the apparent sample size if treated as separate data points. This problem, sometimes called pseudoreplication, has been documented across cell culture and animal research, and it was the subject of a widely cited 2018 analysis in PLOS Biology by Stacy Lazic and colleagues asking what exactly the N means in such experiments.
If the 14 and 19 values in the zirconia–silver study represent histological sections or images rather than implants, the effective sample size could be far smaller than the table suggests, and the reported confidence intervals and P values would be optimistically narrow. Yang and Zhang also note that the original paper does not explain how within-animal dependence was handled, meaning whether observations from implants placed in the same rat were statistically linked. Modern guidelines for reporting animal research, including the ARRIVE 2.0 framework published in PLOS Biology in 2020, explicitly require researchers to specify the experimental unit and account for clustering, precisely because these choices determine whether a study’s statistics are interpretable at all.
The second pillar of the letter concerns a subtler but equally important point about how null results are described. The original study reported no significant difference in bone-to-implant contact between coated and control implants, with a P value of 0.84. On its face, that looks like strong evidence that the coating did nothing harmful. But Yang and Zhang point out that a non-significant difference test with a broad confidence interval does not, by itself, establish that the coating preserves osseointegration. A high P value simply means the data did not detect a difference; it does not mean the data excluded the possibility of a meaningful one. This is the classic statistical trap memorably summarized by Douglas Altman and Martin Bland in a 1995 British Medical Journal editorial: absence of evidence is not evidence of absence.
The proper tool for the claim the original authors wanted to make is a noninferiority test, in which researchers prespecify an acceptable margin, the largest reduction in bone-to-implant contact that would still be considered clinically tolerable, and then demonstrate statistically that the true difference is very likely smaller than that margin. Equivalence and noninferiority testing, as explained in a widely used primer by Eugene Walker and Amy Nowacki, requires the margin to be defined before the data are collected and the analysis to be framed around it. Without such a prespecified margin, concluding that a coating preserves osseointegration from a simple failure-to-reject analysis conflates two very different states of knowledge: failing to find a harm, and demonstrating that harm is absent or acceptably small.
What makes this exchange notable is its constructive tone and its practical implications. Yang and Zhang emphasize that the clarifications they seek require no new experiments, no additional animals, and no re-analysis beyond a transparent accounting of what was measured and how. If the 14 and 19 observations are indeed sections or images, the authors could reanalyze the data with the implant or animal as the unit, report the clustered structure, and, if appropriate, conduct a formal noninferiority analysis with a justified margin. Such a response would either shore up the original conclusion or appropriately temper it, and either outcome would strengthen the translational value of a coating technology that many in the implant field are watching closely.
The broader lesson extends well beyond zirconia–silver coatings. Implant surface research is a crowded and competitive field, with antimicrobial coatings, nanostructured textures, and bioactive chemistry all vying for clinical translation. Every one of these technologies must eventually clear the same bar: proof that the functional addition, whether silver ions or surface topography, does not compromise the biological integration that keeps an implant anchored for decades. Histomorphometric studies in rodents are the standard early evidence, and their statistical integrity depends on getting the experimental unit right. When sample sizes are small, as they inevitably are in animal work, the difference between counting implants and counting images can flip a conclusion from convincing to questionable.
Yang and Zhang close their letter by framing the issue as one that affects readers planning future implant surface evaluations, and their argument is likely to resonate with reviewers, editors, and funders who have pushed for stricter statistical hygiene in preclinical research. The letter, which received no external funding and declares no competing interests, was reviewed under Associate Editor Joel Stitzel and accepted within eleven days of submission, a pace suggesting the journal viewed the methodological point as substantive. Whether the original authors respond with a reanalysis or a clarification, the exchange serves as a compact teaching case for the field: report what your N truly is, account for the animals behind your numbers, and never mistake a wide confidence interval for a clean bill of health. In implant research, where the endpoint is measured in years of patient function, that distinction is not pedantry. It is the difference between a coating that is genuinely ready for the clinic and one that simply has not yet been proven otherwise.
Subject of Research: Statistical methodology and experimental unit analysis in a zirconia–silver titanium implant coating study
Article Title: Clarifying the Experimental Unit and the Evidence for Preserved Osseointegration in a Zirconia–Silver Coating Study
Article References: Clarifying the Experimental Unit and the Evidence for Preserved Osseointegration in a Zirconia–Silver Coating Study. (n.d.). https://doi.org/10.1007/s10439-026-04410-4
Image Credits: AI Generated
DOI: 10.1007/s10439-026-04410-4
Keywords: titanium implants, zirconia-silver coating, osseointegration, experimental unit, pseudoreplication, noninferiority testing, bone-to-implant contact, antibacterial coatings, ARRIVE guidelines, animal research statistics, implant-associated infection, Annals of Biomedical Engineering
Cite Scienmag News
Neil Sanderson. (October 1, 2026). Statisticians Challenge How an Antibacterial Implant Coating Study Counted Its Evidence. Scienmag. https://scienmag.com/statisticians-challenge-how-an-antibacterial-implant-coating-study-counted-its-evidence/
Neil Sanderson. "Statisticians Challenge How an Antibacterial Implant Coating Study Counted Its Evidence." Scienmag, 1 October 2026, https://scienmag.com/statisticians-challenge-how-an-antibacterial-implant-coating-study-counted-its-evidence/. Accessed 1 October 2026.
Neil Sanderson. "Statisticians Challenge How an Antibacterial Implant Coating Study Counted Its Evidence." Scienmag. October 1, 2026. https://scienmag.com/statisticians-challenge-how-an-antibacterial-implant-coating-study-counted-its-evidence/

