When regulators weigh whether a new drug is safe or effective, they increasingly rely not on the tidy, controlled environment of a randomized clinical trial but on real-world data: insurance claims, electronic health records, disease registries, and pharmacy dispensing logs generated in the course of ordinary care. These data sources hold enormous promise for answering questions that trials cannot, yet they arrive fragmented, inconsistent, and scattered across institutions. Before any of it can support a regulatory decision, records belonging to the same patient must be stitched together across databases, a process known as record linkage. A new research project at George Mason University aims to tackle one of the most underappreciated hazards in that process: what happens to the credibility of the final analysis when the linkage itself is uncertain.
Brenda Betancourt, a Term Associate Professor of Statistics in George Mason’s College of Engineering and Computing, has received $286,130 from the U.S. Food and Drug Administration for a project titled Regulatory Grade Frameworks for Assessing Record Linkage Uncertainty in Real-World Evidence. The funding began in September 2026 and runs through September 2028. Over those two years, Betancourt will evaluate how uncertainty in record linkage affects the reliability, interpretability, and evidentiary confidence of analyses based on linked real-world data that are used to support FDA regulatory decision-making. In other words, the project asks a deceptively simple question: if the links between records are not certain, how certain can the conclusions drawn from them be?
The stakes are considerable. Record linkage is the invisible plumbing of real-world evidence. When a researcher wants to know whether patients exposed to a medication later experienced a particular adverse event, the exposure records typically live in one database, perhaps pharmacy claims, while the outcome records live in another, perhaps a hospital discharge registry. Joining those datasets requires deciding which rows refer to the same individual, often without a shared unique identifier such as a national patient number. Probabilistic matching algorithms compare names, dates of birth, addresses, and other fields, assigning scores that indicate how likely two records are to describe the same person. Every one of those assignments carries a probability of error, and those errors propagate silently into every downstream estimate.
Statisticians have long understood that linkage errors can bias results in either direction. False matches, in which records from two different people are merged, can dilute genuine signals or manufacture spurious associations. Missed matches, in which records from the same person remain split, can fragment a patient’s medical history and undercount exposures or outcomes. The direction and magnitude of the bias depend on the nature of the analysis, the matching variables, and the population being studied. What has been missing, particularly in the regulatory context, is a standardized framework that quantifies this uncertainty in a way that meets the evidentiary bar regulators demand when a decision affects public health.
That is the gap Betancourt’s project is designed to fill. The phrase regulatory grade in the project title signals the ambition: not merely to measure linkage uncertainty in an academic sense, but to develop frameworks rigorous and transparent enough to be used in submissions that inform FDA decisions. Regulatory science operates under constraints that academic research often does not. Methods must be reproducible, their assumptions explicit, their outputs interpretable by reviewers who must weigh evidence under statutory standards. A framework that cannot communicate how much of a reported effect might be attributable to linkage error is of limited use to an agency deciding whether to approve, restrict, or warn about a therapy.
The mathematical core of the problem is rich. Probabilistic linkage models, descendants of the classical Fellegi-Sunter framework, estimate the likelihood that a pair of records matches based on agreement and disagreement patterns across comparison fields. More modern approaches bring in Bayesian modeling, clustering, and machine learning classifiers, each producing not a single deterministic assignment but a posterior distribution over possible linkages. The challenge is propagation: how does uncertainty in those assignments flow through data cleaning, cohort construction, statistical modeling, and ultimately into confidence intervals and p-values for the regulatory question at hand? Ignoring that propagation can yield analyses that appear far more precise than they actually are, a phenomenon statisticians describe as understated uncertainty.
Several strategies exist in principle for handling this propagation. One is to condition on the most probable linkage and then adjust estimates using sensitivity analyses that vary the assumed error rates. Another is multiple imputation over plausible linkages, generating several alternative versions of the linked dataset and combining the resulting estimates so that the final inference reflects linkage ambiguity. A third embeds the linkage model and the analysis model in a joint Bayesian framework, allowing uncertainty to be integrated out rather than fixed. Each approach carries computational costs and modeling assumptions, and each performs differently depending on data quality, record volume, and the discriminative power of the matching fields. A central task for the project is evaluating when these methods deliver trustworthy inference and where they fall short.
The timing of the work reflects a broader shift in how evidence reaches regulators. The FDA has invested heavily in real-world evidence programs, including the RWE framework established under the twenty-first Century Cures Act, which directed the agency to clarify how real-world data can support approvals and label changes, particularly in areas such as rare diseases, oncology, and post-market safety surveillance where randomized trials are impractical or unethical. As the volume of linked real-world analyses grows, so does the need for standards governing every step of the pipeline. Linkage is among the earliest and least visible of those steps, which makes its errors especially insidious: they are baked in before most analysts ever see the data.
For patients and clinicians, the practical payoff of this research is confidence. If a safety signal emerges from linked claims and hospital data, clinicians and patients deserve to know whether that signal could plausibly be an artifact of mismatched records. Conversely, if a real-world analysis finds no elevated risk, regulators should know how robust that null finding is to plausible linkage errors. Quantified linkage uncertainty turns these questions from matters of intuition into matters of measurement, allowing reviewers to weigh evidence with a clearer sense of its fragility or strength. In an era when real-world evidence increasingly shapes drug labels, coverage decisions, and clinical guidelines, that clarity is not a technical luxury but a foundation of trustworthy medicine.
The project also highlights the growing role of statisticians in regulatory science and the position of universities near the federal policy apparatus in shaping it. George Mason, Virginia’s largest public research university, enrolls more than 40,000 students and sits near Washington, D.C., placing its researchers in close proximity to the agencies whose standards define the evidentiary landscape. Over the next two years, Betancourt’s work will contribute to a question that touches every linked dataset behind every regulatory submission: how sure can we be that the records we joined belong together, and how should that residual doubt be reflected in the evidence we act upon. The answer may help determine how much trust real-world evidence ultimately earns.
Subject of Research: Quantifying record linkage uncertainty in real-world evidence for FDA regulatory decision-making
Article Title: Betancourt studying regulatory grade frameworks for assessing record linkage uncertainty in real-world evidence
Article References: Betancourt studying regulatory grade frameworks for assessing record linkage uncertainty in real-world evidence. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: record linkage, real-world evidence, FDA, regulatory science, statistics, data uncertainty, George Mason University, probabilistic matching, pharmacovigilance, electronic health records, Bayesian methods, drug safety
Cite Scienmag News
Reid Dalton. (October 6, 2026). Statistician Wins FDA Funding to Quantify Record Linkage Uncertainty in Real-World Evidence. Scienmag. https://scienmag.com/statistician-wins-fda-funding-to-quantify-record-linkage-uncertainty-in-real-world-evidence/
Reid Dalton. "Statistician Wins FDA Funding to Quantify Record Linkage Uncertainty in Real-World Evidence." Scienmag, 6 October 2026, https://scienmag.com/statistician-wins-fda-funding-to-quantify-record-linkage-uncertainty-in-real-world-evidence/. Accessed 6 October 2026.
Reid Dalton. "Statistician Wins FDA Funding to Quantify Record Linkage Uncertainty in Real-World Evidence." Scienmag. October 6, 2026. https://scienmag.com/statistician-wins-fda-funding-to-quantify-record-linkage-uncertainty-in-real-world-evidence/








