When epidemiologists wanted to prove that smoking caused lung cancer, they did not need to understand every molecular mechanism of carcinogenesis. They counted cases, standardized their measurements, and looked for statistical patterns across populations. A new concept paper published in AI & Society argues that the same logic could transform how we detect risk in deployed artificial intelligence systems. Rather than prying open the black box of a large language model, the proposal suggests treating expert–AI interactions as reportable events, compressing them into standardized data fields, and scanning the aggregate for signals of trouble — an approach the author, Kit Tempest-Walters, calls “AI epidemiology.”
The core problem the framework addresses is familiar to anyone who has followed the AI governance debate: deployed language models are opaque, and the most sophisticated tools for interpreting them — mechanistic interpretability, feature attribution methods such as SHAP — are expensive, fragile, and in some cases provably limited. Impossibility theorems for feature attribution mean that no technique can reliably reveal why a model produced a given output in every case. Meanwhile, AI incidents continue to accumulate in databases cataloging real-world failures, and audits of large language models remain labor-intensive, point-in-time exercises rather than continuous monitoring. What is missing, the paper argues, is a measurement layer: a way to turn messy, free-form conversations between professionals and AI systems into structured, comparable data that institutions can actually monitor.
The proposed solution is a measurement standardization framework built around a grammar of eight interaction fields. Three of these are input–output fields — mission, conclusion, and justification — that capture what a professional asked the model, what the model recommended, and why. The paper illustrates how these fields transfer across domains: in a clinical setting, the mission might be recommending a diagnostic workup for a persistent cough in a heavy smoker, with the conclusion being urgent imaging justified by red-flag features for malignancy; in lending, the mission might be assessing a mortgage application with a low credit score, ending in a rejection justified by underwriting thresholds; in law, evaluating a settlement offer against comparable precedent awards. Because the fields carry the same structure regardless of domain, scores produced in one sector can in principle be compared with those in another.
The remaining fields feed a scoring engine. A large language model acting as a judge — the now well-established “LLM-as-a-judge” paradigm — evaluates each standardized interaction along dimensions such as risk level, policy alignment, and evidential alignment. Policy alignment measures whether the AI’s recommendation conforms to applicable guidelines and regulations; evidential alignment measures whether the justification is actually supported by the evidence cited. Crucially, the judge does not need access to the model’s internals. It works entirely from the observable interaction, which means the framework can be applied to any deployed system, including proprietary models whose weights are hidden even from the institutions using them.
Of course, using one black-box model to grade the outputs of another black-box model introduces what the paper candidly names structural circularity. LLM judges are known to suffer from systematic biases: sycophancy, the tendency to agree with whatever a user asserts; self-preference, the tendency to favor outputs resembling their own generations; and verbosity bias, the tendency to reward longer answers regardless of quality. There is also the problem of non-determinism — even supposedly deterministic settings can produce different judgments across runs, and small prompt changes can have butterfly-effect consequences for model performance. The framework confronts these weaknesses head-on rather than pretending they do not exist.
Its answer is a set of bounded conditions designed to reduce measurement inconsistency: explicit scoring rubrics, structured chain-of-thought reasoning before each verdict, reference documents that anchor judgments to authoritative guidelines, and low-temperature generation to suppress randomness. On top of these, the paper specifies a reliability verification procedure to detect and quantify the residual biases. The statistical machinery is equally explicit: paired bootstrap inference for the main comparisons, DeLong’s test for paired areas under the receiver-operating-characteristic curve as a sensitivity check, a pre-specified one-sided non-inferiority margin of 0.05, and Holm–Bonferroni correction to control for multiple testing. Agreement statistics drawn from the reliability literature, such as intraclass correlation coefficients and the Landis–Koch benchmarks for observer agreement, provide the yardsticks for judging whether the automated judge is consistent enough to be trusted.
To demonstrate feasibility, the paper applies the protocol in a minimal form to a published corpus of expert–AI interactions: the Rheum2Guide vignette study, in which ChatGPT’s treatment recommendations for rheumatic patients were compared against specialist decisions across nineteen clinical cases covering conditions from rheumatoid arthritis to giant cell arteritis. The judge scored each interaction for risk, policy alignment, and evidential alignment, using EULAR guideline documents as the reference corpus. The key finding is narrow but important: the judge reproduced its policy and evidential alignment scores across two independent runs under the specified conditions. In other words, the measurement instrument was stable. The author is careful to stress what this does not show — reliability at scale, across domains and models, remains a question for future empirical work, and the population-level claims the framework is designed to support are the subject of a staged research program, not results claimed in this paper.
If that research program succeeds, the payoff comes in three stages. The first is a reliability claim: under bounded conditions, language models can produce dependable, standardized assessments of expert–AI interactions. The second is a governance claim: alignment scores give experts an immediate signal during deployment — a warning light that flashes when a model’s advice drifts from policy or evidence — while giving institutions a basis for monitoring alignment patterns across mission types, models, and domains. The third, and most ambitious, is the outcome validation claim: once measurement is standardized, aggregate alignment scores could be correlated with downstream outcomes in regulated professional settings, exactly as cholesterol measurements were standardized and then linked to heart disease risk in the Framingham study, and as the lipid standardization programs of the CDC and National Heart, Lung and Blood Institute made population-scale cardiovascular epidemiology possible.
This is where the epidemiological analogy does its heaviest lifting. John Snow did not need to identify the cholera vibrio to map deaths around the Broad Street pump; Doll and Hill built the case against tobacco on correlated variables long before the molecular mechanisms of smoking-induced carcinogenesis were worked out. Hill’s famous 1965 criteria for distinguishing association from causation — strength, consistency, specificity, temporality, and the rest — were tools for reasoning rigorously in the absence of mechanistic certainty. The paper proposes importing precisely this style of reasoning into AI safety: instead of waiting for interpretability research to make model internals transparent, collect standardized measurements at the interface, look for statistical associations with harm, and intervene on the patterns that emerge. It is risk detection by surveillance rather than by dissection.
The framework is honest about its own limits. The demonstration involved a single judge, a single domain, and a small corpus; the circularity of black-box judging is mitigated, not eliminated; and unfaithful chain-of-thought reasoning — language models that do not always say what they think — remains a live concern even for structured scoring pipelines. But the conceptual move is significant. AI governance has long oscillated between two poles: demanding impossible transparency from opaque systems, or resigning itself to anecdotal incident reports. A standardized measurement layer offers a third path, one that medicine took a century to build and that could give institutions a shared language for describing, comparing, and ultimately predicting AI risk. Whether AI epidemiology can graduate from concept paper to working surveillance system will depend on the empirical studies the author has now carefully specified — but the blueprint for the experiment is on the table.
Subject of Research: A measurement standardization framework for detecting risks in deployed AI systems using epidemiological methods
Article Title: Toward AI epidemiology: a measurement standardization framework for prospective risk detection
Article References: Tempest-Walters, K. (2026). Toward AI epidemiology: a measurement standardization framework for prospective risk detection. AI & SOCIETY. https://doi.org/10.1007/s00146-026-03276-3
Image Credits: AI Generated
DOI: 10.1007/s00146-026-03276-3
Keywords: AI epidemiology, measurement standardization, LLM-as-a-judge, AI governance, risk detection, alignment, large language models, epidemiological methods, AI auditing, sycophancy, reliability, AI & Society
Cite Scienmag News
Phoebe Ingram. (October 6, 2026). AI epidemiology: borrowing public health’s playbook to spot risky chatbot behavior. Scienmag. https://scienmag.com/ai-epidemiology-borrowing-public-healths-playbook-to-spot-risky-chatbot-behavior/
Phoebe Ingram. "AI epidemiology: borrowing public health’s playbook to spot risky chatbot behavior." Scienmag, 6 October 2026, https://scienmag.com/ai-epidemiology-borrowing-public-healths-playbook-to-spot-risky-chatbot-behavior/. Accessed 6 October 2026.
Phoebe Ingram. "AI epidemiology: borrowing public health’s playbook to spot risky chatbot behavior." Scienmag. October 6, 2026. https://scienmag.com/ai-epidemiology-borrowing-public-healths-playbook-to-spot-risky-chatbot-behavior/

