A sweeping series of computer simulations has delivered an uncomfortable verdict to cognitive scientists: the most widely used measures of memory performance are, in many circumstances, statistically incapable of doing the job they were designed for. In a study published in Behavior Research Methods, Adva Levi, Raz Danino, Yonatan Goshen-Gottstein and colleagues at Tel Aviv University, together with Caren M. Rotello of the University of Massachusetts Amherst, ran nearly two thousand Monte Carlo simulations in which the true state of memory was known in advance. The results show that familiar indices such as d′ and corrected recognition can manufacture strong, perfectly replicable evidence for memory differences that do not exist, while a little-known measure called da stays honest in almost every scenario tested.
The heart of the problem lies in the distinction between two psychological quantities that are easily confused. In an old–new recognition task, participants judge whether each test item was previously studied. Researchers usually want to know about sensitivity, the participant’s true ability to discriminate studied targets from unstated lures. But performance is also shaped by bias, the participant’s tendency to say “old” regardless of the evidence. A liberal responder says “old” often and racks up both hits and false alarms; a conservative responder says “old” rarely. A valid sensitivity measure must ignore this bias entirely. If two experimental conditions genuinely have identical memory accuracy but differ in how willing participants are to respond, the correct conclusion is that no sensitivity difference exists, and a statistical test should produce a significant result only about 5 percent of the time, matching the conventional alpha level.
Previous simulation work had already raised alarms. Rotello and colleagues showed in 2008 that common single-point measures, which compute sensitivity from a single pair of hit rate and false alarm rate, systematically confuse bias with sensitivity. Yet those findings, cited only around 110 times, failed to change practice. In a review of 253 recognition memory papers published across 40 Springer Nature journals between 2018 and 2024, the new study found that percent correct and corrected recognition together accounted for nearly a quarter of reported measures, with d′ close behind. Not one of the dominant measures had ever been validated in a way that escaped circularity, because in real experiments the true sensitivity is unknowable except through the very measures being tested.
The theoretical foundation of the crisis is a mismatch between two models. Every dataset is generated by some underlying population of memory signals, described by a data-generating model. Three major families compete to describe recognition memory: continuous Gaussian signal detection models, including the unequal-variance signal detection model, or UVSD; discrete threshold models such as the double-high-threshold model, which underlie corrected recognition; and mixture models like the dual-process signal detection model, which combines an all-or-none recollection process with a continuous familiarity process. Meanwhile, every sensitivity measure embodies a measurement model, a set of assumptions about the relationship between hits and false alarms that can be visualized as a receiver operating characteristic curve. When the measurement model does not match the data-generating model, two operating points that truly reflect identical sensitivity can appear to fall on different curves, and bias masquerades as a genuine memory difference.
The simulations confronted this problem directly. Across 1,962 parameter combinations, each replicated 10,000 times, the researchers sampled memory signals from distributions corresponding to four data-generating models, created two conditions that were identical in sensitivity but differed in decision criteria, and then tested each condition-pair with t-tests for five different measures: corrected recognition, A′, d′, the geometric area under the ROC curve, and da. Whenever a measure was valid, false positive rates should hover near 5 percent. The single-point measures failed almost everywhere outside the narrow case where their own assumptions happened to hold: d′ survived only under equal-variance Gaussian distributions, and corrected recognition only under rectangular distributions. Crucially, under unequal-variance Gaussians, which most experts consider the most plausible account of recognition data, all of these measures produced false discovery rates far above 5 percent, climbing steadily as sample size grew and reaching a staggering 100 percent in the largest, longest simulated experiments.
That last number deserves a pause. A 100 percent Type I error rate means that every single experiment would find a significant sensitivity difference where none exists, and every replication would confirm it. This inverts the usual hope that large samples and many trials protect science from error. With invalid measures, more data simply amplify the confound, producing results of enormous psychometric reliability and near-zero validity. The authors draw a pointed parallel to the replication crisis: highly replicable findings can still be false discoveries, and decades of research effort have been spent chasing effects, from the “revelation effect” to the apparent memory advantage of emotional words, that later analyses attributed to bias rather than sensitivity. Similar measurement failures have been documented in eyewitness lineup research, reasoning, social psychology, and child welfare.
The solution the researchers champion is da, a signal detection measure that, like d′, transforms hit and false alarm rates into z-scores and takes their difference, but that abandons the indefensible equal-variance assumption. Instead, da measures the distance between target and lure distributions in units of their root-mean-square standard deviation, using the zROC slope, denoted S, to estimate the actual variance ratio for each participant. Empirical studies consistently find that target distributions are roughly 25 percent wider than lure distributions, with slopes averaging about 0.8 but varying widely across individuals, from around 0.5 to above 1. The new version of da exploits this by collecting confidence ratings, which yield multiple operating points per participant, and estimating S by linear regression on each participant’s zROC curve, requiring at least three usable points per condition. Only one parametric assumption remains: that the underlying distributions are Gaussian.
The simulations vindicated this choice. For both equal- and unequal-variance Gaussian data, da held Type I error rates at approximately 5 percent across every sample size, number of trials, criterion placement, and distributional separation tested. Under the dual-process mixture model, da also stayed near 5 percent in typical scenarios, because the ROC curve implied by that model closely resembles the UVSD curve, though errors rose toward 16 percent in extreme parameterizations with very large recollection contributions. Only for rectangular threshold distributions, where the Gaussian assumption is simply wrong, did da break down, as theory predicts. The one caveat concerned precision: with few trials and many participants, noisy estimates of S could push error rates to about 10 percent, and experiments that excluded many participants for having too few operating points showed similar inflation. The remedy, the authors argue, is straightforward: run enough trials per participant, closer to 128 than to 64, so that the slope estimate converges on the true variance ratio.
Even the seemingly assumption-free alternative fared poorly. The geometric area under the empirical ROC curve, popular in machine learning and often promoted as non-parametric, was computed by connecting adjacent operating points with straight lines and summing triangles and trapezoids, which omits area whose size depends on where the points fall. Because that omitted area shifts with bias, the measure was biased too, reaching near-total false positive rates under large criterion shifts with big samples. The researchers also decline to endorse full mathematical modeling as a universal fix, noting that in a blinded expert challenge, experienced modelers reached strikingly inconsistent conclusions about which experiments had manipulated sensitivity, bias, or neither, and that modelers in the literature overwhelmingly choose threshold models whose fit is inferior to continuous alternatives. A simple scalar measure available to every researcher, they argue, beats a fragile modeling pipeline.
The implications reach beyond recognition memory to the entire family of single-interval tasks, from perceptual discrimination and attentional cueing to lexical decision and metacognitive judgment, wherever binary judgments against a criterion are made. For the existing recognition literature, the news is sobering: a substantial share of published sensitivity effects based on binary responses and standard indices may be uninterpretable, though the authors offer limited escape routes, such as mirror effects, in which higher hit rates coexist with lower false alarm rates in a way bias cannot explain, and designs where within-list criterion shifts are implausible. For future work, the prescription is sharper. Collect confidence ratings, compute da with an individually estimated variance ratio, and let reviewers and editors retire measures whose false positives replicate perfectly forever. In the ongoing struggle to make psychology’s discoveries both replicable and true, the study suggests, the battle begins with choosing the right ruler.
Subject of Research: Validation of bias-independent sensitivity measures for recognition memory using Monte Carlo simulations
Article Title: Reject common measures like d′ and Corrected Recognition, embrace da: Simulation explorations of single- and multi-point recognition measures of sensitivity
Article References: Levi, A., Danino, R., Rotello, C. M., & Goshen-Gottstein, Y. (2026). Reject common measures like d′ and Corrected Recognition, embrace da: Simulation explorations of single- and multi-point recognition measures of sensitivity. Behavior Research Methods, 58(10), Article 287. https://doi.org/10.3758/s13428-026-03130-w
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03130-w
Keywords: recognition memory, sensitivity, response bias, signal detection theory, Monte Carlo simulations, measurement crisis, da measure, ROC curves, unequal variance model, Type I error, Behavior Research Methods, replication crisis
Cite Scienmag News
Glenn Wilkins. (September 13, 2026). Simulations show only da reliably separates memory accuracy from response bias. Scienmag. https://scienmag.com/simulations-show-only-da-reliably-separates-memory-accuracy-from-response-bias/
Glenn Wilkins. "Simulations show only da reliably separates memory accuracy from response bias." Scienmag, 13 September 2026, https://scienmag.com/simulations-show-only-da-reliably-separates-memory-accuracy-from-response-bias/. Accessed 13 September 2026.
Glenn Wilkins. "Simulations show only da reliably separates memory accuracy from response bias." Scienmag. September 13, 2026. https://scienmag.com/simulations-show-only-da-reliably-separates-memory-accuracy-from-response-bias/

