For more than a century, the study of human cognition has largely depended on a single snapshot: a participant walks into a laboratory, completes a battery of tests under fluorescent lights, and walks out. Researchers then treat the resulting scores as stable truths about that person’s mind. But a growing body of evidence suggests this approach misses something fundamental. Cognitive performance is not a fixed trait. It wobbles from hour to hour and day to day, shaped by stress, sleep, mood, social encounters, and the simple fact of whether someone is having a good morning or a bad one. A new study published in Behavior Research Methods takes aim at this problem with an unusually ambitious solution: a battery of nine brief cognitive tests designed to be taken dozens of times in everyday settings, complete with the psychometric groundwork researchers need to use them properly.
The research team, led by Andrew J. Aschenbrenner of Washington University in St. Louis together with colleagues including Joshua J. Jackson at the University of Georgia, calls the general approach “high-frequency cognitive assessment,” or HFCA. The idea is simple in principle: instead of measuring cognition once, sample it repeatedly and remotely, in the participant’s own environment, using tests that take only a minute or two each. Such designs have already produced striking findings. Repeated cognitive testing has revealed that subtle effects of Alzheimer’s disease pathology may be exaggerated at certain times of day, that environmental distractions and social contexts change how well people perform, and that momentary loneliness, stress, and motivation all leave measurable fingerprints on thinking speed and accuracy.
The tricky part has always been test selection. Because participants must engage with testing multiple times per day, sometimes for weeks, researchers typically keep each session under five minutes to avoid burnout. That constraint forces a choice: include several tests that each tap a different cognitive domain, or include multiple measures of just one domain. Without psychometric data on how these brief, repeated tasks actually behave, researchers have been flying somewhat blind. Is a momentary lapse on a memory test a sign of a broadly sluggish cognitive system, a memory-specific failure, or an artifact of the particular task’s stimuli or instructions? Standard laboratory batteries can answer such questions because each measure has decades of psychometric vetting behind it. For high-frequency remote testing, almost nothing of the sort existed.
Enter the Cognitive Variability Battery, or CVB, the centerpiece of the new study. The battery consists of nine tests organized into three cognitive domains: attentional control, episodic memory, and processing speed. The attentional control tasks are updated versions of classics like the Stroop and Flanker paradigms, collectively known as the “Squared” tasks because they were engineered to overcome the notoriously poor psychometric properties of traditional conflict tasks. In the Stroop task, participants name the ink color of color words while ignoring the word itself, a measure of the ability to prioritize goal-relevant information over competing stimuli. The episodic memory tests include free recall, spatial memory, and paired associates learning, while the processing speed tests feature a symbol-digit task, a number comparison task, and mental rotation. All nine tasks were coded on the Gorilla web platform and released to an open science repository, making them available to any researcher at essentially no cost, a deliberate contrast to the specialized, often expensive infrastructure that custom HFCA batteries have traditionally required.
To validate the battery, the team administered it to a heterogeneous lifespan sample of adults recruited across the United States through listservs, social media, and word of mouth. The only requirements were being at least eighteen years old and having an internet-connected device, preferably a smartphone. Participants completed cognitive tests three times per day for three full weeks. To keep any single session short, the nine tests were rotated across assessments, meaning each individual task was administered up to sixty times per participant. That volume of data is what makes the psychometric analysis possible: with fewer than a handful of observations per task, questions about reliability and variability would be statistically unanswerable.
The analyses went far beyond simple averages. For each task, the researchers calculated means, skewness, kurtosis, specific quantiles, and several distinct reliability statistics. Among the most important are between-person reliability, which describes whether a task rank-orders different people consistently over time, and within-person reliability, which captures whether an individual’s own fluctuations across occasions are systematic rather than random noise. They also computed between-person and within-person correlations, revealing whether people who excel on one task tend to excel on another, and whether a given person’s performance on two tasks rises and falls together at the same moment. Finally, they used the root mean square of successive differences to quantify the magnitude of change from one testing occasion to the next. Together, these statistics allow researchers to select tasks based purely on psychometric fit: stable, highly reliable measures for distinguishing groups by average performance, or more fluctuation-prone measures for tracking moment-to-moment changes.
That distinction matters more than it might seem. The authors emphasize that the right test depends entirely on the research question. A scientist trying to separate cognitively healthy individuals from those with impairment wants a test showing minimal fluctuation over time, so that a bad afternoon does not drag down a person’s score. A scientist studying how sleep, stress, or affect shape thinking in real time wants the opposite: a test with relatively large, systematic fluctuations. And a scientist hoping to use variability itself as a biomarker, perhaps for early neurodegenerative disease, needs tasks that fluctuate in at-risk groups but remain stable in healthy ones. The study’s context is sobering here, because prior work has shown that high cognitive variability can signal cognitive impairment or genetic risk for Alzheimer’s disease even before any overt deficits appear. Avoiding ceiling and floor effects is also critical, since a test lasting only a minute or two can artificially restrict the range of possible scores, biasing estimates of both means and variability.
To demonstrate that the battery is sensitive to real-world context, the team tested whether task performance tracked three well-established situational factors: stress, negative affect, and social engagement. These were chosen partly because their effects on cognition are already documented in the literature, allowing the researchers to check that the CVB replicates known findings, and partly because they represent modifiable lifestyle factors that may ultimately reduce dementia risk, giving the measures potential clinical relevance as targets for behavioral intervention. Alongside the cognitive tests, participants answered questions about their emotional state, social interactions, personality traits, and even completed additional measures such as logical reasoning, all of which are included in the shared dataset.
The final contribution is perhaps the most practical: a series of bootstrapped power analyses specifying, for each of the nine tests, how many participants and how many repeated assessments are needed to reliably detect effects of interest. Anyone who has designed an intensive longitudinal study knows the difficulty of these calculations, which depend on the reliability and variability structure of the specific measures used. By publishing these estimates, the team has effectively handed other researchers a design manual, indicating which tests to choose and how much data to collect, whether the goal is detecting the cognitive cost of a stressful day or charting the earliest cognitive signatures of neurodegenerative disease.
The broader significance of the work lies in its challenge to the single-assessment tradition. As the authors note, a one-time test may simply capture a person on an unrepresentative day, and even large panel studies with widely spaced assessments cannot disentangle stable traits from momentary states. Interestingly, prior research suggests that age effects on cognitive variability depend heavily on the timescale examined: older adults show more variable reaction times across trials within a single task, yet tend to be less variable than younger adults across different days, possibly because they are less able to switch response strategies mid-study, trading flexibility for consistency. Disentangling such phenomena requires exactly the kind of dense, psychometrically grounded sampling the CVB provides. With an open-source toolkit now freely available, the study’s authors hope that measuring the mind as it actually lives, fluctuating through the ordinary noise of daily life, will become the norm rather than the exception.
Cite Scienmag News
Glenn Wilkins. (September 11, 2026). New psychological test battery tracks everyday cognition. Scienmag. https://scienmag.com/new-psychological-test-battery-tracks-everyday-cognition/
Glenn Wilkins. "New psychological test battery tracks everyday cognition." Scienmag, 11 September 2026, https://scienmag.com/new-psychological-test-battery-tracks-everyday-cognition/. Accessed 11 September 2026.
Glenn Wilkins. "New psychological test battery tracks everyday cognition." Scienmag. September 11, 2026. https://scienmag.com/new-psychological-test-battery-tracks-everyday-cognition/








