Executive functioning — the umbrella term for the mental machinery behind working memory, cognitive flexibility, and inhibitory control — shapes nearly everything we do, from solving a math problem to resisting the pull of a smartphone notification. Yet measuring it has always been slow, noisy, and inefficient. A new study published in Behavior Research Methods by Robert Kasumba, Dennis L. Barbour, and colleagues at Washington University in St. Louis and their collaborators now shows, through rigorously controlled simulations, that a pair of machine learning techniques can estimate a person’s executive function profile with a fraction of the data that conventional testing demands. The findings could reshape how psychologists, clinicians, and educators assess cognition, potentially turning hour-long test batteries into brief, adaptive sessions tailored to each individual brain.
The core problem with traditional cognitive assessment is statistical. Standard batteries such as the Corsi Block-Tapping task for working memory or the Stroop test for inhibitory control treat each task as an isolated measurement. Researchers typically average across repeated trials — mean response times on the Stroop, for example — and treat trial-to-trial variability as mere noise. Structural equation models can link tasks together, but only under assumptions of linearity that the authors argue oversimplify the true architecture of cognition. The result is a testing paradigm that requires many observations per task before estimates stabilize, and one that fails entirely when a task is skipped or data are missing.
The Washington University team had previously introduced an alternative: the distributional latent variable model, or DLVM. Instead of summarizing performance with a single average, DLVM models the full distribution of an individual’s behavior on each task — capturing both central tendency and variability — and compresses that information into a low-dimensional latent space learned by a neural network. Crucially, every single observation contributes to the estimation of multiple constructs at once, because the model exploits the nonlinear dependencies that link performance across tasks. A person’s position in this learned latent space constitutes their cognitive profile, and the trained model can even generate plausible behavioral data for hypothetical profiles, functioning as a generative oracle for simulation.
Paired with DLVM is a second innovation: distributional active learning, or DALE, a Bayesian algorithm that decides which test item to administer next. DALE frames cognitive testing as sequential Bayesian inference. After each response, it updates a posterior distribution over the individual’s latent position and then selects the next trial by maximizing expected mutual information — in plain terms, it always asks the question that will do the most to shrink its own uncertainty. The algorithm can be primed with a small batch of observations spanning all tasks, in this case two samples per task, before active selection kicks in. This lineage descends from Bayesian active learning methods that have already transformed audiology and vision testing, where adaptive stimulus selection dramatically reduced the trials needed to map perceptual thresholds.
What has been missing until now is ground truth. In the team’s earlier human study, real participants produced real data, but the true underlying cognitive parameters were unknown, making it impossible to measure estimation accuracy precisely. The new paper solves this with an elegant simulation strategy. The researchers trained DLVM models on a retrospective dataset — 88 valid testing sessions in which 18 participants completed up to ten sessions of an eight-task battery over ten days via a mobile app, covering tasks such as the Paced Auditory Serial Addition Test, Countermanding, Running Span, Numerical Stroop, Magnitude Comparison, Corsi Simple and Complex Span, and Cancellation. They then sampled 88 points systematically across the learned latent space and used the model to generate the corresponding ground-truth distributional parameters, from which 240 trial-level observations per task were simulated. Every estimate could now be checked against a known answer.
The first set of analyses pitted DLVM against independent maximum likelihood estimation, or IMLE, the optimal approach if tasks truly were independent. Both models received identical data under equal allocation. The verdict was striking. With only two observations per task, DLVM with two latent dimensions achieved Kullback–Leibler divergence values below 0.200 across all tasks, consistently beating IMLE, with the biggest gains on the sigmoid-shaped span tasks that are notoriously data-hungry. DLVM held its advantage until roughly seven observations per task. More dramatically, in validation analyses DLVM needed only about 20 observations per task — 160 total — to reach near-perfect accuracy in recovering marginal distributions, whereas IMLE required about 100 per task, or 800 total. DLVM could even estimate parameters for tasks that were never administered, something IMLE fundamentally cannot do.
The second stage asked how the way data are collected changes the picture. The team compared six configurations: DLVM or IMLE, each fed by DALE’s adaptive sampling, uniform random sampling, or a traditional fixed test battery delivered in sequential blocks. DALE combined with DLVM was the clear winner in the sparse-data regime, driving KLD below 0.05 by roughly 80 observations. The adaptive algorithm concentrated trials on the complex distributional tasks that carried the most information while allocating fewer trials to simpler accuracy-based tasks, and each simulated session received its own unique, personalized battery. The contrast with conventional practice was stark: at 80 total observations, when a fixed battery had covered only three tasks, DLVM with the battery achieved a KLD of 0.148 while IMLE with the same battery sat at a catastrophic 39.8, simply because it could not infer anything about tasks it had not yet reached.
The authors were careful to probe their own assumptions. Because the simulated data were generated from a DLVM-learned latent space, DLVM might enjoy a structural home-field advantage. So they repeated the analysis using an IMLE-based generative process instead. The qualitative pattern held: IMLE performed best when recovering parameters from data generated under its own specification with abundant data, but DLVM and DALE retained their advantage whenever data were sparse. The team also examined DALE’s trajectories through latent space, finding that the algorithm made large corrections in the first trials, converged to a localized region after about 30 observations, and reliably landed in regions of high probability — even when those regions did not coincide exactly with the true latent position, a consequence of the nonlinear latent space admitting multiple equally plausible solutions. Mean root mean squared error across all 88 sessions was 1.02, and only seven sessions converged to positions with normalized negative log probability above 0.05.
Perhaps the most counterintuitive insight concerns the trade-off between model flexibility and data hunger. In most of machine learning, highly flexible models like deep neural networks need enormous datasets to converge. Here the logic inverts: IMLE, the more flexible estimator, only overtakes DLVM once more than 800 observations are available under these testing conditions, while DLVM’s deliberately constrained low-dimensional embedding extracts meaningful inference from a handful of trials. The authors even observed that random sampling eventually surpassed active learning at very large sample counts, suggesting that switching to random sampling once DALE plateaus — or refining its acquisition function — could reveal additional structure in the data. Both of the study’s pre-registered hypotheses were supported by the results.
The practical implications are considerable. DALE could support adaptive assessment in clinical screening, longitudinal monitoring of cognitive change, and large-scale educational studies where testing time is limited and participant burden matters. One examinee might receive extra Stroop trials while another gets more span or countermanding items, depending on where their performance remains most uncertain, and the output is an uncertainty-aware summary of observable performance rather than a brittle point estimate. The authors argue that future behavioral tasks should be designed to be multidimensional and fully featured so that adaptive algorithms never need to repeat an item, since unsampled regions of a feature space are almost always more informative than sampled ones. For a field long anchored to rigid, hour-long batteries, the message is clear: cognition can be measured faster, smarter, and more personally than ever before.
Subject of Research: Bayesian distributional latent variable modeling and adaptive active learning for efficient assessment of executive functioning
Article Title: Bayesian distributional models of executive functioning
Article References: Bayesian distributional models of executive functioning. (n.d.). https://doi.org/10.3758/s13428-026-03191-x
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03191-x
Keywords: executive function, Bayesian inference, active learning, latent variable models, machine learning, cognitive assessment, working memory, inhibitory control, cognitive flexibility, psychometrics, neural networks, adaptive testing
Cite Scienmag News
Glenn Wilkins. (October 5, 2026). AI-Powered Bayesian Models Slash the Time Needed to Measure Executive Function. Scienmag. https://scienmag.com/ai-powered-bayesian-models-slash-the-time-needed-to-measure-executive-function/
Glenn Wilkins. "AI-Powered Bayesian Models Slash the Time Needed to Measure Executive Function." Scienmag, 5 October 2026, https://scienmag.com/ai-powered-bayesian-models-slash-the-time-needed-to-measure-executive-function/. Accessed 5 October 2026.
Glenn Wilkins. "AI-Powered Bayesian Models Slash the Time Needed to Measure Executive Function." Scienmag. October 5, 2026. https://scienmag.com/ai-powered-bayesian-models-slash-the-time-needed-to-measure-executive-function/

