Working memory — the mind’s scratchpad for holding and manipulating information — has long been measured with pen-and-paper style tests that reveal what a person can do but say little about how the brain and body actually get there. Now a team at the Indian Institute of Technology Kharagpur has shown that the answer may be written in our eyes and our skin. By recording eye movements and electrodermal activity from 89 young adults performing a demanding memory task, and feeding those signals into machine learning models, the researchers could reliably sort people into high and low working memory groups — with accuracy reaching 97 percent in cross-validation and 94 percent on an independent held-out cohort. The study, published in the journal Cognitive Computation, offers one of the most complete demonstrations yet that everyday physiological signals can serve as objective, noninvasive biomarkers of a core cognitive trait.
The experiment centered on the dual n-back task, a notoriously taxing benchmark of working memory. Participants watched a red square flicker among nine positions in a three-by-three grid while simultaneously hearing letters through headphones. In the easier one-back condition they had to indicate whether the current stimulus matched the one presented a single trial earlier; in the harder two-back condition the comparison reached two trials back, forcing the brain to continuously update and juggle two streams of information. Those who scored at least 80 percent accuracy on their first attempt advanced to the harder condition and were labeled high working memory, while those who fell short across attempts were labeled low working memory. This adaptive structure produced two complementary classification schemes — one based on first attempts alone and one on aggregate performance — that the researchers used to test the robustness of their models.
While participants wrestled with the task, a Tobii eye tracker sampling at 120 hertz captured an unusually rich portrait of their gaze behavior: fixation durations, saccade amplitudes and velocities, blink rates, and pupil diameter. At the same time, a Shimmer galvanic skin response device recorded electrodermal activity from the non-dominant hand, tracking the tiny sweat-gland responses that betray autonomic arousal. The two data streams were synchronized within the same recording environment, and each session began with a two-minute rest period to establish individual baselines for pupil size and skin conductance. The result was a multimodal time series for every participant, capturing both the attentional mechanics of the eyes and the arousal dynamics of the sympathetic nervous system as they unfolded moment by moment.
Raw physiological data of this kind poses a serious problem for machine learning: high-capacity participants simply performed longer tasks, so their recordings contained more time points. A naive model could cheat by using recording length as a proxy for ability. The researchers engineered an elegant solution — a hybrid sequence transformation in which each participant’s time series was divided into an adaptive number of equal-length sliding windows, with features averaged within each window. This preserved the local temporal dynamics of gaze and arousal while standardizing sequence length across individuals. Features directly tied to duration, such as total fixation time and reaction time, were deliberately discarded to eliminate any residual leakage. The team then benchmarked four classifiers spanning different modeling philosophies: logistic regression, support vector machines, random forests, and XGBoost, each tuned by grid search and evaluated with ten-fold cross-validation.
The results were striking. When electrodermal features were added to eye-tracking data, the random forest model achieved a mean cross-validation accuracy of 97.14 percent for first-attempt classification, and both logistic regression and random forest reached 94.12 percent on the independent test set — a cohort of 17 participants collected separately with an identical protocol. When both task attempts were included, linear and kernel models frequently reached 94.29 percent in validation, though test accuracy dipped slightly to 88.24 percent, likely because repeated exposure dampened the unique physiological variance that distinguishes individuals. Across all setups, performance improved as temporal resolution increased from five to about fifteen windows, then plateaued — a sweet spot where the models captured enough temporal structure without drowning in noise. The convergence of linear and ensemble methods on similar accuracy suggests the underlying discriminative signal is genuinely robust, not an artifact of any single algorithm.
Interpretability came next, and here the researchers extended a standard technique called permutation feature importance. Because each physiological feature appears once per temporal window, permuting individual columns would artificially inflate importance scores by treating repeated instances as independent. Instead, the team grouped all occurrences of each feature across windows and permuted them together, measuring how much the model’s loss increased when that feature’s statistical structure was destroyed. Averaged across folds and window configurations, the analysis produced a stable ranking of biomarkers. Average fixation duration emerged as the single most powerful predictor, consistent with decades of cognitive science showing that longer fixations reflect deeper encoding and heavier memory load. Pupil diameter and saccadic velocities ranked close behind, while electrodermal measures — mean response amplitude, response proportion, and response frequency — acquired substantial weight once arousal signals entered the model.
The division of labor between the two modalities tells a coherent neuroscientific story. Eye-tracking features capture the attentional dimension of working memory: pupil dilation indexes the allocation of cognitive resources, and saccadic dynamics reflect the efficiency with which information is scanned and updated. Electrodermal activity, by contrast, taps the arousal dimension — the sympathetic nervous system’s response to cognitive effort, governed in part by the locus coeruleus–norepinephrine system and consistent with the classic Yerkes–Dodson law relating arousal to performance. When only eye data were available, pupil diameter served as the best available proxy for arousal; once direct EDA measurements were added, the model reallocated predictive weight toward them, revealing a hierarchical organization in which sympathetic markers differentiate high and low capacity more sensitively. The findings thus ground computational prediction in established frameworks, from Baddeley and Hitch’s multicomponent model to modern neuro-visceral integration theories.
Perhaps the most provocative result came from a counterfactual ablation experiment designed to probe causality rather than mere correlation. The researchers took the six most important physiological features and iteratively nudged each participant’s values toward the statistical profile of the opposite group, watching the classifier’s output probability shift over one hundred steps. An asymmetry emerged: low-capacity participants could frequently be pushed across the decision boundary and reclassified as high capacity, while high-capacity participants resisted the manipulation, their prediction probabilities declining but never dropping below the classification threshold. In other words, the physiological signatures of high working memory appear trait-like and stable, whereas those of low working memory are more malleable — at least as represented within the model’s learned feature space.
The authors are careful to frame this asymmetry as a hypothesis rather than proof that physiological interventions could durably raise working memory. Still, the interpretation resonates with independent evidence: high-capacity individuals show stronger top-down attentional control and resistance to distraction, developmental gains in working memory are underpinned by strengthened frontoparietal connectivity, and training studies suggest attentional control can be improved with practice. If the malleability observed in the model reflects real physiological flexibility, individuals with lower capacity may be the most promising candidates for targeted cognitive training or neuromodulatory interventions, while stable gaze-and-arousal signatures could serve as reliable screening markers.
The practical implications extend well beyond the laboratory. Because eye trackers and skin conductance sensors are increasingly portable and affordable, the framework points toward passive, scalable assessment of cognitive capacity in classrooms, clinics, and workplaces — no verbal answers or manual test booklets required. The researchers have released their code and pipeline openly, offering a reproducible template for applying machine learning to noisy, multimodal behavioral data. Limitations remain: the cohort consisted of healthy young adults, the counterfactual manipulation operates on model representations rather than real physiology, and no formal power analysis preceded recruitment. Yet as a proof of concept, the study marks a compelling step toward a future in which the capacity of the mind’s scratchpad can be read not from what we say, but from how our eyes move and our bodies respond.
Subject of Research: Machine learning classification of working memory capacity from eye-tracking and electrodermal activity signals
Article Title: Decoding Working Memory Capacity from Gaze and Arousal Dynamics using Machine Learning
Article References: Sharma, S., Hassan, A., & Guha, R. (2026). Decoding Working Memory Capacity from Gaze and Arousal Dynamics using Machine Learning. Cognitive Computation, 18(1), Article 116. https://doi.org/10.1007/s12559-026-10664-w
Image Credits: AI Generated
DOI: 10.1007/s12559-026-10664-w
Keywords: working memory, machine learning, eye-tracking, electrodermal activity, pupillometry, dual n-back, cognitive load, arousal, feature importance, counterfactual ablation, cognitive assessment, random forest
Cite Scienmag News
Denise Maddox. (October 7, 2026). Your Eyes and Skin May Reveal How Powerful Your Working Memory Is. Scienmag. https://scienmag.com/your-eyes-and-skin-may-reveal-how-powerful-your-working-memory-is/
Denise Maddox. "Your Eyes and Skin May Reveal How Powerful Your Working Memory Is." Scienmag, 7 October 2026, https://scienmag.com/your-eyes-and-skin-may-reveal-how-powerful-your-working-memory-is/. Accessed 7 October 2026.
Denise Maddox. "Your Eyes and Skin May Reveal How Powerful Your Working Memory Is." Scienmag. October 7, 2026. https://scienmag.com/your-eyes-and-skin-may-reveal-how-powerful-your-working-memory-is/

