One of the most quietly consequential findings in brain-computer interface research has just been published in Multimedia Tools and Applications, and its message is deceptively simple: a basic statistical transformation, applied at the right moment in the data pipeline, can be the difference between a state-of-the-art artificial intelligence system that decodes human movement from brainwaves and one that performs no better than a coin flip. Enrico Mattei and Daniele Lozzi of the University of L’Aquila systematically tested how Z-score normalization affects Transformer-based deep learning models tasked with classifying motor imagery and motor execution from electroencephalography, and the results should make every laboratory building EEG decoding systems re-examine its preprocessing choices.
The study used the Upper Limb Movement EEG Dataset, recorded at Graz University of Technology, which the authors argue is uniquely suited to this kind of controlled investigation. The dataset contains recordings from fifteen healthy subjects, aged 22 to 40, who performed six distinct movements of the right upper limb, including elbow flexion and extension, forearm supination and pronation, and hand opening and closing, with the arm supported by an exoskeleton to prevent muscle fatigue. Crucially, the data were collected with sixty-one active wet electrodes at a high sampling rate of 512 Hz and distributed in a completely raw format. That raw provenance matters enormously here: many widely used EEG benchmarks arrive with hidden preprocessing already applied, which makes it impossible to disentangle what the model is actually learning from what the pipeline has already cleaned up for it.
The experimental design was deliberately ablative. The researchers trained four architectures, the convolutional EEGNet alongside three Transformer models designed for EEG work: ContraNet, built for motor imagery decoding; Conformer, designed for motor imagery and emotion decoding; and EEGDeformer, formulated for cognitive attention detection. Each model was tested on binary classification, distinguishing movement from rest, and on a harder three-class problem separating hand opening, hand closing, and rest, under four combinations of preprocessing and normalization. Preprocessing involved a rigorous pipeline: the PREP procedure to remove and interpolate noisy channels, robust common average referencing, band-pass filtering between 1 and 100 Hz, notch filtering of power-line noise and its harmonics, and artifact removal through Independent Component Analysis using the Extended Infomax algorithm with automatic classification of artifactual components by ICLabel. No epochs were discarded, preserving perfect class balance.
The headline numbers are striking. In motor execution binary classification under the subject-mixed protocol, EEGDeformer improved from 0.51 accuracy without normalization and preprocessing to 0.89 with both applied, a relative gain of roughly 75 percent. Conformer and ContraNet showed similar trajectories, climbing from near-chance levels of around 0.52 to 0.87 and 0.85 respectively. Even EEGNet, the compact convolutional baseline, jumped from 0.55 to 0.86. Without normalization, models frequently failed to exceed the statistical significance threshold of 56 percent established for the binary task, meaning their apparent learning was indistinguishable from chance. The pattern repeated in motor imagery, though with smaller absolute gains, and in the three-class problems, where executed movements reached up to 0.66 accuracy with ContraNet while imagined movements proved considerably harder.
Why should a Transformer care so much about whether its inputs have zero mean and unit variance? The authors point to the mathematics of attention itself. Transformer models compute attention weights by passing scaled dot products of query and key vectors through a softmax function. EEG signals vary wildly in amplitude across subjects because of differences in skull thickness, electrode impedance, and neural activation strength. Feed such unstandardized signals into the attention mechanism and the dot products become extreme, saturating the softmax so that nearly all weight collapses onto a single token while the rest of the spatiotemporal sequence is effectively ignored. The model is then blind to the distributed patterns that characterize motor activity in the brain. Z-score normalization, computed channel by channel as the raw value minus the mean divided by the standard deviation, keeps those dot products in comparable ranges across subjects, allowing nuanced attention patterns to emerge.
Perhaps the most important finding, however, is that normalization alone is not enough. When Z-score scaling was applied to non-preprocessed, artifact-contaminated signals, it provided minimal benefit and occasionally made things worse. The explanation is proportional compression: when artifacts dominate the variance of a signal, the normalization parameters are determined primarily by noise rather than by neural patterns, effectively lowering the signal-to-noise ratio at the model’s input. Only after rigorous artifact removal through ICA and the PREP pipeline do the normalization statistics reflect genuine neural variability. Preprocessing and normalization thus form a synergistic pair, a two-stage sequence in which artifacts are removed first to improve signal quality, and the cleaned signals are then standardized across subjects while preserving their relative temporal and spatial structure. Neither step alone suffices; together they enabled every model to reach statistically significant performance.
The study also distinguished between two validation regimes with important practical implications. In the subject-mixed protocol, data from all participants were pooled and split, allowing intra-subject information to leak across subsets and representing a theoretical upper bound. In the stricter subject-independent protocol, twenty percent of participants were held out as a test group, and normalization parameters were fitted exclusively on training subjects and applied to validation folds, preventing any cross-subject leakage. For final evaluation on unseen test subjects, Z-score parameters were derived from the test subjects’ own unlabeled data, simulating the standard unsupervised calibration phase a new BCI user would undergo. Encouragingly, the improvements held under the strict protocol, with EEGDeformer reaching 0.85 in motor execution binary classification and all models surpassing their significance thresholds, confirming that normalization facilitates learning of genuinely generalizable, subject-independent features.
Training dynamics told a consistent story. Models trained with normalized data converged faster, with training loss dropping more steeply in early epochs, and their validation loss curves showed less oscillation and a smaller gap between training and validation performance, indicating reduced overfitting. Architecture mattered too: EEGDeformer, with its dense connections that propagate input scaling effects throughout the network, showed the largest relative gains, while ContraNet’s hybrid convolutional-transformer design delivered the best absolute multiclass motor execution accuracy. Notably, the benefits extended beyond attention-based models, with the convolutional EEGNet improving by 51 percent relative in motor execution binary classification, suggesting that input standardization aids learning across architecture families, though Transformers appear especially sensitive to it because their learned positional encodings can be disrupted by unstandardized inputs.
The authors are candid about limitations. The analysis relied on a single dataset, albeit one chosen precisely because its raw format enables a rigorous ablative study, and only a binary choice of applying or not applying Z-score was tested, leaving alternatives such as min-max scaling, robust scaling, or frequency-band-specific normalization unexplored. The normalization strategies also assume either offline processing or an unsupervised calibration buffer, and fully online pipelines with adaptive Z-score calibration remain future work. Still, the central conclusion stands with unusual force: Z-score normalization is not an optional technical detail but a critical requirement, on par with architecture selection itself, for Transformer-based EEG motor classification. For a field racing toward practical brain-controlled prosthetics and assistive devices, the message is that the humblest step in the pipeline may deserve the most attention.
Subject of Research: Effects of Z-score data normalization on Transformer-based EEG motor imagery and motor execution classification
Article Title: Effects of EEG-data normalization on EEG-Transformer-based motor classification
Article References: Mattei, E., & Lozzi, D. (2026). Effects of EEG-data normalization on EEG-Transformer-based motor classification. Multimedia Tools and Applications, 85(10), Article 796. https://doi.org/10.1007/s11042-026-21929-9
Image Credits: AI Generated
DOI: 10.1007/s11042-026-21929-9
Keywords: EEG, brain-computer interface, Transformer, Z-score normalization, motor imagery, motor execution, deep learning, preprocessing, independent component analysis, subject-independent validation, EEGDeformer, Conformer
Cite Scienmag News
Cassandra Pierce. (October 7, 2026). A Simple Scaling Step Makes or Breaks AI That Reads Brain Signals for Movement. Scienmag. https://scienmag.com/a-simple-scaling-step-makes-or-breaks-ai-that-reads-brain-signals-for-movement/
Cassandra Pierce. "A Simple Scaling Step Makes or Breaks AI That Reads Brain Signals for Movement." Scienmag, 7 October 2026, https://scienmag.com/a-simple-scaling-step-makes-or-breaks-ai-that-reads-brain-signals-for-movement/. Accessed 7 October 2026.
Cassandra Pierce. "A Simple Scaling Step Makes or Breaks AI That Reads Brain Signals for Movement." Scienmag. October 7, 2026. https://scienmag.com/a-simple-scaling-step-makes-or-breaks-ai-that-reads-brain-signals-for-movement/

