Automated machine learning has long promised to hand the power of deep learning to scientists who never trained as programmers, yet most of these tools deliver a finished model with little explanation of how it was built. A new open-source software package called xAutoGA, described in the journal Multimedia Tools and Applications, takes a different path: it not only designs neural networks automatically for sensor-based time-series data but also walks its users through the reasoning behind every choice. Developed by Seyed Mojtaba Mohasel of Montana State University and Alireza Afzal Aghaei, the framework combines a genetic algorithm, a graphical interface, rule extraction, and a locally deployed large language model into a single system that aims to educate as much as it optimizes.
The motivation comes from a persistent gap in the automated machine learning landscape. Building an accurate deep learning model requires navigating a multi-stage pipeline: formulating the problem, splitting data into training, validation, and test sets, preprocessing temporal signals, selecting architectures such as convolutional and LSTM networks, tuning hyperparameters, and finally deploying the model. Errors at any stage can compromise validity, and most existing frameworks concentrate almost exclusively on the modeling phase, leaving problem formulation and data preparation underexplored. General-purpose tools also fail to capture domain-specific nuances, and their black-box nature hinders adoption in scientific fields where transparency and reproducibility are non-negotiable.
xAutoGA targets multivariate time-series data from wearable and laboratory sensors, a staple of biomechanics and movement disorder research. The software automates participant-level splitting of data into training, validation, and test sets, a strategy shown to give more realistic estimates of how well a model generalizes to new people. It then jointly optimizes the sliding window size and overlap used to segment sensor signals together with the neural network architecture itself, searching across both time-domain and frequency-domain representations. The frequency stream applies a short-time Fourier transform with a Hann window to convert one-dimensional signals into spectrogram-like images, which are processed by a two-dimensional convolutional network, while a parallel one-dimensional network extracts temporal features; the two streams are concatenated before classification.
The optimization engine is a genetic algorithm that encodes the entire pipeline into a fixed-length chromosome. Gene blocks govern data segmentation, with window durations ranging from 0.25 to 10 seconds and overlaps of 25 to 75 percent; convolutional blocks with binary existence flags that allow layers to be included or bypassed; filter counts, kernel sizes, strides, activations, and batch normalization choices; and training parameters including batch size, learning rate, and the loss function itself. To handle imbalanced classes, a common problem in activity recognition, the chromosome can select between weighted cross-entropy and focal loss with tunable class weights. Each candidate model is evaluated by its macro-averaged F1-score on validation data, averaged over three training trials, and populations of 50 models evolve over 30 generations through tournament selection, uniform crossover, and mutation.
Benchmarks against two established automated deep learning systems, McFly and DRAGON, showed that xAutoGA achieved the highest F1-scores on the AReM, WISDM, and Opportunity datasets, with statistically significant improvements over McFly on three of five public datasets according to Wilcoxon signed-rank tests. DRAGON retained the edge on PAMAP2, a shortfall the authors attribute to fixed genetic algorithm settings that were not retuned per dataset. In controlled comparisons of optimization strategies on a search space of roughly 16,000 configurations, the genetic algorithm showed the best convergence stability, with the lowest interquartile range of final scores, and its population-based structure allows parallel evaluation of candidates, unlike the sequential Bayesian optimization baseline.
The team also stress-tested the algorithm with a deliberately deceptive synthetic landscape they call the Synchronization Trap, where only perfectly matched dual-stream architectures reach the global optimum. Here the genetic algorithm’s survival-of-the-fittest mechanism became a liability: tournament selection quickly eliminated risky dual-stream candidates in favor of a mediocre single-stream local optimum, and the algorithm found the true optimum in only 58 percent of runs, compared with 100 percent for random search and 96 percent for Bayesian optimization. The authors note that real biomechanical datasets tend to have more continuous fitness landscapes, which explains why the genetic algorithm excelled in the primary experiments despite this theoretical weakness.
What sets xAutoGA apart is its educational machinery. Hovering over any hyperparameter in the interface reveals its definition, purpose, default values, and literature references, so users learn while configuring experiments. During optimization, the software displays gene domination plots that borrow the logic of natural selection: each model is an organism, its hyperparameters are genes, and hyperparameter values that confer higher fitness visibly spread through the population over generations, much as longer giraffe necks once did. In one example, batch size 64 started as one option among five and gradually dominated the population until it was the only survivor by generation 12, giving users an intuitive picture of why the final model chose what it chose.
Beyond visualization, the system mines its own optimization history. Every model the genetic algorithm produced becomes a data point, with hyperparameter settings as features and F1-score as the label, and a linear programming-based rule learner extracts interpretable if-then rules that separate top-performing from underperforming configurations. On the AReM dataset, for instance, smaller batch sizes, larger windows and overlaps, and shallower networks with fewer neurons distinguished the best models. A locally deployed language model, chosen as Qwen3-4B to keep user data private, then translates these machine rules into plain-language rationales, explaining for example that smaller batches allow more frequent parameter updates. The authors deliberately position the language model as an explanatory agent rather than a search engine, avoiding the computational overhead and hallucination risks seen when language models drive optimization directly.
A usability survey at Montana State University, involving 19 students from biomechanics, statistics, computer science, and other fields plus two machine learning experts, offered cautious encouragement. Most participants rated the visualizations as moderately to very effective at showing how the final model was created and at increasing trust, and System Usability Scale responses revealed strong agreement that the software is well integrated and not cumbersome, though responses were split on how much prior learning is needed to get started. Experts flagged that gene domination plots could mislead if the population converges prematurely, prompting plans for restart mechanisms and diversity alerts. Future work will also validate whether extracted rules transfer across datasets, integrate additional explainability techniques such as SHAP and Grad-CAM, and formally assess how much users actually learn. With the source code publicly available on GitHub, the researchers hope the framework will help biomechanics researchers, clinicians, and engineers build trustworthy deep learning models while understanding exactly what the machine did on their behalf.
Subject of Research: Explainable automated deep learning for multivariate time-series sensor data using genetic algorithms and large language models
Article Title: xAutoGA: an automated deep learning software to educate users
Article References: Mohasel, S. M., & Aghaei, A. A. (2026). xAutoGA: an automated deep learning software to educate users. Multimedia Tools and Applications, 85(10), Article 791. https://doi.org/10.1007/s11042-026-21942-y
Image Credits: AI Generated
DOI: 10.1007/s11042-026-21942-y
Keywords: automated machine learning, explainable AI, genetic algorithm, neural architecture search, time-series classification, large language model, biomechanics, activity recognition, hyperparameter optimization, user education, wearable sensors, deep learning
Cite Scienmag News
Blake Davidson. (October 2, 2026). New Explainable AutoML Tool Teaches Users While It Builds Their Deep Learning Models. Scienmag. https://scienmag.com/new-explainable-automl-tool-teaches-users-while-it-builds-their-deep-learning-models/
Blake Davidson. "New Explainable AutoML Tool Teaches Users While It Builds Their Deep Learning Models." Scienmag, 2 October 2026, https://scienmag.com/new-explainable-automl-tool-teaches-users-while-it-builds-their-deep-learning-models/. Accessed 2 October 2026.
Blake Davidson. "New Explainable AutoML Tool Teaches Users While It Builds Their Deep Learning Models." Scienmag. October 2, 2026. https://scienmag.com/new-explainable-automl-tool-teaches-users-while-it-builds-their-deep-learning-models/

