Amyotrophic lateral sclerosis has long been one of the most frustrating diagnoses in medicine. The disease destroys motor neurons in the cortex, brainstem and spinal cord, producing muscle weakness, atrophy and eventual paralysis, yet clinicians still rely largely on physical signs and electromyography to confirm it. Because early symptoms vary widely and overlap with other late-age neurodegenerative disorders, patients typically wait ten to sixteen months from their first symptoms to a confirmed diagnosis. By that point, electromyography only detects lower motor neuron damage after substantial neuronal loss, while upper motor neuron degeneration remains electrophysiologically invisible. A new computational study, published in Discover Artificial Intelligence, argues that the answers to earlier detection may already be circulating in a routine blood sample, if researchers can read the data correctly.
The research team, led by Renu Yadav and colleagues at the Indian Institute of Technology (BHU) Varanasi and Symbiosis Institute of Technology Hyderabad, built a pipeline that merges transcriptomics with machine learning to hunt for ALS biomarkers. Their motivation is grounded in a practical clinical reality: cerebrospinal fluid, though molecularly rich, requires an invasive lumbar puncture with risks of infection and bleeding, whereas blood is readily accessible and well suited to repeated, longitudinal monitoring. Existing blood-based biomarkers have disappointed, however. Inflammatory markers such as IL-6, IL-8, TNF-alpha, MCP-1 and CRP show high variability across studies, with inconsistent reports of upregulation, downregulation or no change at all, and neurofilament proteins lack ALS specificity, reflecting late-stage axonal loss rather than early disease.
To sidestep these limitations, the team turned to RNA sequencing, which converts total RNA into cDNA and reads it out with next-generation sequencing technology. Compared with microarrays, which can only detect pre-designed sequences and miss novel transcripts, RNA-seq offers higher resolution, a broader detection range and lower technical variability, and it captures both coding and non-coding RNAs without prior knowledge of the transcriptome. The researchers drew on two public datasets from the NCBI Gene Expression Omnibus: GSE277709, containing 84 whole blood samples equally split between 42 motor neuron disease patients and 42 healthy controls, and GSE234297, comprising 144 peripheral blood samples of which 96 were sporadic ALS cases and 48 were controls. After quality-based filtering, 132 samples with 39,378 genes each were retained from the second dataset.
Combining datasets from different laboratories introduces a notorious problem: batch effects. Differences in sample processing, sequencing platforms, reagent lots and laboratory conditions can masquerade as biological differences, generating false biomarkers. The team addressed this with ComBat-seq, a method built on a negative binomial regression model that estimates and removes batch effects while preserving the integer nature of RNA-seq count data. That detail matters, because the older ComBat method assumes Gaussian distributions, which can produce non-integer or even negative adjusted expression values that undermine downstream interpretability. After harmonizing gene identifiers, filtering low-expression genes and merging the common 12,281 genes into a unified dataset of 216 samples, the researchers validated the correction with principal component analysis and ANOVA-based variance decomposition.
The results were striking. Before correction, the first principal component accounted for 70.2 percent of the variance, with samples clustering sharply by dataset of origin rather than by disease status. After ComBat-seq, that figure dropped to 19.9 percent, and the samples mixed in a way that allowed genuine biological differences between ALS and control groups to emerge. Quantitatively, ANOVA-based R-squared analysis showed batch-associated variance falling from 0.092 to 0.012, an 86.65 percent reduction, while biological variance remained essentially unchanged at roughly 0.036. In other words, the procedure stripped away technical noise without erasing the disease signal, a balance that is far harder to achieve than it sounds.
With a clean, unified dataset in hand, the team deployed four machine learning classifiers to identify differentially expressed genes: logistic regression, support vector machine, random forest and eXtreme Gradient Boosting. Each model ran under stratified fivefold cross-validation, with feature selection performed independently within each training fold to prevent information leakage from the test partition. Genes were ranked by feature importance, and the top 20 in each fold fed into the final model. The support vector machine came out on top, achieving a balanced accuracy, sensitivity and specificity of 80 percent, with precision of 81 percent and an F1-score of 80.54 percent. Logistic regression, random forest and XGBoost followed with balanced accuracies of 77.46, 77 and 75.82 percent respectively. The authors attribute the SVM’s edge to its strength in high-dimensional, small-sample transcriptomic settings, where maximizing the margin between classes reduces overfitting.
The machine learning hits were then cross-validated against a classical statistical analysis using DESeq2, with a significance threshold of p-value below 0.05 and an absolute log2 fold change of at least 0.6. DESeq2 alone flagged 284 significant genes from the 12,281 tested, of which only 8 were upregulated and 276 downregulated. Comparing the statistical results with the model-derived gene lists yielded 29 coding genes overall, and ultimately 18 unique coding genes supported by both approaches: COL6A1, CCND1, MXRA8, MYL9, TMTC1, GRIP2, MORC3, NPIPB5, MADCAM1, AOC3, TMEM144, MSMP, ACTR3C, CCDC17, CCDC30, RILPL1, NPIPB13 and CKLF-CMTM1. Three of these, COL6A1, CCND1 and MXRA8, were upregulated, while the remaining fifteen were downregulated. Notably, NPIPB5 appeared across all four machine learning models, and five of the eighteen genes had prior, if indirect, links to ALS in the literature.
Functional enrichment analysis using the Enrichr web tool, together with Gene Ontology, KEGG and Reactome databases, connected ten of the differentially expressed genes to 34 ALS-relevant biological pathways. The picture that emerged spans nearly every hallmark of the disease. MXRA8 was enriched in pathways governing blood-brain barrier establishment and glial cell development, hinting at neurovascular dysfunction. MORC3, a chromatin regulator, showed enrichment in cellular senescence and interferon-beta regulation, consistent with the epigenetic dysregulation documented in ALS spinal cord. COL6A1 tied to skeletal muscle fiber development and myotube formation, echoing earlier findings that the gene marks perivascular fibroblast accumulation in presymptomatic sporadic ALS. MSMP and MADCAM1 linked to lymphocyte chemotaxis and leukocyte migration, pointing to immune cell recruitment, while GRIP2’s altered expression may reflect disrupted AMPA receptor homeostasis, a central mechanism in motor neuron excitotoxicity.
Two genes stood out as central hubs. CCND1, a cyclin D1 cell-cycle regulator, was enriched in CDK4/CDK6 inhibition, PTK6-regulated cell cycle, RUNX3-regulated WNT signaling and the p14-ARF pathway, suggesting that post-mitotic motor neurons may undergo pathological cell-cycle re-entry, a process widely implicated in neurodegeneration. MYL9 connected to focal adhesion, actin cytoskeleton regulation and EPHA-mediated growth cone collapse, pathways tied to cytoskeletal disorganization, axonal guidance and neuromuscular integrity. The authors propose CCND1 and MYL9 as candidate blood-based biomarkers, though they are careful to frame them as candidates rather than validated diagnostics. The study’s limitations are acknowledged: only two datasets were used, and the team calls for validation in larger multi-center cohorts, comparisons with alternative batch correction methods, and experimental work to probe the hub genes’ functional roles. Even so, the work demonstrates that a disciplined marriage of transcriptomics and artificial intelligence, anchored by rigorous batch correction and statistical cross-validation, can surface molecular leads that clinical observation alone has missed, and it offers a reproducible template for biomarker discovery in other hard-to-diagnose neurodegenerative diseases.
Subject of Research: Artificial intelligence and transcriptomics for blood-based biomarker discovery in amyotrophic lateral sclerosis
Article Title: Computational models for biomarker identification in amyotrophic lateral sclerosis using transcriptomics and artificial intelligence
Article References: Yadav, R., Sriram Kumar, P., Pragya, P., & Ronickom, J. F. A. (2026). Computational models for biomarker identification in amyotrophic lateral sclerosis using transcriptomics and artificial intelligence. Discover Artificial Intelligence, 6(1), Article 1421. https://doi.org/10.1007/s44163-026-02415-5
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02415-5
Keywords: amyotrophic lateral sclerosis, biomarkers, transcriptomics, RNA-seq, machine learning, ComBat-seq, DESeq2, batch correction, CCND1, MYL9, gene enrichment analysis, neurodegeneration
Cite Scienmag News
Juliet Wilcox. (October 10, 2026). AI Models Mine Blood Gene Data to Uncover New ALS Biomarkers. Scienmag. https://scienmag.com/ai-models-mine-blood-gene-data-to-uncover-new-als-biomarkers/
Juliet Wilcox. "AI Models Mine Blood Gene Data to Uncover New ALS Biomarkers." Scienmag, 10 October 2026, https://scienmag.com/ai-models-mine-blood-gene-data-to-uncover-new-als-biomarkers/. Accessed 10 October 2026.
Juliet Wilcox. "AI Models Mine Blood Gene Data to Uncover New ALS Biomarkers." Scienmag. October 10, 2026. https://scienmag.com/ai-models-mine-blood-gene-data-to-uncover-new-als-biomarkers/

