A sweeping new analysis of artificial intelligence research in mental and cognitive healthcare has revealed a field that is simultaneously promising and strikingly narrow. A scoping review published in Discover Artificial Intelligence examined studies in which AI systems fuse raw image data, such as MRI scans or facial photographs, with non-image data such as genetic profiles, clinical records, or physiological signals. The review, conducted by Sebastian Unger of Witten/Herdecke University together with Yara Hasan and Laura Anderle of Westphalian University of Applied Science, found just 19 eligible studies, and only two of them had ever been validated against the performance of real clinical professionals. For a technology frequently touted as the future of diagnosis, the gap between laboratory performance and clinical reality appears to be wide indeed.
The motivation behind the review lies in a fundamental mismatch between how AI is typically built and how medicine is actually practiced. Clinicians rarely make decisions based on a single source of information; they synthesize imaging, patient history, laboratory results, and behavioral observations. Yet most medical AI remains stubbornly unimodal. In Alzheimer’s disease research specifically, a prior review cited in the study found that multimodal data was used in only 27 percent of cases. The authors of the new review argue that this matters even more than it first appears, because the term multimodal is often applied loosely. Combining structural MRI with functional MRI, or MRI with PET imaging, merges formats that are at least broadly similar. Truly heterogeneous fusion, joining raw pixels or voxels with speech, genetic markers, or questionnaire scores, is a fundamentally different technical challenge.
Herein lies the review’s central technical distinction. Many multimodal frameworks avoid processing raw image data altogether, relying instead on pre-extracted features such as regional brain volumes or cortical thickness measurements. While mathematically convenient, this homogenization can obscure the complex spatial structure inherent in the original data. The review deliberately set a high bar: studies qualified only if they fed raw image data, whether clinical scans or non-clinical photographs and video, directly into the model alongside at least one non-image modality. Following the Joanna Briggs Institute methodology, the team searched PubMed, PsycArticles, and IEEE Xplore, screening 373 identified records down through title, abstract, and full-text stages to reach the final 19 articles, all published between 2022 and 2025.
The geographic and topical concentration of the resulting literature is remarkable. China contributed eight of the 19 studies and the United States six, with single studies each from New Zealand, the United Kingdom, India, Mexico, and Italy. More striking still is the thematic monoculture. Fourteen of the 19 articles, nearly three quarters, targeted clinical pathologies, dominated by Alzheimer’s disease and mild cognitive impairment, with additional work on Parkinson’s disease, schizophrenia, vascular cognitive impairment, and depression severity. The remainder split between neurobiological markers such as brain age and brain legibility, and mental states, covered exclusively by emotion recognition studies. Conditions like anxiety disorders, substance use, and most of psychiatry simply did not appear in the raw-image multimodal literature at all.
The authors trace this narrowness to a single root cause: dataset dependence. Fourteen of the 19 studies relied on publicly available data, and the Alzheimer’s Disease Neuroimaging Initiative, or ADNI, alone supplied data for 11 of them. ADNI has appeared in more than 6,000 publications, a triumph of open science that has inadvertently created a feedback loop. Because the dataset is standardized, well-curated, and easily accessible, researchers gravitate toward it, which in turn concentrates methodological attention on Alzheimer’s disease and MRI-based imaging. The review warns that this repeated reuse introduces real risks, including benchmark overfitting and degraded performance when models encounter dataset shift in real-world clinical environments, a phenomenon well documented during external validation of AI systems in other domains.
On the technical side, the review mapped how these systems actually combine their inputs, distinguishing three fusion levels. Early fusion merges raw data at the input stage, for example through pixel concatenation or element-wise multiplication. Late fusion combines the separate outputs of unimodal models, through vector concatenation or ensemble methods such as weighted averaging. Intermediate fusion, by far the dominant strategy at 15 of 19 studies, merges representations partway through processing. Within this intermediate category, the mechanisms were exceptionally diverse: feature-level aggregation, attention-based schemes including multi-head self-attention and cross-modal transformers, interaction-based approaches that explicitly model dependencies between modalities, and joint co-learning strategies. The authors suggest this variety reflects the absence of any consensus on optimal fusion architecture, with choices driven more by individual researcher preference than standardized evidence.
Performance results were, on the surface, impressive. Models targeting pathological conditions, particularly those trained on AD-related datasets, frequently reported accuracy exceeding 95 percent, and multimodal approaches generally outperformed unimodal baselines. Three studies achieved performance comparable to medical professionals. Yet the pattern was uneven. Emotion recognition models using facial images with text or physiological signals, and a schizophrenia classifier using multimodal MRI with genetic data, reported noticeably lower accuracy. The authors conclude that the targeted clinical domain, rather than the specific choice of data modalities, appears to be the primary driver of performance, though the overwhelming predominance of MRI-based research makes definitive comparisons impossible at this stage.
The validation gap emerges as perhaps the review’s most sobering finding. Only two studies tested their frameworks against clinical professionals in any real-world sense, one using institutional data and one using public data. The authors note that while explainable AI techniques were sometimes deployed to build trust, rigorous external clinical validation offers a valuable complement, or a viable alternative where technical transparency is difficult to achieve. They also observed that the search captured no modern vision-language models or large language models in this specific domain, suggesting that heterogeneous fusion of raw images with non-image data still rests on conventional machine learning and established neural architectures. As multimodal foundation models evolve, adapting them to these heterogeneous fusion tasks represents an obvious frontier.
The review’s limitations are candidly acknowledged. Inconsistent terminology across the included studies complicated data extraction, the database selection may have missed work published in specialized or regional repositories, and the search strategy deliberately omitted terms such as deep learning, transformer, and psychiatry in favor of broader high-yield keywords. Even so, the authors’ prescriptions are clear. Future research should prioritize rigorous, head-to-head comparison of fusion levels and mechanisms to establish which approaches, whether cross-attention or joint co-learning, deliver the most robust performance for specific data pairings. Just as importantly, the field must diversify its data foundations, embracing resources such as the UK Biobank and the Parkinson’s Progression Markers Initiative, each of which appeared only once in this review. Only by moving beyond the gravitational pull of ADNI, the authors argue, can image-centric multimodal AI mature from a collection of impressive benchmark results into clinically trusted tools spanning the full breadth of mental and cognitive health.
Subject of Research: Multimodal AI fusing raw image and non-image data for mental and cognitive healthcare
Article Title: Scoping review of multimodal AI incorporating raw images for heterogeneous data fusion in mental and cognitive healthcare
Article References: Unger, S., Hasan, Y., & Anderle, L. (2026). Scoping review of multimodal AI incorporating raw images for heterogeneous data fusion in mental and cognitive healthcare. Discover Artificial Intelligence, 6(1), Article 1419. https://doi.org/10.1007/s44163-026-02444-0
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02444-0
Keywords: multimodal AI, data fusion, mental health, cognitive healthcare, Alzheimer's disease, MRI, ADNI, intermediate fusion, scoping review, clinical validation, neurodegenerative disease, machine learning
Cite Scienmag News
Glenn Wilkins. (October 9, 2026). Multimodal AI in Mental Health Care Leans Heavily on a Single Dataset, Review Finds. Scienmag. https://scienmag.com/multimodal-ai-in-mental-health-care-leans-heavily-on-a-single-dataset-review-finds/
Glenn Wilkins. "Multimodal AI in Mental Health Care Leans Heavily on a Single Dataset, Review Finds." Scienmag, 9 October 2026, https://scienmag.com/multimodal-ai-in-mental-health-care-leans-heavily-on-a-single-dataset-review-finds/. Accessed 9 October 2026.
Glenn Wilkins. "Multimodal AI in Mental Health Care Leans Heavily on a Single Dataset, Review Finds." Scienmag. October 9, 2026. https://scienmag.com/multimodal-ai-in-mental-health-care-leans-heavily-on-a-single-dataset-review-finds/

