Foundation models are promising to become the “universal translators” of biomedical data—but a new perspective argues that medicine should resist treating them as all-knowing digital doctors. In a review published in Nature Biomedical Engineering, researchers describe how large, adaptable artificial intelligence systems are reshaping biomedical imaging while warning that impressive benchmark scores can conceal serious weaknesses in clinical environments. Their central message is direct: foundation models are most likely to improve healthcare by augmenting specialists, not replacing them.
Foundation models are designed to learn broad representations from enormous and diverse datasets before being adapted to specific tasks. In biomedical imaging, those data may include magnetic resonance imaging, computed tomography, X-rays, ultrasound, digital pathology slides and ophthalmic images. The same underlying model could theoretically identify tumors, segment organs, estimate disease risk, retrieve similar cases and generate clinical reports. The vision becomes even broader when imaging is combined with pathology, electronic health records and genomic information, creating a composite system intended to analyze a patient across multiple biological scales.
That ambition reflects a major change from traditional medical AI. Earlier systems were generally built for one narrowly defined task, such as detecting pneumonia on chest radiographs or outlining a brain lesion on MRI scans. Foundation models instead attempt to create a reusable backbone that can be transferred across diseases, hospitals and imaging technologies. Technically, they learn statistical patterns in high-dimensional data, often through self-supervised training in which the system predicts missing, transformed or associated information. Afterward, developers fine-tune or prompt the model for particular clinical applications.
But biomedical imaging is not a single, uniform world. Images differ according to scanner manufacturer, acquisition protocol, patient population, disease prevalence and local clinical practice. A model trained primarily on data from large academic hospitals may perform very differently in community clinics, rural settings or countries with limited equipment. This problem, known as domain shift, can occur even when images appear visually similar. Small changes in image quality, contrast, hardware or patient demographics may alter the statistical distribution learned by the model and undermine its predictions.
The authors introduce a framework called real-world evaluation and assessment of foundation models, or REAL-FM, to examine whether these systems are ready for practical use. Rather than focusing only on accuracy scores, REAL-FM considers several dimensions at once, including the quality and representativeness of training data, technical readiness, clinical value, integration into existing workflows and responsible artificial intelligence. The framework is intended to help clinicians interpret claims about new models and to push developers toward evaluations that resemble actual medical practice.
One of the sharpest distinctions in the perspective is between pattern recognition and causal reasoning. Foundation models can become extraordinarily skilled at recognizing visual associations—for example, linking a particular texture or anatomical feature with a diagnosis present in their training data. Yet association is not the same as understanding why a disease occurs, how it will progress or whether a treatment will benefit an individual patient. A model may identify a correlation that is valid in one hospital but reflects a hidden confounder, such as a scanner type, reporting convention or patient-selection pattern, rather than a biological signal.
This limitation becomes especially dangerous when models are moved beyond simplified benchmarks. Many benchmark datasets provide carefully curated images, clear labels and narrowly defined tasks. Clinical care is messier: scans can be incomplete, diagnoses can be uncertain, records may contain contradictory information and several conditions may coexist. The perspective highlights a lack of verified generalization across such conditions, along with a shortage of prospective, outcome-based validation. In a prospective study, an AI system would be evaluated while care is actually being delivered, with researchers measuring whether it improves diagnostic accuracy, treatment decisions, patient outcomes or workflow efficiency.
Data scarcity is another obstacle. The largest foundation models require vast quantities of data, but medical information is difficult to collect, standardize and share. Patient records are fragmented across institutions, imaging data may lack reliable annotations and rare diseases are inherently underrepresented. Privacy regulations and governance requirements further complicate the creation of centralized datasets. As a result, a model may appear broadly capable while remaining poorly tested in the populations and clinical circumstances where errors could have the greatest consequences.
The researchers argue that human oversight therefore remains indispensable. In practical terms, foundation models may be most useful as clinical assistants that prioritize images for review, highlight suspicious regions, summarize longitudinal records or offer a second opinion that specialists can interrogate. Safe deployment will require transparent reporting of uncertainty, monitoring for performance drift, mechanisms for correcting errors and clear accountability when recommendations influence care. The future envisioned by the authors is not a single monolithic medical oracle, but a coordinated ecosystem of specialized AI systems, each evaluated for a defined clinical role and connected to expert-led workflows.
The perspective arrives as the biomedical AI field races toward increasingly general-purpose systems. Its warning is not that foundation models lack value, but that technical scale alone cannot establish clinical reliability. Before these models can be trusted with consequential decisions, they must demonstrate robustness across domains, usefulness in real workflows, safety under unexpected conditions and measurable benefits for patients. The proposed REAL-FM framework offers a way to separate viral demonstrations from durable medical progress—by asking not only what a model can recognize, but where it works, why it works and whether it makes care better.
Subject of Research: Foundation models and their real-world evaluation in biomedical imaging and clinical medicine.
Article Title: Foundation models in biomedical imaging: turning hype into reality
Article References: Muneer, A., Zhang, K., Hamdi, I. et al. Foundation models in biomedical imaging: turning hype into reality. Nature Biomedical Engineering 10, 1557–1575 (2026). https://doi.org/10.1038/s41551-026-01762-z
Image Credits: AI Generated
DOI: 10.1038/s41551-026-01762-z
Keywords: foundation models, biomedical imaging, medical artificial intelligence, clinical AI, multimodal AI, domain shift, causal reasoning, clinical validation, responsible AI, REAL-FM

