Abstract
Foundation models (FMs) are driving a prominent shift in biomedical imaging, from task-specific models to unified backbone models for diverse tasks. This opens an avenue to integrate imaging, pathology, clinical records and genomics data into a composite system. However, this vision contrasts sharply with modern medicine’s trajectory towards more granular sub-specialization. This tension, coupled with data scarcity, domain heterogeneity and limited interpretability, creates a gap between benchmark success and real-world clinical value. We argue that the immediate role of FMs lies in augmenting, not replacing, clinical expertise. To separate hype from reality, we introduce real-world evaluation and assessment of FMs (REAL-FM), a multi-dimensional framework assessing data, technical readiness, clinical value, workflow integration and responsible artificial intelligence. Using REAL-FM, we find that although FMs excel in pattern recognition they fall short on causal reasoning, domain robustness and safety. Clinical translation is hindered by scarce representative data for model training, unverified generalization beyond over-simplified benchmark settings and a lack of prospective outcome-based validation. This Perspective provides clinicians with a practical way to interpret FM claims, identify where these systems may safely support imaging workflows and recognize why human oversight remains indispensable. For developers, it defines the validation, workflow, safety and governance requirements that must be met before FMs can become clinically reliable tools. We envision that the path forward lies not in a monolithic medical oracle, but in coordinated subspecialist AI systems that are transparent, safe and clinically grounded.
This is a preview of subscription content, access
Access options
Access through your institution
- Purchase on SpringerLink
- Instant access to the full article PDF.
39,95 €
Prices may be subject to local taxes which are calculated during checkout
Subjects
- Translational research
- Medical imaging
- Preclinical research
Data availability
Source data are provided with this paper.
References
Awais, M. et al. Foundation models defining a New Era in vision: a survey and outlook. IEEE Trans. Pattern Anal. Mach. Intell.47, 2245–2264 (2025)
Niu, C. et al. Medical multimodal multitask foundation model for lung cancer screening. Nat. Commun.16, 1523 (2025)
Huang, J. et al. Foundation models and intelligent decision-making: progress, challenges, and perspectives. Innovation6, 100948 (2025)
De Fauw, J. et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat. Med.24, 1342–1350 (2018)
Ma, J. et al. Segment anything in medical images. Nat. Commun.15, 654 (2024)
Hirsch, L. et al. High-performance open-ol. Artif. Intell.7, e240550 (2025)
Ker, J., Wang, L., Rao, J. & Lim, T. Deep learning applications in medical image analysis. IEEE Access6, 9375–9389 (2018)
Zhou, J., Park, S., Dong, S., Tang, X. & Wei, X. Artificial intelligence-driven transformative applications in disease diagnosis technology. Med. Rev.5, 353–377 (2025)
Sun, Y., Wang, L., Li, G., Lin, W. & Wang, L. A foundation model for enhancing magnetic resonance images and downstream segmentation, registration and diagnostic tasks. Nat. Biomed. Eng.9, 521–538 (2025)
Wang, J. et al. Self-improving generative foundation model for synthetic medical image generation and clinical applications. Nat. Med.31, 609–617 (2025)
Ranschaert, E. Artificial intelligence in radiology: hype or hope? J. Belg. Soc. Radiol.102, 20 (2018)
Ng, D. et al. Today’s radiologists meet tomorrow’s AI: the promises, pitfalls, and unbridled potential. Quant. Imaging Med. Surg.11, 2775–2779 (2021)
Zhang, S. & Metaxas, D. On the challenges and perspectives of foundation models for medical image analysis. Med. Image Anal.91, 102996 (2024)
Schäfer, R. et al. Overcoming data scarcity in biomedical imaging with a foundational multi-task model. Nat. Comput. Sci.4, 495–509 (2024)
Von Eschenbach, W. J. Transparency and the black box problem: why we do not trust AI. Philos. Technol.34, 1607–1622 (2021)
Schneider, J., Meske, C. & Kuss, P. Foundation models: a new paradigm for artificial intelligence. Bus. Inf. Syst. Eng.66, 221–231 (2024)
Hassija, V. et al. Interpreting black-box models: a review on explainable artificial intelligence. Cogn. Comput.16, 45–74 (2024)
Muneer, A. et al. From classical machine learning to emerging foundation models: review on multimodal data integration for cancer research. Artif. Intell. Rev.59, 119 (2026)
Collins, G. S. et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. Br. Med. J.385, e078378 (2024)
Liu, X. et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat. Med.26, 1364–1374 (2020)
Cruz Rivera, S. et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat. Med.26, 1351–1363 (2020)
Tejani, A. S. et al. Checklist for artificial intelligence in medical imaging (CLAIM): 2024 update. Radiol. Artif. Intell.6, e240300 (2024)
Lekadir, K. et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. Br. Med. J.388, e081554 (2025)
Zambrano Chaves, J. M. et al. A clinically accessible small multimodal radiology model and evaluation metric for chest X-ray findings. Nat. Commun.16, 3108 (2025)
Azad, B. et al. Foundational models in medical imaging: A comprehensive survey and future vision. Preprint at https://doi.org/10.48550/arXiv.2310.18689 (2023)
Bian, Y., Li, J., Ye, C., Jia, X. & Yang, Q. Artificial intelligence in medical imaging: from task-specific models to large-scale foundation models. Chin. Med. J.138, 651–663 (2025)
Wang, H. et al. SAM-Med3D: a vision foundation model for general-purpose segmentation on volumetric medical images. IEEE Trans. Neural Netw. Learn. Syst.36, 17599–17612 (2025)
Zhang, S. et al. A generalist foundation model and database for open-world medical image segmentation. Nat. Biomed. Eng.10, 1026–1041 (2026)
Zhang, S. et al. A multimodal biomedical foundation model trained from fifteen million image–text pairs. N. Engl. J. Med. AI2, AIoa2400640 (2025)
Kirillov, A. et al. Segment Anything. In Proc. IEEE/CVF International Conference on Computer Vision 4015–4026 (IEEE, 2023)
Cheng, J. et al. SAM-Med2D. Preprint at https://doi.org/10.48550/arXiv.2308.16184 (2023)
Koleilat, T., Asgariandehkordi, H., Rivaz, H. & Xiao, Y. MedCLIP-SAMv2: towards universal text-driven medical image segmentation. Med. Image Anal.106, 103749 (2025)
Aleem, S. et al. Test-time adaptation with SaLIP: a cascade of SAM and CLIP for zero-shot medical image segmentation. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops 5184–5193 (IEEE, 2024)
Dai, H. et al. SAMAug: point prompt augmentation for segment anything model. Preprint at https://doi.org/10.48550/arXiv.2307.01187 (2023)
Koleilat, T., Asgariandehkordi, H., Rivaz, H. & Xiao, Y. MedCLIP-SAM: bridging text and image towards universal medical image segmentation. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 (eds Linguraru, M. G. et al.) 643–653 (Springer, 2024)
Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J. & Maier-Hein, K. H. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods18, 203–211 (2021)
Archit, A. et al. Segment Anything for Microscopy. Nat. Methods22, 579–591 (2025)
Marks, M. et al. CellSAM: a foundation model for cell segmentation. Nat. Methods22, 2585–2593 (2025)
Tiu, E. et al. Expert-level detection of pathologies from unannotated chest X-ray images2)
Ju, L. et al. Delving into out-of-distribution detection with medical vision–language models. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025 (eds Gee, J. C. et al.) 133–143 (Springer, 2025)
Taher, M. R. H., Gotway, M. B. & Liang, J. Towards foundation models learned from anatomy in medical imagingansfer: 5th MICCAI Workshop, DART 2023 (eds Koch, L. et al.) 94–104 (Springer, 2024)
Zhou, Y. et al. A foundation model for generalizable disease detection from retinal images. Nature622, 156–163 (2023)
Chen, Z. et al. Masked image modeling advances 3D medical image analysis. In Proc. 2023 IEEE/CVF Winter Conference on Applications of Computer Vision 1970–1980 (IEEE, 2023); https://doi.org/10.1109/wacv56688.2023.00201
Tu, T. et al. Towards generalist biomedical AI. N. Engl. J. Med. AI1, AIoa2300138 (2024)
Prosperi, M. et al. Causal inference and counterfactual prediction in machine learning for actionable healthcare. Nat. Mach. Intell.2, 369–375 (2020)
Lau, J. J., Gayen, S., Ben Abacha, A. & Demner-Fushman, D. A dataset of clinically generated visual questions and answers about radiology images. Sci. Data5, 180251 (2018)
He, X., Zhang, Y., Mou, L., Xing, E. & Xie, P. PathVQA: 30000+ questions for medical visual question answering. In Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing(Volume 2: Short Papers) 708–718 (Association for Computational Linguistics, 2021)
Bannur, S. et al. Learning to exploit temporal structure for biomedical vision–language processing. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition 15016–15027 (IEEE, 2023)
Wu, L. et al. A universal foundation model for grounded biomedical image interpretation. Nat. Commun. 17, 7173 (2026)
Li, C. et al. LLaVA-med: training a large language-and-vision assistant for biomedicine in one day. In Advances in Neural Information Processing Systems 36 (eds Oh, A. et al.) 28541–28564 (NeurIPS, 2023)
Yang, X. et al. A large language model for electronic health records. npj Digit. Med.5, 194 (2022)
Sellergren, A. et al. MedGemma technical report. Preprint at https://doi.org/10.48550/arXiv.2507.05201 (2025)
Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature634, 466–473 (2024)
Wu, C. et al. Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. Nat. Commun.16, 7866 (2025)
Kim, Y. et al. Medical hallucinations in foundation models and their impact on healthcare. Preprint at https://doi.org/10.48550/arXiv.2503.05777 (2025)
Nguyen, D. et al. Localizing before answering: a benchmark for grounded medical visual question answering. In Proc. 34th International Joint Conference on Artificial Intelligence (ed. Kwok, J.) 7670–7678 (IJCAI, 2025)
Chen, Y. et al. MIMO: a medical vision language model with visual referring multimodal input and pixel grounding multimodal output. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition 24732–24741 (IEEE, 2025)
Yoon, S.-Y., Lee, K. S., Bezuidenhout, A. F. & Kruskal, J. B. Spectrum of cognitive biases in diagnostic radiology. Radiographics44, e230059 (2024)
Moëll, B., Sand Aronsson, F. & Akbar, S. Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1. Front. Artif. Intell.8, 1616145 (2025)
Feng, G. et al. Towards revealing the mystery behind chain of thought: a theoretical perspective. In Advances in Neural Information Processing Systems 36 (eds Oh, A. et al.) 70757–70798 (NeurIPS, 2023)
Zhang, W. et al. A sepsis diagnosis method based on chain-of-thought reasoning using large language models. Biocybern. Biomed. Eng.45, 269–277 (2025)
Rückert, J. et al. ROCOv2: radiology objects in context version 2, an updated multimodal image dataset. Sci. Data11, 688 (2024)
Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med.31, 943–950 (2025)
Liu, J., Wang, Y., Du, J., Zhou, J. T. & Liu, Z. MedCoT: medical chain of thoughtethods in Natural Language Processing (eds Al-Onaizan, Y. et al.) 17371–17389 (Association for Computational Linguistics, 2024)
Wang, X. et al. CoVT-CXR: building chain of visual thought for interpretable chest X-ray diagnosis. Preprint at https://openreview.net/forum?id=myZNJSpiK1 (2024)
Lam, H. Y. I., Ong, X. E. & Mutwil, M. Large language models in plant biology. Trends Plant Sci.29, 1145–1155 (2024)
Hazra, R., Venturato, G., Zuidberg dos Martires, P. & De Raedt, L. Can large language models reason? A characterizationICLR, 2025)
Wu, Y. et al. When more is less: understanding chain-of-thought length in LLMs. In ICLR 2026: 14th International Conference on Learning Representations (ICLR, 2026)
Hussain, S. et al. Modern diagnostic imaging technique applications and risk factors in the medical field: a review. BioMed Res. Int.2022, 5164970 (2022)
Borden, N. M., Forseen, S. E. & Stefan, C. Imaging Anatomy of the Human Brain: A Comprehensive Atlas Including Adjacent Structures (Demos Medical, 2015)
Sanders, J., Hurwitz, L. & Samei, E. Patient-specific quantification of image quality: an automated method for measuring spatial resolution in clinical CT images. Med. Phys.43, 5330–5338 (2016)
Mühler, K., Tietjen, C., Ritter, F. & Preim, B. The medical exploration toolkit: an efficient support for visual computing in surgical planning and training. IEEE Trans. Vis. Comput. Graph.16, 133–146 (2010)
Hafner, D., Pasukonis, J., Ba, J. & Lillicrap, T. Mastering diverse control tasks through world models. Nature640, 647–653 (2025)
Wang, A. Q. et al. A framework for interpretability in machine learning for medical imaging. IEEE Access12, 53277–53292 (2024)
Trinh, Q.-H., Nguyen, M.-V., Zeng, J., Jha, D. & Bagci, U. PRS-Med: position reasoning segmentation in medical imaging. In 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops 6814–6824 (IEEE, 2026)
Thiéry, A. H., Braeu, F., Tun, T. A., Aung, T. & Girard, M. J. A. Medical application of geometric deep learning for the diagnosis of glaucoma. Transl. Vis. Sci. Technol.12, 23 (2023)
Luo, L. et al. Building universal foundation models for medical image analysis with spatially adaptive networks. Preprint at https://doi.org/10.48550/arXiv.2312.07630 (2023)
Zhang, X. et al. A foundation model for lesion segmentation on brain MRI with Mixture of Modality Experts. IEEE Trans. Med. Imaging44, 2594–2604 (2025)
Qureshi, R. et al. Thinking beyond tokens: from brain-inspired intelligence to cognitive foundations for artificial general intelligence and its societal impact. Preprint at https://doi.org/10.48550/arXiv.2507.00951 (2025)
Azizi, S. et al. Robust and data-efficient generalization of self-supervised machine learning for diagnostic imaging. Nat. Biomed. Eng.7, 756–779 (2023)
Chen, R. J. et al. Towards a general-purpose foundation model for computational pathology. Nat. Med.30, 850–862 (2024)
Juyal, D. et al. PLUTO: Pathology-Universal Transformer. In ICML 2024 Workshop on Accessible and Efficient Foundation Models for Biological Discovery (ICML, 2024)
Ma, J. et al. A generalizable pathology foundation model using a unified knowledge distillation pretraining framework. Nat. Biomed. Eng.10, 545–564 (2026)
Dippel, J. et al. RudolfV: A foundation model by pathologists for pathologists. Preprint at https://doi.org/10.48550/arXiv.2401.04079 (2024)
Shaikovski, G. et al. PRISM: a multi-modal generative foundation model for slide-level histopathology. Preprint at https://doi.org/10.48550/arXiv.2405.10254 (2024)
Lu, M. Y. et al. A visual-language foundation model for computational pathology. Nat. Med.30, 863–874 (2024)
Ding, T. et al. A multimodal whole-slide foundation model for pathology. Nat. Med.31, 3749–3761 (2025)
Sun, Y. et al. CPath-Omni: a unified multimodal foundation model for patch and whole slide image analysis in computational pathology. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition 10360–10371 (IEEE, 2025)
Xiang, J. et al. A vision–language foundation model for precision oncology. Nature638, 769–778 (2025)
Hamamci, I. E. et al. Generalist foundation models from a multimodal dataset for 3D computed tomography. Nat. Biomed. Eng.https://doi.org/10.1038/s41551-025-01599-y (2026)
Pai, S. et al. Vision foundation models for computed tomography. Preprint at https://doi.org/10.48550/arXiv.2501.09001 (2025)
Shao, H. & Hou, Q. MedSeg-R: medical image segmentation with clinical reasoning. Preprint at https://doi.org/10.48550/arXiv.2506.18669 (2025)
Wu, X. et al. Causal inference in the medical domain: a survey. Appl. Intell.54, 4911–4934 (2024)
Acharya, K., Velasquez, A. & Song, H. H. A survey on symbolic knowledge distillation of large language models. IEEE Trans. Artif. Intell.5, 5928–5948 (2024)
Wang, X. et al. MEDMKG: Benchmarking medical knowledge exploitation with multimodal knowledge graph. Preprint at https://doi.org/10.48550/arXiv.2505.17214 (2025)
Li, L. et al. Real-world data medical knowledge graph: construction and applications. Artif. Intell. Med.103, 101817 (2020)
Chen, X. et al. Knowledge boosting: rethinking medical contrastive vision–language pre-training. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2023: 26th International Conference (eds Greenspan, H. et al.) 405–415 (Springer, 2023)
Jha, D. et al. Ethical framework for responsible foundational models in medical imaging. Front. Med.12, 1544501 (2025)
Zhang, X., Acosta, J. N., Zhou, H.-Y. & Rajpurkar, P. Uncovering knowledge gaps in radiology report generation models through knowledge graphs. In Proc. 6th Conference on Health, Inference, and Learning (eds Xu, X. O. et al.) 30–42 (PMLR, 2025)
Marquis, M., Bossenko, I. & Ross, P. RadLex and SNOMED CT integration: a pilot study for standardising radiology classification. Insights Imaging16, 58 (2025)
Sai, A. B., Mohankumar, A. K. & Khapra, M. M. A survey of evaluation metrics used for NLG systems. ACM Comput. Surv.55, 26 (2023)
Zhou, H.-Y. et al. MedVersa: a generalist foundation model for diverse medical imaging tasks. N. Engl. J. Med. AI3, AIoa2500595 (2026)
Wang, Z., Wu, Z., Agarwal, D. & Sun, J. MedCLIP: contrastive learning from unpaired medical images and text. In Proc. 2022 Conference on Empirical Methods in Natural Language Processing (eds Goldberg, Y. et al.) 3876–3887 (Association for Computational Linguistics, 2022)
Zou, J. & Topol, E. J. The rise of agentic AI teammates in medicine. Lancet405, 457 (2025)
Ferber, D. et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat. Cancer6, 1337–1349 (2025)
Rood, J. E., Hupalowska, A. & Regev, A. Toward a foundation model of causal cell and tissue biology with a perturbation cell and tissue atlas. Cell187, 4520–4545 (2024)
Gupta, T. et al. The essential role of causality in foundation world models for embodied AI. Preprint at https://doi.org/10.48550/arXiv.2402.06665 (2024)
Zečević, M., Willig, M., Dhami, D. S. & Kersting, K. Causal parrots: large language models may talk causality but are not causal. Trans. Mach. Learn. Res. 1142 (2023)
Zhang, K., Xie, S., Ng, I. & Zheng, Y. Causal representation learning from multiple distributions: a general setting. In Proc. 41st International Conference on Machine Learning (eds Salakhutdinov, R. et al.) 60057–60075 (PMLR, 2024)
Glymour, C., Zhang, K. & Spirtes, P. Review of causal discovery methods based on graphical models. Front. Genet.10, 524 (2019)
Schölkopf, B. in Probabilistic and Causal Inference: The Works of Judea Pearl (eds Geffner, H. et al.) 765–804 (Association for Computing Machinery, 2022)
Zheng, Y. et al. causal-learn: causal discovery in Python. J. Mach. Learn. Res.25, 1–8 (2024)
CAS
Google ScholarShi, C., Rezai, R., Yang, J., Dou, Q. & Li, X. A survey on trustworthiness in foundation models for medical image analysis. Preprint at https://doi.org/10.48550/arXiv.2407.15851 (2024)
Ihongbe, I. E., Fouad, S., Mahmoud, T. F., Rajasekaran, A. & Bhatia, B. Evaluating explainable artificial intelligence (XAI) techniques in chest radiology imaging through a human-centered lens. PLoS ONE19, e0308758 (2024)
Chen, R. J. et al. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat. Biomed. Eng.7, 719–742 (2023)
Frazer, J. et al. Disease variant prediction with deep generative models of evolutionary data. Nature599, 91–95 (2021)
Goktas, P. & Grzybowski, A. Shaping the future of healthcare: ethical clinical challenges and pathways to trustworthy AI. J. Clin. Med.14, 1605 (2025)
Zhang, J. et al. Revisiting the trustworthiness of saliency methods in radiology AI. Radiol. Artif. Intell.6, e220221 (2024)
Zhu, E. et al. Progress and challenges of artificial intelligence in lung cancer clinical translation. npj Precis. Oncol.9, 210 (2025)
Esmaeilzadeh, P. Challenges and strategies for wide-scale artificial intelligence (AI) deployment in healthcare practices: a perspective for healthcare organizations. Artif. Intell. Med.151, 102861 (2024)
Center for Connected Medicine & KLAS Research. How Health Systems Are Navigating the Complexities of AI (Center for Connected Medicine, 2024); https://connectedmed.com/re
Ali, H., Gronlund, C. & Shah, Z. Leveraging GANs for data scarcity of COVID-19: beyond the hype. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition 659–667 (IEEE, 2023)
Roberts, M. et al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat. Mach. Intell.3, 199–217 (2021)
Finlayson, S. G. et al. Adversarial attacks on medical machine learning. Science363, 1287–1289 (2019)
Qayyum, A., Qadir, J., Bilal, M. & Al-Fuqaha, A. Secure and robust machine learning for healthcare: a survey. IEEE Rev. Biomed. Eng.14, 156–180 (2021)
Renieris, E. M., Kiron, D. & Mills, S. Building robust RAI programs as third-party AI tools proliferate. MIT Sloan Management Reviewhttps://sloanreview.mit.edu/projects/building-robust-rai-programs-as-third-party-ai-tools-proliferate/ (2023)
Fouad, S. et al. CHERIE: user-centred development of an XAI system for chest radiology through co-design. In Proc. 37th International BCS Human-Computer Interaction Conference (eds Fitton, D. & Horton, M.) 205–210 (Association for Computing Machinery, 2024)
Woodward, M. et al. How to co-design a prototype of a clinical practice tool: a framework with practical guidance and a case study. BMJ Qual. Saf.33, 258–270 (2024)
Clarkson, P. J., Coleman, R., Keates, S. & Lebbon, C. Inclusive Design: Design for the Whole Population (Springer, 2003)
Aldea, M. et al. ESMO basic requirements for AI-based biomarkers in oncology (EBAI). Ann. Oncol.37, 414–430 (2026)
Yu, K.-H., Healey, E., Leong, T.-Y., Kohane, I. S. & Manrai, A. K. Medical artificial intelligence and human values. N. Engl. J. Med.390, 1895–1904 (2024)
Kohane, I. S. & Manrai, A. K. The missing value of medical artificial intelligence. Nat. Med.31, 3962–3963 (2025)
Weiss, A. D. et al. A deep learning framework for causal inference in clinical trial design: the clinical trials uncovering real efficacy artificial intelligence large clinicogenomic foundation model. AI Precis. Oncol.2, 129–141 (2025)
Pérez-García, F. et al. Exploring scalable medical image encoders beyond text supervision. Nat. Mach. Intell.7, 119–130 (2025)
He, Y. et al. VISTA3D: a unified segmentation foundation model for 3D medical imaging. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition 20863–20873 (IEEE, 2025)
Eslami, S., Meinel, C. & de Melo, G. PubMedCLIP: how much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023 (eds Vlachos, A. & Augenstein, I.) 1181–1193 (Association for Computational Linguistics, 2023)
Zhang, X., Wu, C., Zhang, Y., Xie, W. & Wang, Y. Knowledge-enhanced visual-language pre-training on chest radiology images. Nat. Commun.14, 4542 (2023)
US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices (FDA, accessed 8 January 2026); https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
Funding
This research was partially supported by Cancer Prevention and Research Institute of Texas grant RP240117. The fundingterpretation or manuscript preparation
Author information
Authors and Affiliations
Department of Imaging Physics, The University of Texas MD Anderson Cancer Center, Houston, TX, USA
Amgad Muneer, Kai Zhang, Ibraheem Hamdi, Muhammad Waqas & Jia Wu
Pediatric Surgical Research Laboratories, Massachusetts General Hospital, Boston, MA, USA
Rizwan Qureshi
Department of Computer Science, Salim Habib University, Karachi, Pakistan
Rizwan Qureshi
School of Computer Science and Digital Technologies, Aston Centre for Artificial Intelligence Research and Application, Aston University, Birmingham, UK
Shereen Fouad
School of Computing Data and Mathematical Sciences, University of Stirling, Stirling, UK
Hazrat Ali
School of Medicine and Health Sciences, George Washington University, Washington, DC, USA
Syed Muhammad Anwar
Sheikh Zayed Institute for Pediatric Surgical Innovation, Children’s National Hospital, Washington, DC, USA
Syed Muhammad Anwar
Department of Thoracic/Head and Neck Medical Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX, USA
Jia Wu
Authors
- Kai ZhangView author publications
Search author on:PubMed Google Scholar
- Ibraheem HamdiView author publications
Search author on:PubMed Google Scholar
- Rizwan QureshiView author publications
Search author on:PubMed Google Scholar
- Hazrat AliView author publications
Search author on:PubMed Google Scholar
- Jia WuView author publications
Search author on:PubMed Google Scholar
Contributions
A.M., K.Z., R.Q., S.M.A. and J.W. conceived of the scope and central thesis of this Perspective. A.M. led the conceptual development of the manuscript, including the critical analysis of FMs, taxonomy of reasoning and discussion of causality, trustworthiness and deployment challenges in biomedical imaging. A.M., K.Z., I.H., R.Q. and S.M.A. contributed to the evaluation of current FM paradigms, limitations and emerging trends across imaging modalities. A.M., S.F., H.A., S.M.A. and M.W. provided domain expertise on clinical relevance, validation practices and translational considerations. J.W. supervised the overall direction of the work and provided strategic guidance on clinical and methodological framing. All authors contributed to the writing, critical revision and intellectual refinement of the manuscript and approved the final version for publication.
Ethics declarations
Competing interests
The authors declare no competing interests
Peer review
Peer review information
Nature Biomedical Engineering thanks Todd Hollon, Guangyu Wang and Munib Mesinovic for their contribution to the peer review of this work
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations
Supplementary information
Supplementary Information (download PDF )
Supplementary Figs. 1–3 and Notes 1–3
Supplementary Data 1 (download XLSX )
REAL-FM scoring rubric and FM evaluation workbook, including criterion-level scores, reviewer agreement statistics, adjudicated final scores and written justifications
Source data
Source Data Fig. 2 (download XLSX )
Statisticalluation
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law
About this article
Cite this article
Muneer, A., Zhang, K., Hamdi, I. et al. Foundation models in biomedical imaging: turning hype into reality.
Nat. Biomed. Eng10, 1557–1575 (2026). https://doi.org/10.1038/s41551-026-01762-z
Received:16 January 2026
Accepted:03 July 2026
Published:11 August 2026
Version of record:11 August 2026
Issue date:August 2026
DOI
:https://doi.org/10.1038/s41551-026-01762-z


