Gommers, J. et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. Lancet 407, 505–514 (2026).
Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 643, 466–473 (2024).
Singh, R., Bapna, M., Diab, A. R., Ruiz, E. S. & Lotter, W. How AI is used in FDA-authorized medical devices: a taxonomy across 1,016 authorizations. NPJ Digit. Med. 8, 388 (2025).
Griot, M., Vanderdonckt, J. & Yuksel, D. Implementation of large language models in electronic health records. PLOS Digital Health https://doi.org/10.1371/journal.pdig.0001141 (2025).
Armitage, H. Clinicians can ‘chat’ with medical records through new AI software, ChatEHR. Stanford Medicine https://med.stanford.edu/news/all-news/2025/06/chatehr.html (2025).
Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023).
Clusmann, J. et al. The future landscape of large language models in medicine. Commun. Med. 3, 141 (2023).
Bubeck, S. et al. Sparks of Artificial General Intelligence: Early experiments with GPT-4. Preprint at https://doi.org/10.48550/arXiv.2303.12712 (2023).
Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med. 31, 943–950 (2025).
Gu, Y. et al. The illusion of readiness: stress testing large frontier models on multimodal medical benchmarks. Preprint at https://doi.org/10.48550/arXiv.2509.18234 (2025). This study reveals that medical LLM benchmarks may overestimate real-world readiness as benchmarks fail to capture brittleness and reasoning flaws.
... continue reading