Skip to content
Tech News
← Back to articles

Safety and security of large language models in healthcare

read original more articles
Why This Matters

The integration of large language models (LLMs) in healthcare is transforming medical diagnostics, record management, and decision support, offering potential improvements in accuracy and efficiency. However, concerns around safety, security, and robustness highlight the need for rigorous testing and validation before widespread adoption. Ensuring these models are reliable and secure is crucial for protecting patient safety and maintaining trust in AI-driven healthcare solutions.

Key Takeaways

Gommers, J. et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. Lancet 407, 505–514 (2026).

Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 643, 466–473 (2024).

Singh, R., Bapna, M., Diab, A. R., Ruiz, E. S. & Lotter, W. How AI is used in FDA-authorized medical devices: a taxonomy across 1,016 authorizations. NPJ Digit. Med. 8, 388 (2025).

Griot, M., Vanderdonckt, J. & Yuksel, D. Implementation of large language models in electronic health records. PLOS Digital Health https://doi.org/10.1371/journal.pdig.0001141 (2025).

Armitage, H. Clinicians can ‘chat’ with medical records through new AI software, ChatEHR. Stanford Medicine https://med.stanford.edu/news/all-news/2025/06/chatehr.html (2025).

Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023).

Clusmann, J. et al. The future landscape of large language models in medicine. Commun. Med. 3, 141 (2023).

Bubeck, S. et al. Sparks of Artificial General Intelligence: Early experiments with GPT-4. Preprint at https://doi.org/10.48550/arXiv.2303.12712 (2023).

Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med. 31, 943–950 (2025).

Gu, Y. et al. The illusion of readiness: stress testing large frontier models on multimodal medical benchmarks. Preprint at https://doi.org/10.48550/arXiv.2509.18234 (2025). This study reveals that medical LLM benchmarks may overestimate real-world readiness as benchmarks fail to capture brittleness and reasoning flaws.

... continue reading