Abstract
Objective: To evaluate the extent to which clinical artificial intelligence (AI) and machine learning (ML) models prioritize updating, transparency, and demographic reporting in the published literature.
Patients and Methods: This study conducted a systematic review of clinical AI/ML models using PRISMA guidelines from March 2020 until December 2021. A new checklist and scoring system were introduced to assess model quality, with additional evaluation of demographic reporting, particularly by ethnicity and race. A comprehensive search was performed across six major databases, including Ovid Embase, MEDLINE, and Cochrane Library. Across various study designs, eligible studies included human-based predictive or prognostic AI/ML models using supervised learning and at least two predictors. Studies not meeting these criteria were excluded.
Results: Out of 390 AI/ML studies reviewed, only 9% mentioned plans or methods for future model updates. The vast majority (98%) of models were still in the research phase, and only 2% had reached production. Additionally, only 12% adhered to best practices in model development, and 84% failed to report demographic composition by race or ethnicity.
Conclusion: These findings highlight key limitations in the current clinical AI landscape—especially a lack of transparency, limited readiness for deployment, and minimal consideration for inclusivity or generalizability. Greater focus on model updating, adherence to development standards, and demographic transparency is essential to improve the safety, reliability, and equity of clinical AI/ML models.