Generative artificial intelligence for clinical reasoning assessment in medical students: a systematic review of early experimental evidence

Inteligencia artificial generativa para la evaluación del razonamiento clínico en estudiantes de medicina: una revisión sistemática de la evidencia experimental temprana

Introduction

Clinical reasoning is a core competency in undergraduate medical education and is typically developed through repeated exposure to clinical scenarios with structured feedback from experienced clinicians. However, access to individualized feedback is often limited by structural constraints within medical education systems. Artificial intelligence (AI)–based educational tools have emerged as a potential strategy to support the development and assessment of clinical reasoning by providing scalable automated feedback and simulation-based learning environments. This study evaluated the effectiveness of AI-based educational interventions in enhancing clinical reasoning skills among undergraduate medical students.

Material and methods

A systematic search of PubMed, Scopus, ERIC, and IEEE Xplore was conducted for studies published between January 2010 and January 2025. Eligible studies included undergraduate medical students exposed to AI-based tools designed to assess or support clinical reasoning. Methodological quality was assessed using the Mixed Methods Appraisal Tool (MMAT) 2018. When sufficient data were available, standardized mean differences (Hedges’ g) were calculated. Due to methodological heterogeneity, findings were synthesized narratively following the SWiM framework.

Results

Four studies met the inclusion criteria, including 423 undergraduate medical students and five clinical experts. AI interventions included automated formative feedback systems, AI-simulated patients, and AI-generated assessment tools. Across studies, AI-mediated feedback produced learning outcomes comparable to expert-generated feedback, and one quasi-experimental study suggested accelerated reasoning development among novice learners.

Conclusion

Current evidence suggests that AI-based educational interventions may support clinical reasoning development in undergraduate medical education. However, the evidence base remains limited, highlighting the need for further high-quality and longitudinal studies.

Enlazar con artículo