AOR

The Archives of Ophthalmological Research aims to publish issues related to publish articles of the highest scientific and clinical value at an international level, and accepts articles on these topics. The target audience of the journal included specialists and physicians working in ophthalmology, and other health professionals interested in these fields.

EndNote Style
Index
Original Article
Multimodal retinal imaging–based assessment of large language models for the differential diagnosis of retinal vascular occlusion
Aims: To compare the diagnostic and treatment recommendation performance of three large language models (LLMs)-ChatGPT-4o, Gemini, and Microsoft Copilot-in the differential diagnosis of retinal vascular occlusions using multimodal retinal imaging.
Methods: This observational study included 75 patients, comprising 50 retinal vein occlusion (RVO; 25 branch and 25 central) and 25 retinal artery occlusion (RAO; 6 branch and 19 central) cases. Each LLM independently evaluated three imaging datasets: (1) color fundus photographs, (2) fundus photographs combined with optical coherence tomography (OCT), and (3) fundus photographs combined with OCT and fundus fluorescein angiography (FFA). A standardized prompt was used for all analyses. Diagnostic accuracy and treatment recommendation accuracy were compared among the models using Cochran’s Q and post-hoc McNemar tests.
Results: Significant differences in diagnostic accuracy were observed among the three models under all imaging conditions (all p<0.001). Using fundus photographs alone, diagnostic accuracy was 21.3% for ChatGPT-4o, 12.0% for Gemini, and 1.3% for Copilot. After adding OCT images, ChatGPT-4o and Gemini each achieved an accuracy of 21.3%, whereas Copilot remained at 1.3%. With multimodal imaging (fundus photographs+OCT+FFA), diagnostic accuracy increased to 28.0% for ChatGPT-4o and 37.3% for Gemini, while Copilot again remained unchanged at 1.3%. Pairwise comparisons demonstrated no significant difference between ChatGPT-4o and Gemini across imaging conditions, whereas both significantly outperformed Copilot. For treatment recommendations, Gemini achieved the highest accuracy (57.3%), followed by ChatGPT-4o (26.7%) and Copilot (10.7%) (p<0.001). Gemini performed significantly better than both ChatGPT-4o and Copilot, and ChatGPT-4o also outperformed Copilot.
Conclusion: Although multimodal retinal imaging improved the diagnostic performance of current LLMs, their overall diagnostic accuracy remained inadequate for independent clinical use in retinal vascular occlusions. Gemini demonstrated the greatest benefit from multimodal image integration and provided the most accurate treatment recommendations. These findings suggest that current LLMs should be regarded as supportive clinical tools rather than autonomous diagnostic systems, while future multimodal Artificial Intelligence models trained on ophthalmology-specific datasets may achieve substantially improved performance.


1. Saeedi P, Petersohn I, Salpea P, et al. Global and regional diabetes prevalence estimates for 2019 and projections for 2030 and 2045: results from the International Diabetes Federation Diabetes Atlas, 9th edition. Diabetes Res Clin Pract. 2019;157:107843. doi:10.1016/j.diabres.2019.107843
2. Gulshan V, Peng L, Coram M, et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA. 2016;316(22):2402-2410. doi:10.1001/jama.2016.17216
3. Huang XM, Yang BF, Zheng WL, et al. Cost-effectiveness of Artificial Intelligence screening for diabetic retinopathy in rural China. BMC Health Serv Res. 2022;22(1):260. doi:10.1186/s12913-022-07655-6
4. Peng Y, Dharssi S, Chen Q, et al. DeepSeeNet: A deep learning model for automated classification of patient-based age-related macular degeneration severity from color fundus photographs. Ophthalmology. 2019;126(4):565-575. doi:10.1016/j.ophtha.2018.11.015
5. Li Z, Wang L, Wu X, et al. Artificial Intelligence in ophthalmology: the path to the real-world clinic. Cell Rep Med. 2023;4(7):101095. doi:10.1016/j.xcrm.2023.101095
6. Ting DSW, Cheung CY, Lim G, et al. Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes. JAMA. 2017;318(22):2211-2223. doi:10.1001/jama.2017.18152
7. Wong DCS, Kiew G, Jeon S, Ting D. Singapore eye lesions analyzer (SELENA): the deep learning system for retinal diseases. In: Grzybowski A, editor. Artificial Intelligence in Ophthalmology. Cham: Springer; 2021:177-185. doi:10.1007/978-3-030-78601-4_13
8. Xie Y, Nguyen QD, Hamzah H, et al. Artificial Intelligence for teleophthalmology-based diabetic retinopathy screening in a national programme: an economic analysis modelling study. Lancet Digit Health. 2020;2(5):240-249. doi:10.1016/S2589-7500(20)30060-1
9. Dismuke C. Progress in examining cost-effectiveness of AI in diabetic retinopathy screening. Lancet Digit Health. 2020;2(5):e212-e213. doi:10.1016/S2589-7500(20)30077-7
10. Leavitt JA, Larson TA, Hodge DO, Gullerud RE. The incidence of central retinal artery occlusion in Olmsted County, Minnesota. Am J Ophthalmol. 2011;152(5):820-823. doi:10.1016/j.ajo.2011.05.005
11. Song P, Xu Y, Zha M, Zhang Y, Rudan I. Global epidemiology of retinal vein occlusion: a systematic review and meta-analysis of prevalence, incidence, and risk factors. J Glob Health. 2019;9(1):010427. doi:10.7189/jogh.09.010427
12. Ponto KA, Elbaz H, Peto T, et al. Prevalence and risk factors of retinal vein occlusion: the Gutenberg health study. J Thromb Haemost. 2015;13(7):1254-1263. doi:10.1111/jth.12982
13. Frederiksen KH, Stokholm L, Frederiksen PH, et al. Cardiovascular morbidity and all-cause mortality in patients with retinal vein occlusion: a Danish nationwide cohort study. Br J Ophthalmol. 2023;107(9):1324-1330. doi:10.1136/bjophthalmol-2022-321225
Volume 3, Issue 3, 2026
Page : 52-56
_Footer