Document resource
Background Compare the performance of three artificial intelligence (AI) chatbots, including ChatGPT o4-mini, Gemini 2.5 flash and DeepSeek-R1, in diagnosis, treatment suggestion and visual prognosis of age-related macular degeneration (AMD).Methods A retrospective study was conducted on 54 participants with AMD from Shenzhen Eye Hospital and Huizhou Third People’s Hospital (from January 2024 to June 2025). Each participant was analysed by three AI chatbots and an ophthalmology resident using standardised prompts. The main outcome measures included diagnosis agreement, treatment suggestion and visual prognosis. Global Quality Score (GQS, 1–5) was assessed to evaluate response quality. Statistical analysis was conducted using R (version 4.4.1), with significance set at p<0.05.Results Diagnosis agreement was 96.3% for ChatGPT o4-mini, 94.4% for Gemini 2.5 flash, 90.7% for DeepSeek-R1 and resident (all p>0.05). Treatment suggestion agreement was 90.7% for Gemini 2.5 flash, 88.9% for ChatGPT o4-mini, DeepSeek-R1 and resident (all p>0.05). Visual prognosis agreement was 66.7% for ChatGPT o4-mini, 48.1% for Gemini 2.5 flash, 40.7% for DeepSeek-R1 and 37.0% for resident. ChatGPT o4-mini significantly outperformed DeepSeek-R1 (p=0.011) and the resident (p=0.006). GQS ratings favoured ChatGPT o4-mini, significantly higher than Gemini 2.5 flash and DeepSeek-R1, but comparable to the resident (all p>0.05).Conclusions Three AI chatbots showed strong capabilities in neovascular AMD diagnosis and treatment suggestions, with ChatGPT o4-mini superior in visual prognosis prediction. However, real-world clinical settings require large-scale studies, and integrative use of multiple large language models may improve robustness and clinical reliability.