Document resource
Background/aims Ocular surface infections remain a major cause of visual loss worldwide, yet diagnosis often relies on slow or insensitive microbiological techniques. Artificial intelligence may complement emerging molecular tools by supporting rapid triage and diagnostic reasoning. This study benchmarked publicly available multimodal large language models (LLMs) against corneal specialists for the diagnosis, treatment and urgency triage of infectious keratitis and conjunctivitis.Methods A single-centre diagnostic-accuracy study included 60 microbiologically confirmed infectious keratitis and conjunctivitis cases, each comprising a slit-lamp photograph and a paired clinical vignette. Six multimodal LLMs (GPT-4o, GPT-5, Gemini, Claude, Perplexity and Grok) were evaluated for diagnosis, treatment and urgency triage under three input conditions (image-only, text-only and image+text). Outputs were compared with two corneal specialists.Results LLM performance depended strongly on input modality. Image-only accuracy was lowest (best GPT-5, 61.4%; κ=0.38) with frequent misclassification of fungal and Acanthamoeba keratitis and hallucinations confined to this setting. Text input improved results (GPT-5, 83.3%; κ=0.78), though accuracy remained below specialists (87–90%; κ≈0.8). Combined image+text achieved near-human accuracy without consistently surpassing corneal specialists (Perplexity 96.7%; κ=0.95; GPT-5 91.7%; κ=0.87). Treatment accuracy remained lower (81–85% vs 90–98%), while urgency triage matched experts in multimodal input.Conclusion Publicly accessible multimodal LLMs can approach expert-level performance in diagnosis and triage when provided with clinical context and slit-lamp images. Gaps in therapeutic reasoning and rare pathogen recognition underscore the need for targeted refinement and validation. These models may complement specialist care, supporting rapid triage and integration with molecular or metagenomic diagnostics, especially in resource-limited settings.