
The comparative study evaluated the performance of 15 large language models (LLMs) in analyzing 20 difficult ophthalmic cases. The cases were presented in open-ended text format, and LLMs provided responses on differential diagnoses, diagnostic tests, and recommended treatments. The main outcome, accuracy, showed a mean score of 19 ± 9 out of 60. Three models—ChatGPT 3.5, Claude Pro, and Copilot Pro—scored significantly higher, with Copilot Pro outperforming the others. While readability exceeded the American Medical Association’s recommended level for public understanding, the content was suitable for ophthalmologists. Despite promising performance as clinical assistants, LLMs' accuracy is insufficient for independent patient care decisions, though they could aid in reducing oversights.
Like
Save
Share