
Bing Chat and Claude 3.5 Sonnet outperformed others in accuracy (5.78 ± 0.48, 5.75 ± 0.54) and competence (2.65 ± 0.58, 2.80 ± 0.41). Updated models (Gemini 2.0 Flash Experimental, ChatGPT-4o mini) performed comparably. While LLMs provide reliable CNLD/HOT information, domain-specific variability and misinformation risks highlight the need for expert oversight before clinical use.
Like
Save
Share