Skip to main navigation Skip to search Skip to main content

Evaluating Artificial Intelligence Chatbots for Patient Education on Thyroid Radiofrequency Ablation: An Analysis of Accuracy, Quality, and Readability

  • Jacob Beiriger
  • , Leonard E. Estephan
  • , Daniel Campbell
  • , Vaninder Kaur Dhillon
  • , David Goldenberg
  • , Julia Noel
  • , Lisa A. Orloff
  • , Marika Russell
  • , Jonathon O. Russell
  • , Catherine F. Sinclair
  • , Elizabeth E. Cottrill

Research output: Contribution to journalArticlepeer-review

Abstract

BACKGROUND: Artificial intelligence (AI) chatbots are increasingly being used by patients to obtain medical information. Comparison between platforms with specialty-specific physician assessment remains limited. This study compares the quality, factual accuracy, readability, and consistency of responses generated by four publicly available AI chatbots when answering patient-centered questions about thyroid radiofrequency ablation (RFA). METHODS: We conducted a cross-sectional analysis of chatbot-generated responses using 20 standardized clinical questions about thyroid RFA. Responses from ChatGPT-4, Gemini, Copilot, and Perplexity were evaluated by six blinded physician reviewers experienced in thyroid RFA using 5-point Likert scales for global quality and factual accuracy. Higher Likert scale scores indicated better performance. Readability and response length were analyzed with established metrics. Statistical significance was defined as p < 0.05. RESULTS: Gemini achieved the highest mean scores for global quality (4.08 ± 0.87) and accuracy (3.76 ± 1.05), with significantly better performance than ChatGPT and Copilot (p < 0.005). ChatGPT responses were significantly longer and more readable. Score variability across questions was lowest for Gemini. Copilot and Perplexity ranked lowest across most domains. Question-level analysis identified specific prompts that best discriminated between platforms. CONCLUSIONS: AI chatbot performance varied across platforms for thyroid RFA queries. Chatbots were generally reliable for straightforward factual information but were less dependable for judgment or context-dependent assessments. These AI tools should supplement, not replace, clinician-vetted patient education and institutional materials.

Original languageEnglish
Pages (from-to)162-168
Number of pages7
JournalThyroid
Volume36
Issue number2
DOIs
StatePublished - 1 Feb 2026
Externally publishedYes

Keywords

  • artificial intelligence
  • chatbot
  • large language model
  • patient education
  • radiofrequency ablation
  • thyroid nodule

Fingerprint

Dive into the research topics of 'Evaluating Artificial Intelligence Chatbots for Patient Education on Thyroid Radiofrequency Ablation: An Analysis of Accuracy, Quality, and Readability'. Together they form a unique fingerprint.

Cite this