Dr. Shauly
Research
Devices
Digital Curbside
Innovation
Podcast
About
Contact

Dr. Orr Shauly

Plastic & Reconstructive Surgery

ResearchDevicesDigital CurbsideInnovationPodcastAboutContact

© 2026 Dr. Orr Shauly

LinkedInInstagramX / Twitter
← All episodes

Season 1, Episode 13

May 11, 2026 · Innovation

MEDICAL CHATBOTS

This landmark article from Dr. Shauly evaluates the efficacy of AI chatbots integrated into top-ranking plastic surgery websites, specifically focusing on their ability to perform medical triage. While the study found these tools highly successful at managing administrative tasks and elective inquiries, they demonstrated a dangerous inability to recognize emergent medical complications. Specifically, the models failed to identify four out of five life-threatening scenarios, often defaulting to generic scripts instead of necessary physician escalation. The data suggest that although chatbots provide significant cost savings and logistical efficiency, they currently lack the specialized training required for safe clinical decision-making. Consequently, the authors conclude that these tools should remain limited to scheduling and education until more robust, specialty-specific models are developed.

Nip Talk: Medical Chatbots

Your browser does not support the audio element.

Featured article

Nip Talk: Medical Chatbots

Written by Orr Shauly

Comprehensive Study Guide

Short Answer Questions

Please answer the following questions in 2-3 sentences each.

  1. What was the primary objective of this research study?

    AnswerThe study aimed to evaluate the triage classification accuracy, escalation patterns, and quality of patient interactions for AI chatbots on plastic surgery websites. It sought to determine if these tools effectively fulfill their administrative potential while maintaining clinical safety.

  2. Describe the methodology used to select the websites and chatbots analyzed in this study.

    AnswerResearchers identified the top twenty plastic surgery websites using search engine optimization (SEO) rankings through an incognito Google search with location services disabled. Only websites containing embedded, publicly accessible, and functional interactive chatbots were included in the analysis.

  3. How did the study define and categorize the "clinical scenarios" presented to the chatbots?

    AnswerThe chatbots were tested using 60 standardized clinical interactions divided into three urgency levels: emergent, urgent, and elective (20 scenarios each). These scenarios were pre-approved by a physician and submitted verbatim to each chatbot using a standardized fictitious patient profile.

  4. What does the 20% sensitivity rate for emergent scenarios indicate about chatbot performance?

    AnswerA 20% sensitivity rate indicates that the chatbots failed to correctly identify 80% of emergent cases (4 out of 5 cases missed). This high false-negative rate suggests that chatbots are currently ill-equipped to handle high-acuity patient needs and may pose safety risks.

  5. What is "escalation" in the context of this study, and how frequently did it occur in emergent cases?

    AnswerEscalation was defined as any instance where the chatbot referred the user to a live provider or forwarded their information to one. In emergent cases, escalation occurred 80% of the time, though often these escalations happened after a delay or following misclassification.

  6. How did classification accuracy affect the user experience, according to CUQ scores?

    AnswerMisclassified interactions were associated with significantly lower usability scores, with a mean Chatbot Usability Questionnaire (CUQ) score of 49.1 compared to 60.8 for correctly classified interactions. This suggests that technical accuracy is a primary driver of patient satisfaction and perceived utility.

  7. What were the findings regarding the use of medical disclaimers on the analyzed websites?

    AnswerThe study found that only one of the twenty websites included a disclaimer stating that the chatbot cannot replace in-person medical care. Most other "disclaimers" were standard legal or privacy notices that did not clarify the chatbot’s limited clinical capacity.

  8. Compare the error rates of rule-based (button/flow) chatbots versus AI/NLP-based systems.

    AnswerRule-based (button/flow) chatbots demonstrated the lowest error rate at 41%. In contrast, more complex hybrid models and AI/NLP-based systems had significantly higher error rates, both exceeding 65%.

  9. What are "Small Language Models" (SLMs), and why are they suggested as a future direction?

    AnswerSLMs are compact AI tools tailored to specific clinical domains, such as aesthetic surgical concerns. They are suggested as a way to integrate specialty-specific knowledge and aesthetic judgment heuristics that current general-purpose models lack.

  10. Explain the meaning of the Cohen’s kappa score of 0.47 reported in the results.

    AnswerA Cohen’s kappa score of 0.47 indicates "moderate agreement" between the chatbot’s triage classification and the true physician-determined classification. This confirms that while there is some alignment, there is still a significant discrepancy in how AI and human experts evaluate clinical urgency.

Key Terms

Chatbot Usability Questionnaire (CUQ)
A metric utilized to quantify and assess the user experience and satisfaction of patients interacting with chatbots.
Emergent Classification
A high-acuity triage category for critical medical situations, which chatbots poorly identified with a sensitivity of only 20% and an 80% false negative rate.
Natural Language Processing (NLP)
A machine learning technology used by chatbots to nurture cosmetic surgery leads and handle inquiries, though NLP-based systems demonstrated high error rates exceeding 65% in the studied clinical scenarios.
Rule-Based Chatbots
Button or flow-based automated systems that demonstrated the lowest error rate (41%) compared to hybrid or AI-based models.
Small Language Models (SLMs)
Compact artificial intelligence tools tailored to specific clinical domains that are recommended as a future direction to improve the emergency triage capabilities of medical chatbots.
Triage Classification
The process of categorizing clinical inquiries as emergent, urgent, or elective based on severity to determine the appropriate scheduling, advice, or escalation response.
Visual Analog Scale (VAS)
A 10-point scale used to evaluate patient satisfaction and user experience, with higher scores correlating to chatbots that had lower error rates and more appropriate triage behaviors

Previous

NECKLIFT GLAND RESECTION