Automating evaluation of LLM-generated responses to patient questions about rare diseases.
Patients with rare diseases often struggle to find accurate medical information, and large language model (LLM)-based chatbots may help meet this need. However, evaluating LLM-generated free-text answers typically requires physician review, which is time-consuming and difficult to scale. This study compared traditional natural language processing (NLP) metrics to emerging LLM-based evaluation approaches for assessing answer quality in the context of Complex Lymphatic Anomalies (CLAs).
Author(s): Zhao, Min, Oh, Inez Y, Gupta, Aditi, Cohen-Cutler, Sally, Harmoney, Kathryn M, Lai, Albert M, Sisk, Bryan A
DOI: 10.1093/jamiaopen/ooag054