Skip to main content

Webinar Library

From Clinical Notes and Patient Narratives to Real-World Evidence: Case Studies and Evaluation Approaches for LLM-Based Clinical NLP Pipelines

Large language models are transforming clinical NLP, but real-world healthcare pipelines still combine LLMs with traditional methods to turn clinical notes and patient narratives into structured evidence. This talk covers two active projects using this approach — one studying medication adherence across five diseases, another improving drug safety prediction — and discusses best practices for evaluating LLM systems in clinical settings.
Default image for courses.

Data-Driven Scientific Hypothesis Generation in Clinical Research: The Roles of Human and AI

Hypothesis generation is an early and critical step in any hypothesis-driven clinical research project. However, the process of hypothesis generation is less understood. In this talk, Dr. Jing will briefly introduce the VIADS (a Visual Interactive Analytic tool for filtering and summarizing large health Data Sets coded with hierarchical terminologies) project and its results* in the last decade. The focus will be on a study that recruited clinical researchers who used (or did not use) VIADS for data-driven hypothesis generation in the clinical research context. The talk will summarize the experiments and principal findings and share the lessons learned. The role of humans and/or AI in the process and challenges in measuring the processes and results will also be discussed. * Selected publications— PMID: 39819516, 40768504, 40417518, 39211055, 37011112, 38384898, 35736798, 30764811, 24727931, 22195119 Presenter
Default image for courses.

Graph-based Prediction of Spatio-Temporal Vaccine Hesitancy from Insurance Claims Data

The VaxHesSTL framework combines Graph and Recurrent Neural Networks to predict vaccine hesitancy at the ZIP code level by capturing both spatial relationships and historical trends, outperforming existing models when trained on a six-year, five-million-person insurance claims dataset from Virginia. To address the high cost of such data, the research also explores an active learning approach to optimize which ZIP codes are selected for training.
Default image for courses.