Learning statistical models of phenotypes using noisy labeled training data.
Traditionally, patient groups with a phenotype are selected through rule-based definitions whose creation and validation are time-consuming. Machine learning approaches to electronic phenotyping are limited by the paucity of labeled training datasets. We demonstrate the feasibility of utilizing semi-automatically labeled training sets to create phenotype models via machine learning, using a comprehensive representation of the patient medical record.
Author(s): Agarwal, Vibhu, Podchiyska, Tanya, Banda, Juan M, Goel, Veena, Leung, Tiffany I, Minty, Evan P, Sweeney, Timothy E, Gyang, Elsie, Shah, Nigam H
DOI: 10.1093/jamia/ocw028