Is one run enough? Reproducibility of flagship large language models across temperature and reasoning settings in biomedical text processing.
To quantify run-to-run reproducibility of Gemini 3 Flash Preview and GPT-5.2 for trial-success classification across temperature and reasoning/thinking settings and determine whether single-run reporting suffices.
Author(s): Windisch, Paul, Koechli, Carole, Dennstädt, Fabio, Aebersold, Daniel M, Zwahlen, Daniel R, Förster, Robert, Schröder, Christina
DOI: 10.1093/jamia/ocag039