Searched for: in-biosketch:true
person:moyl02
Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM): 2025 Updates
Park, Seong Ho; Suh, Chong Hyun; Lee, Jeong Hyun; Tejani, Ali S; You, Seng Chan; Kahn, Charles E; Moy, Linda
Recent systematic reviews have raised concerns about the quality of reporting in studies evaluating the accuracy of large language models (LLMs) in medical applications. Incomplete and inconsistent reporting hampers the ability of reviewers and readers to assess study methodology, interpret results, and evaluate reproducibility. To address this issue, the MInimum reporting items for CLear Evaluation of Accuracy Reports of Large Language Models in healthcare (MI-CLEAR-LLM) checklist was developed. This article presents an extensively updated version. While the original version focused on proprietary LLMs accessed via web-based chatbot interfaces, the updated checklist incorporates considerations relevant to application programming interfaces and self-managed models, typically based on open-source LLMs. As before, the revised MI-CLEAR-LLM focuses on reporting practices specific to LLM accuracy evaluations: specifically, the reporting of how LLMs are specified, accessed, adapted, and applied in testing, with special attention to methodological factors that influence outputs. The checklist includes essential items across categories such as model identification, access mode, input data type, adaptation strategy, prompt optimization, prompt execution, stochasticity management, and test data independence. This article also presents reporting examples from the literature. Adoption of the updated MI-CLEAR-LLM can help ensure transparency in reporting and enable more accurate and meaningful evaluation of studies.
PMID: 41199132
ISSN: 2005-8330
CID: 5960222
Best Practices for the Safe Use of Large Language Models and Other Generative AI in Radiology
Yi, Paul H; Haver, Hana L; Jeudy, Jean J; Kim, Woojin; Kitamura, Felipe C; Oluyemi, Eniola T; Smith, Andrew D; Moy, Linda; Parekh, Vishwa S
As large language models (LLMs) and other generative artificial intelligence (AI) models are rapidly integrated into radiology workflows, unique pitfalls threatening their safe use have emerged. Problems with AI are often identified only after public release, highlighting the need for preventive measures to mitigate negative impacts and ensure safe, effective deployment into clinical settings. This article summarizes best practices for the safe use of LLMs and other generative AI models in radiology, focusing on three key areas that can lead to pitfalls if overlooked: regulatory issues, data privacy, and bias. To address these areas and minimize risk to patients, radiologists must examine all potential failure modes and ensure vendor transparency. These best practices are based on the best available evidence and the experiences of leaders in the field. Ultimately, this article provides actionable guidelines for radiologists, radiology departments, and vendors using and integrating generative AI into radiology workflows, offering a framework to prevent these problems.
PMID: 40985835
ISSN: 1527-1315
CID: 5937652
Evaluating Breast Cancer Intravoxel Incoherent Motion MRI Biomarkers across Software Platforms
Sigmund, Eric E; Cho, Gene Y; Basukala, Dibash; Sutton, Olivia M; Horvat, Joao V; Mikheev, Artem; Rusinek, Henry; Gilani, Nima; Li, Xiaochun; Babb, James S; Goldberg, Judith D; Pinker, Katja; Moy, Linda; Thakur, Sunitha B
Purpose To evaluate intravoxel incoherent motion (IVIM) biomarkers across different MRI vendors and software programs for breast cancer characterization in a two-site study. Materials and Methods This institutional review board-approved, Health Insurance Portability and Accountability Act-compliant retrospective study included 106 patients (with 18 benign and 88 malignant lesions) who underwent bilateral diffusion-weighted imaging (DWI) between February 2009 and March 2013. DWI was performed using 1.5-T (n = 6) or 3-T MRI scanners from two vendors using single-shot spin-echo echo-planar imaging or twice-refocused, bipolar gradient single-shot turbo spin-echo readout with multiple b values between 0 and 1000 sec/mm2. IVIM parameters tissue diffusivity (Dt
PMID: 40910883
ISSN: 2638-616x
CID: 5936402
Editorial Opportunities for Radiology Trainees: RSNA's Radiology: In Training Program [Editorial]
Guarnera, Alessia; Yilmaz, Enis C; Marrocchio, Cristina; Prodigios, Joice; Moy, Linda; Chernyak, Victoria
PMID: 40828046
ISSN: 1527-1315
CID: 5908902
Performance of Algorithms Submitted in the 2023 RSNA Screening Mammography Breast Cancer Detection AI Challenge
Chen, Yan; Partridge, George J W; Vazirabad, Maryam; Ball, Robyn L; Trivedi, Hari M; Kitamura, Felipe Campos; Frazer, Helen M L; Retson, Tara A; Yao, Luyan; Darker, Iain T; Kelil, Tatiana; Mongan, John; Mann, Ritse M; Moy, Linda
Background The 2023 RSNA Screening Mammography Breast Cancer Detection AI Challenge invited participants to develop artificial intelligence (AI) models capable of independently interpreting mammograms. Purpose To assess the performance of the submitted algorithms, explore the potential for improving performance by combining the best-performing AI algorithms, and investigate how performance was influenced by the demographic and clinical characteristics of the evaluation cohort. Materials and Methods A total of 1687 AI algorithms were submitted from November 2022 to February 2023. Of these, 1537 algorithms were assessed using an evaluation dataset from two sites-one in the United States and one in Australia. Cancer cases were identified at screening and confirmed with pathologic examination; noncancer cases were followed up for at least 1 year. Results for ensemble models of top algorithms were computed by recalling a case when any of the included algorithms indicated recall. Odds ratios (ORs) were used to investigate differences in AI performance when the dataset was stratified by clinical or demographic characteristics. Results The evaluation dataset consisted of 5415 women (median age, 59 years [IQR, 52-66 years]). Among the 1537 AI algorithms, the median recall rate, sensitivity, specificity, and positive predictive value (PPV) were 1.7%, 27.6%, 98.7%, and 36.9%, respectively. For the top-ranked algorithm, the recall rate, sensitivity, specificity, and PPV were 1.5%, 48.6%, 99.5%, and 64.6%, respectively. Ensemble models of the top 3 and top 10 algorithms had a sensitivity of 60.7% and 67.8%, respectively; the corresponding recall rates were 2.4% and 3.5%, and the corresponding specificities were 98.8% and 97.8%. Lower sensitivity was observed for the U.S. dataset than for the Australian dataset (top 3 ensemble model: 52.0% vs 68.1%; OR = 0.51; P = .02), and greater sensitivity was observed for invasive cancers than for noninvasive cancers (top 3 ensemble model: 68.0% vs 43.8%; OR = 2.73; P = .001). Conclusion The different AI algorithms identified different cancers during screening mammography, and ensemble models had increased sensitivity while maintaining low recall rates. © RSNA, 2025 Supplemental material is available for this article.
PMID: 40793948
ISSN: 1527-1315
CID: 5907052
Pitfalls and Best Practices in Evaluation of AI Algorithmic Biases in Radiology
Yi, Paul H; Bachina, Preetham; Bharti, Beepul; Garin, Sean P; Kanhere, Adway; Kulkarni, Pranav; Li, David; Parekh, Vishwa S; Santomartino, Samantha M; Moy, Linda; Sulam, Jeremias
Despite growing awareness of problems with fairness in artificial intelligence (AI) models in radiology, evaluation of algorithmic biases, or AI biases, remains challenging due to various complexities. These include incomplete reporting of demographic information in medical imaging datasets, variability in definitions of demographic categories, and inconsistent statistical definitions of bias. To guide the appropriate evaluation of AI biases in radiology, this article summarizes the pitfalls in the evaluation and measurement of algorithmic biases. These pitfalls span the spectrum from the technical (eg, how different statistical definitions of bias impact conclusions about whether an AI model is biased) to those associated with social context (eg, how different conventions of race and ethnicity impact identification or masking of biases). Actionable best practices and future directions to avoid these pitfalls are summarized across three key areas: (a) medical imaging datasets, (b) demographic definitions, and (c) statistical evaluations of bias. Although AI bias in radiology has been broadly reviewed in the recent literature, this article focuses specifically on underrecognized potential pitfalls related to the three key areas. By providing awareness of these pitfalls along with actionable practices to avoid them, exciting AI technologies can be used in radiology for the good of all people.
PMID: 40392092
ISSN: 1527-1315
CID: 5852522
AI-generated Podcast Summaries of Radiology Articles: Analysis of Content and Quality
Tejani, Ali S; Khosravi, Bardia; Savage, Cody H; Moy, Linda; Kahn, Charles E; Yi, Paul H
PMCID:11950872
PMID: 40035674
ISSN: 1527-1315
CID: 5842722
Breast Arterial Calcifications on Mammography: A Review of the Literature
Rossi, Joanna; Cho, Leslie; Newell, Mary S; Venta, Luz A; Montgomery, Guy H; Destounis, Stamatia V; Moy, Linda; Brem, Rachel F; Parghi, Chirag; Margolies, Laurie R
Identifying systemic disease with medical imaging studies may improve population health outcomes. Although the pathogenesis of peripheral arterial calcification and coronary artery calcification differ, breast arterial calcification (BAC) on mammography is associated with cardiovascular disease (CVD), a leading cause of death in women. While professional society guidelines on the reporting or management of BAC have not yet been established, and assessment and quantification methods are not yet standardized, the value of reporting BAC is being considered internationally as a possible indicator of subclinical CVD. Furthermore, artificial intelligence (AI) models are being developed to identify and quantify BAC on mammography, as well as to predict the risk of CVD. This review outlines studies evaluating the association of BAC and CVD, introduces the role of preventative cardiology in clinical management, discusses reasons to consider reporting BAC, acknowledges current knowledge gaps and barriers to assessing and reporting calcifications, and provides examples of how AI can be utilized to measure BAC and contribute to cardiovascular risk assessment. Ultimately, reporting BAC on mammography might facilitate earlier mitigation of cardiovascular risk factors in asymptomatic women.
PMID: 40163666
ISSN: 2631-6129
CID: 5818782
Retrospective BReast Intravoxel Incoherent Motion Multisite (BRIMM) multisoftware study
Basukala, Dibash; Mikheev, Artem; Li, Xiaochun; Goldberg, Judith D; Gilani, Nima; Moy, Linda; Pinker, Katja; Partridge, Savannah C; Biswas, Debosmita; Kataoka, Masako; Honda, Maya; Iima, Mami; Thakur, Sunitha B; Sigmund, Eric E
INTRODUCTION/UNASSIGNED:The intravoxel incoherent motion (IVIM) model of diffusion weighted imaging (DWI) provides imaging biomarkers for breast tumor characterization. It has been extensively applied for both diagnostic and prognostic goals in breast cancer, with increasing evidence supporting its clinical relevance. However, variable performance exists in literature owing to the heterogeneity in datasets and quantification methods. METHODS/UNASSIGNED: RESULTS/UNASSIGNED: DISCUSSION/UNASSIGNED:
PMCID:11891049
PMID: 40066090
ISSN: 2234-943x
CID: 5808282
FastMRI Breast: A Publicly Available Radial k-Space Dataset of Breast Dynamic Contrast-enhanced MRI
Solomon, Eddy; Johnson, Patricia M; Tan, Zhengguo; Tibrewala, Radhika; Lui, Yvonne W; Knoll, Florian; Moy, Linda; Kim, Sungheon Gene; Heacock, Laura
The fastMRI breast dataset is the first large-scale dataset of radial k-space and Digital Imaging and Communications in Medicine data for breast dynamic contrast-enhanced MRI with case-level labels, and its public availability aims to advance fast and quantitative machine learning research.
PMCID:11791504
PMID: 39772976
ISSN: 2638-6100
CID: 5805022