Searched for: in-biosketch:yes
person:vjm261
Incidence and prognostic value of electrocardiographic changes in patients with myocardial injury after elective PCI
LaRaja, Alexander; Chakraborty, Ashish; Talmor, Nina; Graves, Claire; Kozloff, Sam; Major, Vincent J; Shah, Binita; Babaev, Anvar; Razzouk, Louai; Rao, Sunil V; Attubato, Michael; Feit, Frederick; Slater, James; Smilowitz, Nathaniel R
BACKGROUND:Myocardial injury after percutaneous coronary intervention (PCI) is common and associated with adverse outcomes. Contemporary definitions of periprocedural myocardial infarction require biomarker elevation with or without ischemic electrocardiographic (ECG) findings. However, the incremental prognostic value of ECG beyond biomarker-defined injury alone remains uncertain. OBJECTIVES/OBJECTIVE:To determine the incidence and independent prognostic value of ischemic ECG changes after elective PCI. METHODS:Consecutive adults age ≥ 18 years undergoing elective PCI at NYU Langone Health between 2011 and 2020 were included. Creatine kinase-myocardial band (CKMB) concentrations were measured at 1 and 3 h post-PCI. Among patients with myocardial injury, baseline and post-PCI ECGs (within 24 h) were reviewed to identify development of ischemic ECG changes (ST segment abnormalities, T wave abnormalities, and Q waves). Relationships between ischemic ECG findings and mortality were evaluated in Cox proportional hazards models adjusted for age, sex, and assay-normalized CKMB. RESULTS:Among 10,735 patients, 1741 (16.2%) developed post-PCI myocardial injury. New ischemic ECG changes occurred in 18.4% of patients with myocardial injury and increased stepwise with higher concentrations of CKMB. New T wave abnormalities were most common (11%), followed by ST depressions (4.9%), Q waves (3.0%), and ST elevations (1.5%). Over a median follow-up of 5.3 years, new ischemic ECG changes were not independently associated with increased mortality among patients with myocardial injury (aHR 1.27, 95% CI 0.82-1.96). CONCLUSIONS:Among patients with myocardial injury after elective PCI, new ischemic ECG changes were uncommon and did not confer independent prognostic value for long-term mortality.
PMID: 42399161
ISSN: 1878-0938
CID: 6063822
Implementing the NYU Electronic Patient Visit Assessment (ePVA)© for head and neck cancer in rural and urban populations: a study protocol for a type 1 hybrid effectiveness-implementation clinical trial
Van Cleave, Janet H; Brody, Abraham A; Schulman-Green, Dena; Hu, Kenneth S; Li, Zujun; Johnson, Stephen B; Major, Vincent J; Lominska, Christopher E; Bauman, Jessica R; Hanania, Alexander N; Tatlonghari, Ghia V; Tsikis, Marcely; Egleston, Brian L
BACKGROUND:for HNC as a digital patient-reported symptom monitoring system that enables early symptom detection and real-time interventions at the point of care. With this study protocol, we aim to test the effectiveness of the ePVA in improving HNC outcomes in real-world settings and to identify implementation strategies optimizing its effectiveness. METHODS:We will conduct a longitudinal mixed-methods hybrid type I study at four National Cancer Institute-designated Comprehensive Cancer Centers serving diverse populations in rural and urban settings (New York University, the University of Kansas Cancer Center, Fox Chase Cancer Center, and Baylor College of Medicine) guided by the Reach, Effectiveness, Adoption, Implementation, and Maintenance (RE-AIM) framework. Patient eligibility criteria include having histologically diagnosed HNC and undergoing radiation therapy with or without chemotherapy for curative intent. We will also interview clinicians caring for patients with HNC at the participating institutions regarding facilitators and barriers to implementing the ePVA. The accrual goal is 270 patients. Aim 1 is to determine the effect of the ePVA on HNC symptoms in a two-arm (usual care vs. ePVA + usual care) trial. The study's primary outcomes are patients' self-reported social function, senses of taste and smell, and swallowing, measured by the European Organization for Research and Treatment of Cancer QLQ-C30 and QLQ-H&N35. For Aim 2, we will interview patients (n = 40) as well as clinicians (n = 30) caring for patients with HNC at the participating institutions regarding facilitators and barriers to implementing the ePVA. In Aim 3, we will integrate Aims 1 and 2 data to identify strategies that optimize the use of the ePVA. DISCUSSION/CONCLUSIONS:The overarching goal of this research is to advance cancer care by identifying implementation standards for effective, widespread use of the ePVA that apply to all patient-reported outcomes in cancer care. TRIAL REGISTRATION/BACKGROUND:ClinicalTrials.gov NCT06030011. Registered on 8 September 2023.
PMCID:12690913
PMID: 41366462
ISSN: 1745-6215
CID: 5977322
Evaluating Hospital Course Summarization by an Electronic Health Record-Based Large Language Model
Small, William R.; Austrian, Jonathan; O\Donnell, Luke; Burk-Rafel, Jesse; Hochman, Katherine A.; Goodman, Adam; Zaretsky, Jonah; Martin, Jacob; Johnson, Stephen; Major, Vincent J.; Jones, Simon; Henke, Christian; Verplanke, Benjamin; Osso, Jwan; Larson, Ian; Saxena, Archana; Mednick, Aron; Simonis, Choumika; Han, Joseph; Kesari, Ravi; Wu, Xinyuan; Heery, Lauren; Desel, Tenzin; Baskharoun, Samuel; Figman, Noah; Farooq, Umar; Shah, Kunal; Jahan, Nusrat; Kim, Jeong Min; Testa, Paul; Feldman, Jonah
ISI:001551557000002
ISSN: 2574-3805
CID: 5974192
Evaluating Hospital Course Summarization by an Electronic Health Record-Based Large Language Model
Small, William R; Austrian, Jonathan; O'Donnell, Luke; Burk-Rafel, Jesse; Hochman, Katherine A; Goodman, Adam; Zaretsky, Jonah; Martin, Jacob; Johnson, Stephen; Major, Vincent J; Jones, Simon; Henke, Christian; Verplanke, Benjamin; Osso, Jwan; Larson, Ian; Saxena, Archana; Mednick, Aron; Simonis, Choumika; Han, Joseph; Kesari, Ravi; Wu, Xinyuan; Heery, Lauren; Desel, Tenzin; Baskharoun, Samuel; Figman, Noah; Farooq, Umar; Shah, Kunal; Jahan, Nusrat; Kim, Jeong Min; Testa, Paul; Feldman, Jonah
IMPORTANCE/UNASSIGNED:Hospital course (HC) summarization represents an increasingly onerous discharge summary component for physicians. Literature supports large language models (LLMs) for HC summarization, but whether physicians can effectively partner with electronic health record-embedded LLMs to draft HCs is unknown. OBJECTIVES/UNASSIGNED:To compare the editing effort required by time-constrained resident physicians to improve LLM- vs physician-generated HCs toward a novel 4Cs (complete, concise, cohesive, and confabulation-free) HC. DESIGN, SETTING, AND PARTICIPANTS/UNASSIGNED:Quality improvement study using a convenience sample of 10 internal medicine resident editors, 8 hospitalist evaluators, and randomly selected general medicine admissions in December 2023 lasting 4 to 8 days at New York University Langone Health. EXPOSURES/UNASSIGNED:Residents and hospitalists reviewed randomly assigned patient medical records for 10 minutes. Residents blinded to author type who edited each HC pair (physician and LLM) for quality in 3 minutes, followed by comparative ratings by attending hospitalists. MAIN OUTCOMES AND MEASURES/UNASSIGNED:Editing effort was quantified by analyzing the edits that occurred on the HC pairs after controlling for length (percentage edited) and the degree to which the original HCs' meaning was altered (semantic change). Hospitalists compared edited HC pairs with A/B testing on the 4Cs (5-point Likert scales converted to 10-point bidirectional scales). RESULTS/UNASSIGNED:Among 100 admissions, compared with physician HCs, residents edited a smaller percentage of LLM HCs (LLM mean [SD], 31.5% [16.6%] vs physicians, 44.8% [20.0%]; P < .001). Additionally, LLM HCs required less semantic change (LLM mean [SD], 2.4% [1.6%] vs physicians, 4.9% [3.5%]; P < .001). Attending physicians deemed LLM HCs to be more complete (mean [SD] difference LLM vs physicians on 10-point bidirectional scale, 3.00 [5.28]; P < .001), similarly concise (mean [SD], -1.02 [6.08]; P = .20), and cohesive (mean [SD], 0.70 [6.14]; P = .60), but with more confabulations (mean [SD], -0.98 [3.53]; P = .002). The composite scores were similar (mean [SD] difference LLM vs physician on 40-point bidirectional scale, 1.70 [14.24]; P = .46). CONCLUSIONS AND RELEVANCE/UNASSIGNED:Electronic health record-embedded LLM HCs required less editing than physician-generated HCs to approach a quality standard, resulting in HCs that were comparably or more complete, concise, and cohesive, but contained more confabulations. Despite the potential influence of artificial time constraints, this study supports the feasibility of a physician-LLM partnership for writing HCs and provides a basis for monitoring LLM HCs in clinical practice.
PMID: 40802185
ISSN: 2574-3805
CID: 5906762
Identifying and mitigating algorithmic bias in the safety net
Mackin, Shaina; Major, Vincent J; Chunara, Rumi; Newton-Dame, Remle
Algorithmic bias occurs when predictive model performance varies meaningfully across sociodemographic classes, exacerbating systemic healthcare disparities. NYC Health + Hospitals, an urban safety net system, assessed bias in two binary classification models in our electronic medical record: one predicting acute visits for asthma and one predicting unplanned readmissions. We evaluated differences in subgroup performance across race/ethnicity, sex, language, and insurance using equal opportunity difference (EOD), a metric comparing false negative rates. The most biased classes (race/ethnicity for asthma, insurance for readmission) were targeted for mitigation using threshold adjustment, which adjusts subgroup thresholds to minimize EOD, and reject option classification, which re-classifies scores near the threshold by subgroup. Successful mitigation was defined as 1) absolute subgroup EODs <5 percentage points, 2) accuracy reduction <10%, and 3) alert rate change <20%. Threshold adjustment met these criteria; reject option classification did not. We introduce a Supplementary Playbook outlining our approach for low-resource bias mitigation.
PMCID:12141433
PMID: 40473916
ISSN: 2398-6352
CID: 5862762
Periprocedural Myocardial Injury Using CKMB Following Elective PCI: Incidence and Associations With Long-Term Mortality
Talmor, Nina; Graves, Claire; Kozloff, Sam; Major, Vincent J; Xia, Yuhe; Shah, Binita; Babaev, Anvar; Razzouk, Louai; Rao, Sunil V; Attubato, Michael; Feit, Frederick; Slater, James; Smilowitz, Nathaniel R
BACKGROUND/UNASSIGNED:Myocardial injury detected after percutaneous coronary intervention (PCI) is associated with increased mortality. Predictors of post-PCI myocardial injury are not well established. The long-term prognostic relevance of post-PCI myocardial injury remains uncertain. METHODS/UNASSIGNED:Consecutive adults aged ≥18 years with stable ischemic heart disease who underwent elective PCI at NYU Langone Health between 2011 and 2020 were included in a retrospective, observational study. Patients with acute myocardial infarction or creatinine kinase myocardial band (CKMB) or troponin concentrations >99% of the upper reference limit before PCI were excluded. All patients had routine measurement of CKMB concentrations at 1 and 3 hours post-PCI. Post-PCI myocardial injury was defined as a peak CKMB concentration >99% upper reference limit. Linear regression models were used to identify clinical factors associated with post-PCI myocardial injury. Cox proportional hazard models were generated to evaluate relationships between post-PCI myocardial injury and all-cause mortality at long-term follow-up. RESULTS/UNASSIGNED:<0.001). After adjustment for demographics and clinical covariates, post-PCI myocardial injury was associated with an excess hazard for long-term mortality (hazard ratio, 1.46 [95% CI, 1.20-1.78]). CONCLUSIONS/UNASSIGNED:Myocardial injury defined by elevated CKMB early after PCI is common and associated with all-cause, long-term mortality. More complex coronary anatomy is predictive of post-PCI myocardial injury.
PMID: 40160098
ISSN: 1941-7632
CID: 5818652
Health system-wide access to generative artificial intelligence: the New York University Langone Health experience
Malhotra, Kiran; Wiesenfeld, Batia; Major, Vincent J; Grover, Himanshu; Aphinyanaphongs, Yindalon; Testa, Paul; Austrian, Jonathan S
OBJECTIVES/OBJECTIVE:The study aimed to assess the usage and impact of a private and secure instance of a generative artificial intelligence (GenAI) application in a large academic health center. The goal was to understand how employees interact with this technology and the influence on their perception of skill and work performance. MATERIALS AND METHODS/METHODS:New York University Langone Health (NYULH) established a secure, private, and managed Azure OpenAI service (GenAI Studio) and granted widespread access to employees. Usage was monitored and users were surveyed about their experiences. RESULTS:Over 6 months, over 1007 individuals applied for access, with high usage among research and clinical departments. Users felt prepared to use the GenAI studio, found it easy to use, and would recommend it to a colleague. Users employed the GenAI studio for diverse tasks such as writing, editing, summarizing, data analysis, and idea generation. Challenges included difficulties in educating the workforce in constructing effective prompts and token and API limitations. DISCUSSION/CONCLUSIONS:The study demonstrated high interest in and extensive use of GenAI in a healthcare setting, with users employing the technology for diverse tasks. While users identified several challenges, they also recognized the potential of GenAI and indicated a need for more instruction and guidance on effective usage. CONCLUSION/CONCLUSIONS:The private GenAI studio provided a useful tool for employees to augment their skills and apply GenAI to their daily tasks. The study underscored the importance of workforce education when implementing system-wide GenAI and provided insights into its strengths and weaknesses.
PMCID:11756645
PMID: 39584477
ISSN: 1527-974x
CID: 5778212
Evaluating Large Language Models in extracting cognitive exam dates and scores
Zhang, Hao; Jethani, Neil; Jones, Simon; Genes, Nicholas; Major, Vincent J; Jaffe, Ian S; Cardillo, Anthony B; Heilenbach, Noah; Ali, Nadia Fazal; Bonanni, Luke J; Clayburn, Andrew J; Khera, Zain; Sadler, Erica C; Prasad, Jaideep; Schlacter, Jamie; Liu, Kevin; Silva, Benjamin; Montgomery, Sophie; Kim, Eric J; Lester, Jacob; Hill, Theodore M; Avoricani, Alba; Chervonski, Ethan; Davydov, James; Small, William; Chakravartty, Eesha; Grover, Himanshu; Dodson, John A; Brody, Abraham A; Aphinyanaphongs, Yindalon; Masurkar, Arjun; Razavian, Narges
Ensuring reliability of Large Language Models (LLMs) in clinical tasks is crucial. Our study assesses two state-of-the-art LLMs (ChatGPT and LlaMA-2) for extracting clinical information, focusing on cognitive tests like MMSE and CDR. Our data consisted of 135,307 clinical notes (Jan 12th, 2010 to May 24th, 2023) mentioning MMSE, CDR, or MoCA. After applying inclusion criteria 34,465 notes remained, of which 765 underwent ChatGPT (GPT-4) and LlaMA-2, and 22 experts reviewed the responses. ChatGPT successfully extracted MMSE and CDR instances with dates from 742 notes. We used 20 notes for fine-tuning and training the reviewers. The remaining 722 were assigned to reviewers, with 309 each assigned to two reviewers simultaneously. Inter-rater-agreement (Fleiss' Kappa), precision, recall, true/false negative rates, and accuracy were calculated. Our study follows TRIPOD reporting guidelines for model validation. For MMSE information extraction, ChatGPT (vs. LlaMA-2) achieved accuracy of 83% (vs. 66.4%), sensitivity of 89.7% (vs. 69.9%), true-negative rates of 96% (vs 60.0%), and precision of 82.7% (vs 62.2%). For CDR the results were lower overall, with accuracy of 87.1% (vs. 74.5%), sensitivity of 84.3% (vs. 39.7%), true-negative rates of 99.8% (98.4%), and precision of 48.3% (vs. 16.1%). We qualitatively evaluated the MMSE errors of ChatGPT and LlaMA-2 on double-reviewed notes. LlaMA-2 errors included 27 cases of total hallucination, 19 cases of reporting other scores instead of MMSE, 25 missed scores, and 23 cases of reporting only the wrong date. In comparison, ChatGPT's errors included only 3 cases of total hallucination, 17 cases of wrong test reported instead of MMSE, and 19 cases of reporting a wrong date. In this diagnostic/prognostic study of ChatGPT and LlaMA-2 for extracting cognitive exam dates and scores from clinical notes, ChatGPT exhibited high accuracy, with better performance compared to LlaMA-2. The use of LLMs could benefit dementia research and clinical care, by identifying eligible patients for treatments initialization or clinical trial enrollments. Rigorous evaluation of LLMs is crucial to understanding their capabilities and limitations.
PMCID:11634005
PMID: 39661652
ISSN: 2767-3170
CID: 5762692
Antibiotic stewardship bundle for uncomplicated gram-negative bacteremia at an academic health system: a quasi-experimental study
DiPietro, Juliana; Dubrovskaya, Yanina; Marsh, Kassandra; Decano, Arnold; Papadopoulos, John; Mazo, Dana; Inglima, Kenneth; Major, Vincent; So, Jonathon; Yuditskiy, Samuel; Siegfried, Justin
OBJECTIVE/UNASSIGNED:To evaluate whether an antimicrobial stewardship bundle (ASB) can safely empower frontline providers in the treatment of gram-negative bloodstream infections (GN-BSI). INTERVENTION AND METHOD/UNASSIGNED:From March 2021 to February 2022, we implemented an ASB intervention for GN-BSI in the electronic medical record (EMR) to guide clinicians at the point of care to optimize their own antibiotic decision-making. We conducted a before-and-after quasi-experimental pre-bundle (preBG) and post-bundle (postBG) study evaluating a composite of in-hospital mortality, infection-related readmission, GN-BSI recurrence, and bundle-related outcomes. SETTING/UNASSIGNED:New York University Langone Health (NYULH), Tisch/Kimmel (T/K) and Brooklyn (BK) campuses, in New York City, New York. PATIENTS/UNASSIGNED:Out of 1097 patients screened, the study included 225 adults aged ≥18 years (101 preBG vs 124 postBG) admitted with at least one positive blood culture with a monomicrobial gram-negative organism. RESULTS/UNASSIGNED:= 0.043. CONCLUSIONS/UNASSIGNED:GN-BSI bundle worked as a nudge-based strategy to guide providers in VAN DC and increased de-escalation to aminopenicillin-based antibiotics without negatively impacting patient outcomes.
PMCID:11474889
PMID: 39411661
ISSN: 2732-494x
CID: 5718532
Evaluating Large Language Models in Extracting Cognitive Exam Dates and Scores
Zhang, Hao; Jethani, Neil; Jones, Simon; Genes, Nicholas; Major, Vincent J; Jaffe, Ian S; Cardillo, Anthony B; Heilenbach, Noah; Ali, Nadia Fazal; Bonanni, Luke J; Clayburn, Andrew J; Khera, Zain; Sadler, Erica C; Prasad, Jaideep; Schlacter, Jamie; Liu, Kevin; Silva, Benjamin; Montgomery, Sophie; Kim, Eric J; Lester, Jacob; Hill, Theodore M; Avoricani, Alba; Chervonski, Ethan; Davydov, James; Small, William; Chakravartty, Eesha; Grover, Himanshu; Dodson, John A; Brody, Abraham A; Aphinyanaphongs, Yindalon; Masurkar, Arjun; Razavian, Narges
IMPORTANCE/UNASSIGNED:Large language models (LLMs) are crucial for medical tasks. Ensuring their reliability is vital to avoid false results. Our study assesses two state-of-the-art LLMs (ChatGPT and LlaMA-2) for extracting clinical information, focusing on cognitive tests like MMSE and CDR. OBJECTIVE/UNASSIGNED:Evaluate ChatGPT and LlaMA-2 performance in extracting MMSE and CDR scores, including their associated dates. METHODS/UNASSIGNED:Our data consisted of 135,307 clinical notes (Jan 12th, 2010 to May 24th, 2023) mentioning MMSE, CDR, or MoCA. After applying inclusion criteria 34,465 notes remained, of which 765 underwent ChatGPT (GPT-4) and LlaMA-2, and 22 experts reviewed the responses. ChatGPT successfully extracted MMSE and CDR instances with dates from 742 notes. We used 20 notes for fine-tuning and training the reviewers. The remaining 722 were assigned to reviewers, with 309 each assigned to two reviewers simultaneously. Inter-rater-agreement (Fleiss' Kappa), precision, recall, true/false negative rates, and accuracy were calculated. Our study follows TRIPOD reporting guidelines for model validation. RESULTS/UNASSIGNED:For MMSE information extraction, ChatGPT (vs. LlaMA-2) achieved accuracy of 83% (vs. 66.4%), sensitivity of 89.7% (vs. 69.9%), true-negative rates of 96% (vs 60.0%), and precision of 82.7% (vs 62.2%). For CDR the results were lower overall, with accuracy of 87.1% (vs. 74.5%), sensitivity of 84.3% (vs. 39.7%), true-negative rates of 99.8% (98.4%), and precision of 48.3% (vs. 16.1%). We qualitatively evaluated the MMSE errors of ChatGPT and LlaMA-2 on double-reviewed notes. LlaMA-2 errors included 27 cases of total hallucination, 19 cases of reporting other scores instead of MMSE, 25 missed scores, and 23 cases of reporting only the wrong date. In comparison, ChatGPT's errors included only 3 cases of total hallucination, 17 cases of wrong test reported instead of MMSE, and 19 cases of reporting a wrong date. CONCLUSIONS/UNASSIGNED:In this diagnostic/prognostic study of ChatGPT and LlaMA-2 for extracting cognitive exam dates and scores from clinical notes, ChatGPT exhibited high accuracy, with better performance compared to LlaMA-2. The use of LLMs could benefit dementia research and clinical care, by identifying eligible patients for treatments initialization or clinical trial enrollments. Rigorous evaluation of LLMs is crucial to understanding their capabilities and limitations.
PMCID:10888985
PMID: 38405784
CID: 5722422