Try a new search

Format these results:

Searched for:

in-biosketch:true

person:smallw03

Total Results:

23


Disappearing Text as a Clinical Decision Support Layer: A Case Series

Silberlust, Jared; Small, William; Shah, Darshi; Chakravartty, Eesha; Moawad, Katherine; Moawad, Andrew; Testa, Paul; Feldman, Jonah
OBJECTIVES/OBJECTIVE:This case series aims to evaluate several applications of inline disappearing text (DT) clinical decision support (CDS) tools within clinician documentation. METHODS:DT blocks were created to prompt documentation for perioperative anticoagulation planning (Scenario 1), pre-discharge intravenous antibiotic planning (Scenario 2), and advanced care planning (Scenario 3). In Scenario 1, DT was the only intervention. In Scenario 2, DT was paired with a documentation SmartList. In Scenario 3, DT was paired with a documentation SmartList and an OurPractice Advisory. The number of documented perioperative anticoagulation plans, pre-discharge intravenous antibiotic plans, and Advanced Care Planning notes were measured pre- and post-intervention and compared using Chi-square analyses. RESULTS:In Scenario 1, there was no statistically significant change in the percentage of perioperative anticoagulation plans documented at 0-24 and 24-48 hours before surgery. In Scenario 2, documentation of antibiotic contingency planning in patients expected to be discharged within 24 hours increased from 60% (54 of 90 notes) to 93% (1,850 of 1,994 notes) X2 (1, N=2,084) = 113.1, p < 0.001. In Scenario 3, ACP note documentation by discharge in patients with a positive mandatory surprise question increased from 43% (821 of 1,909 encounters) to 52% (975 of 1,874 encounters) X2 (1, N=3,783) = 30.5, p < 0.001. CONCLUSIONS:Utilizing DT in conjunction with other forms of CDS was associated with an improvement of documentation quality in pre-discharge IV antibiotics and advanced care planning. A sociotechnical analysis explores how interactions between technology, people, workflow, and culture could contextualize how utilizing DT with other forms of CDS was more effective than DT alone.
PMID: 40763805
ISSN: 1869-0327
CID: 5905032

ACR Appropriateness Criteria® Supplemental Breast Cancer Screening Based on Breast Density: 2024 Update

,; Paulis, Lisa V; Lewin, Alana A; Weinstein, Susan P; Baron, Paul; Dayaratna, Sandra; Dodelzon, Katerina; Dogan, Basak E; Gulati, Abhishek; Kantor, Olga; Kasales, Claudia; Kunjummen, Jean M; Kuzmiak, Cherie M; Newell, Mary S; Salkowski, Lonie R; Sharpe, Richard E; Small, William; Ulaner, Gary A; Slanetz, Priscilla J
Screening mammography has been proven to reduce the mortality from breast cancer by approximately 30%, however, it is less sensitive in women with dense breast tissue and certain risk groups. Supplemental screening may be considered based on the patient's risk level and breast density. In all women, digital breast tomosynthesis improves screening sensitivity. Average-risk women with heterogeneously dense tissue may also benefit from breast MRI, abbreviated breast MRI (AB-MRI) or breast ultrasound (US). In intermediate-risk women with nondense tissue, breast MRI and ABMRI may be appropriate. In intermediate-risk women with heterogeneously dense and extremely dense tissue, breast MRI and AB-MRI are usually appropriate, whereas US and contrast-enhanced mammography (CEM) may be appropriate. Breast MRI or ABMRI is usually appropriate in all high-risk women, regardless of density. Screening breast US or CEM could be considered in this population. The American College of Radiology Appropriateness Criteria are evidence-based guidelines for specific clinical conditions that are reviewed annually by a multidisciplinary expert panel. The guideline development and revision process support the systematic analysis of the medical literature from peer reviewed journals. Established methodology principles such as Grading of Recommendations Assessment, Development, and Evaluation or GRADE are adapted to evaluate the evidence. The RAND/UCLA Appropriateness Method User Manual provides the methodology to determine the appropriateness of imaging and treatment procedures for specific clinical scenarios. In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.
PMID: 40409891
ISSN: 1558-349x
CID: 5853762

ACR Appropriateness Criteria® Ovarian Cancer Screening: 2024 Update

,; Venkatesan, Aradhana M; Kilcoyne, Aoife; Akin, Esma A; Chuang, Linus; Hindman, Nicole M; Huang, Chenchan; McCourt, Carolyn Kay; Rauch, Gaiane M; Sattari, Maryam; Schoenborn, Nancy; Schultz, David; Sertic, Madeleine; Small, William; Stein, Erica B; Suarez-Weiss, Krista; Kang, Stella K
Ovarian cancer remains low in prevalence but has the highest mortality of all gynecologic malignancies. Population-based screening for ovarian cancer remains a topic of interest in contemporary practice, given that the majority of cancers encountered are high-grade aggressive malignancies, for which favorable survival is encountered in the setting of early-stage disease. This document summarizes a review of the available data from randomized and observational trials that have evaluated the role of imaging for ovarian cancer screening in average-risk and high-risk patients. When considering screening using pelvic ultrasound in average-risk patients, we found insufficient published evidence to recommend ovarian cancer screening. Randomized controlled trials have not demonstrated a mortality benefit in this setting. Screening with pelvic ultrasound may be appropriate for select patients at high risk, although the existing data remain limited as large, randomized trials have not been performed in this setting. The American College of Radiology Appropriateness Criteria are evidence-based guidelines for specific clinical conditions that are reviewed annually by a multidisciplinary expert panel. The guideline development and revision process support the systematic analysis of the medical literature from peer reviewed journals. Established methodology principles such as Grading of Recommendations Assessment, Development, and Evaluation or GRADE are adapted to evaluate the evidence. The RAND/UCLA Appropriateness Method User Manual provides the methodology to determine the appropriateness of imaging and treatment procedures for specific clinical scenarios. In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.
PMID: 40409887
ISSN: 1558-349x
CID: 5853732

Classifying Continuous Glucose Monitoring Documents From Electronic Health Records

Zheng, Yaguang; Iturrate, Eduardo; Li, Lehan; Wu, Bei; Small, William R; Zweig, Susan; Fletcher, Jason; Chen, Zhihao; Johnson, Stephen B
BACKGROUND:Clinical use of continuous glucose monitoring (CGM) is increasing storage of CGM-related documents in electronic health records (EHR); however, the standardization of CGM storage is lacking. We aimed to evaluate the sensitivity and specificity of CGM Ambulatory Glucose Profile (AGP) classification criteria. METHODS:We randomly chose 2244 (18.1%) documents from NYU Langone Health. Our document classification algorithm: (1) separated multiple-page documents into a single-page image; (2) rotated all pages into an upright orientation; (3) determined types of devices using optical character recognition; and (4) tested for the presence of particular keywords in the text. Two experts in using CGM for research and clinical practice conducted an independent manual review of 62 (2.8%) reports. We calculated sensitivity (correct classification of CGM AGP report) and specificity (correct classification of non-CGM report) by comparing the classification algorithm against manual review. RESULTS:Among 2244 documents, 1040 (46.5%) were classified as CGM AGP reports (43.3% FreeStyle Libre and 56.7% Dexcom), 1170 (52.1%) non-CGM reports (eg, progress notes, CGM request forms, or physician letters), and 34 (1.5%) uncertain documents. The agreement for the evaluation of the documents between the two experts was 100% for sensitivity and 98.4% for specificity. When comparing the classification result between the algorithm and manual review, the sensitivity and specificity were 95.0% and 91.7%. CONCLUSION/CONCLUSIONS:Nearly half of CGM-related documents were AGP reports, which are useful for clinical practice and diabetes research; however, the remaining half are other clinical documents. Future work needs to standardize the storage of CGM-related documents in the EHR.
PMCID:11904921
PMID: 40071848
ISSN: 1932-2968
CID: 5808452

The Clinical Utility of a 7-Gene Biosignature on Radiation Therapy Decision Making in Patients with Ductal Carcinoma In Situ Following Breast-Conserving Surgery: An Updated Analysis of the DCISionRT® PREDICT Study

Shah, Chirag; Whitworth, Pat; Vicini, Frank A; Narod, Steven; Gerber, Naamit; Jhawar, Sachin R; King, Tari A; Mittendorf, Elizabeth A; Willey, Shawna C; Rabinovich, Rachel; Gold, Linsey; Brown, Eric; Patel, Anushka; Vargo, John; Barry, Parul N; Rock, David; Friedman, Neil; Bedi, Gauri; Templeton, Sandra; Brown, Sheree; Gabordi, Robert; Riley, Lee; Lee, Lucy; Baron, Paul; Majithia, Lonika; Mirabeau-Beale, Kristina L; Reid, Vincent J; Hirsch, Arica; Hwang, Catherine; Pellicane, James; Maganini, Robert; Khan, Sadia; MacDermed, Dhara M; Small, William; Mittal, Karuna; Borgen, Patrick; Cox, Charles; Shivers, Steven C; Bremer, Troy
BACKGROUND:Breast-conserving surgery (BCS) followed by adjuvant radiotherapy (RT) is a standard treatment for ductal carcinoma in situ (DCIS). A low-risk patient subset that does not benefit from RT has not yet been clearly identified. The DCISionRT test provides a clinically validated decision score (DS), which is prognostic of 10-year in-breast recurrence rates (invasive and non-invasive) and is also predictive of RT benefit. This analysis presents final outcomes from the PREDICT prospective registry trial aiming to determine how often the DCISionRT test changes radiation treatment recommendations. METHODS:Overall, 2496 patients were enrolled from February 2018 to January 2022 at 63 academic and community practice sites and received DCISionRT as part of their care plan. Treating physicians reported their treatment recommendations pre- and post-test as well as the patient's preference. The primary endpoint was to identify the percentage of patients where testing led to a change in RT recommendation. The impact of the test on RT treatment recommendation was physician specialty, treatment settings, individual clinical/pathological features and RTOG 9804 like criteria. Multivariate logisitc regression analysis was used to estimate the odds ratio (ORs) for factors associated with the post-test RT recommendations. RESULTS:RT recommendation changed 38% of women, resulting in a 20% decrease in the overall recommendation of RT (p < 0.001). Of those women initially recommended no RT (n = 583), 31% were recommended RT post-test. The recommendation for RT post-test increased with increasing DS, from 29% to 66% to 91% for DS <2, DS 2-4, and DS >4, respectively. On multivariable analysis, DS had the strongest influence on final RT recommendation (odds ratio 22.2, 95% confidence interval 16.3-30.7), which was eightfold greater than clinicopathologic features. Furthermore, there was an overall change in the recommendation to receive RT in 42% of those patients meeting RTOG 9804-like low-risk criteria. CONCLUSIONS:The test results provided information that changes treatment recommendations both for and against RT use in large population of women with DCIS treated in a variety of clinical settings. Overall, clinicians changed their recommendations to include or omit RT for 38% of women based on the test results. Based on published clinical validations and the results from current study, DCISionRT may aid in preventing the over- and undertreatment of clinicopathological 'low-risk' and 'high-risk' DCIS patients. TRIAL REGISTRATION/BACKGROUND:ClinicalTrials.gov identifier: NCT03448926 ( https://clinicaltrials.gov/study/NCT03448926 ).
PMCID:11300542
PMID: 38916700
ISSN: 1534-4681
CID: 5973052

Evaluating Large Language Models in extracting cognitive exam dates and scores

Zhang, Hao; Jethani, Neil; Jones, Simon; Genes, Nicholas; Major, Vincent J; Jaffe, Ian S; Cardillo, Anthony B; Heilenbach, Noah; Ali, Nadia Fazal; Bonanni, Luke J; Clayburn, Andrew J; Khera, Zain; Sadler, Erica C; Prasad, Jaideep; Schlacter, Jamie; Liu, Kevin; Silva, Benjamin; Montgomery, Sophie; Kim, Eric J; Lester, Jacob; Hill, Theodore M; Avoricani, Alba; Chervonski, Ethan; Davydov, James; Small, William; Chakravartty, Eesha; Grover, Himanshu; Dodson, John A; Brody, Abraham A; Aphinyanaphongs, Yindalon; Masurkar, Arjun; Razavian, Narges
Ensuring reliability of Large Language Models (LLMs) in clinical tasks is crucial. Our study assesses two state-of-the-art LLMs (ChatGPT and LlaMA-2) for extracting clinical information, focusing on cognitive tests like MMSE and CDR. Our data consisted of 135,307 clinical notes (Jan 12th, 2010 to May 24th, 2023) mentioning MMSE, CDR, or MoCA. After applying inclusion criteria 34,465 notes remained, of which 765 underwent ChatGPT (GPT-4) and LlaMA-2, and 22 experts reviewed the responses. ChatGPT successfully extracted MMSE and CDR instances with dates from 742 notes. We used 20 notes for fine-tuning and training the reviewers. The remaining 722 were assigned to reviewers, with 309 each assigned to two reviewers simultaneously. Inter-rater-agreement (Fleiss' Kappa), precision, recall, true/false negative rates, and accuracy were calculated. Our study follows TRIPOD reporting guidelines for model validation. For MMSE information extraction, ChatGPT (vs. LlaMA-2) achieved accuracy of 83% (vs. 66.4%), sensitivity of 89.7% (vs. 69.9%), true-negative rates of 96% (vs 60.0%), and precision of 82.7% (vs 62.2%). For CDR the results were lower overall, with accuracy of 87.1% (vs. 74.5%), sensitivity of 84.3% (vs. 39.7%), true-negative rates of 99.8% (98.4%), and precision of 48.3% (vs. 16.1%). We qualitatively evaluated the MMSE errors of ChatGPT and LlaMA-2 on double-reviewed notes. LlaMA-2 errors included 27 cases of total hallucination, 19 cases of reporting other scores instead of MMSE, 25 missed scores, and 23 cases of reporting only the wrong date. In comparison, ChatGPT's errors included only 3 cases of total hallucination, 17 cases of wrong test reported instead of MMSE, and 19 cases of reporting a wrong date. In this diagnostic/prognostic study of ChatGPT and LlaMA-2 for extracting cognitive exam dates and scores from clinical notes, ChatGPT exhibited high accuracy, with better performance compared to LlaMA-2. The use of LLMs could benefit dementia research and clinical care, by identifying eligible patients for treatments initialization or clinical trial enrollments. Rigorous evaluation of LLMs is crucial to understanding their capabilities and limitations.
PMCID:11634005
PMID: 39661652
ISSN: 2767-3170
CID: 5762692

Evaluating Large Language Models in Extracting Cognitive Exam Dates and Scores

Zhang, Hao; Jethani, Neil; Jones, Simon; Genes, Nicholas; Major, Vincent J; Jaffe, Ian S; Cardillo, Anthony B; Heilenbach, Noah; Ali, Nadia Fazal; Bonanni, Luke J; Clayburn, Andrew J; Khera, Zain; Sadler, Erica C; Prasad, Jaideep; Schlacter, Jamie; Liu, Kevin; Silva, Benjamin; Montgomery, Sophie; Kim, Eric J; Lester, Jacob; Hill, Theodore M; Avoricani, Alba; Chervonski, Ethan; Davydov, James; Small, William; Chakravartty, Eesha; Grover, Himanshu; Dodson, John A; Brody, Abraham A; Aphinyanaphongs, Yindalon; Masurkar, Arjun; Razavian, Narges
IMPORTANCE/UNASSIGNED:Large language models (LLMs) are crucial for medical tasks. Ensuring their reliability is vital to avoid false results. Our study assesses two state-of-the-art LLMs (ChatGPT and LlaMA-2) for extracting clinical information, focusing on cognitive tests like MMSE and CDR. OBJECTIVE/UNASSIGNED:Evaluate ChatGPT and LlaMA-2 performance in extracting MMSE and CDR scores, including their associated dates. METHODS/UNASSIGNED:Our data consisted of 135,307 clinical notes (Jan 12th, 2010 to May 24th, 2023) mentioning MMSE, CDR, or MoCA. After applying inclusion criteria 34,465 notes remained, of which 765 underwent ChatGPT (GPT-4) and LlaMA-2, and 22 experts reviewed the responses. ChatGPT successfully extracted MMSE and CDR instances with dates from 742 notes. We used 20 notes for fine-tuning and training the reviewers. The remaining 722 were assigned to reviewers, with 309 each assigned to two reviewers simultaneously. Inter-rater-agreement (Fleiss' Kappa), precision, recall, true/false negative rates, and accuracy were calculated. Our study follows TRIPOD reporting guidelines for model validation. RESULTS/UNASSIGNED:For MMSE information extraction, ChatGPT (vs. LlaMA-2) achieved accuracy of 83% (vs. 66.4%), sensitivity of 89.7% (vs. 69.9%), true-negative rates of 96% (vs 60.0%), and precision of 82.7% (vs 62.2%). For CDR the results were lower overall, with accuracy of 87.1% (vs. 74.5%), sensitivity of 84.3% (vs. 39.7%), true-negative rates of 99.8% (98.4%), and precision of 48.3% (vs. 16.1%). We qualitatively evaluated the MMSE errors of ChatGPT and LlaMA-2 on double-reviewed notes. LlaMA-2 errors included 27 cases of total hallucination, 19 cases of reporting other scores instead of MMSE, 25 missed scores, and 23 cases of reporting only the wrong date. In comparison, ChatGPT's errors included only 3 cases of total hallucination, 17 cases of wrong test reported instead of MMSE, and 19 cases of reporting a wrong date. CONCLUSIONS/UNASSIGNED:In this diagnostic/prognostic study of ChatGPT and LlaMA-2 for extracting cognitive exam dates and scores from clinical notes, ChatGPT exhibited high accuracy, with better performance compared to LlaMA-2. The use of LLMs could benefit dementia research and clinical care, by identifying eligible patients for treatments initialization or clinical trial enrollments. Rigorous evaluation of LLMs is crucial to understanding their capabilities and limitations.
PMCID:10888985
PMID: 38405784
CID: 5722422

Clinical research in endometrial cancer: consensus recommendations from the Gynecologic Cancer InterGroup

Creutzberg, Carien L; Kim, Jae-Weon; Eminowicz, Gemma; Allanson, Emma; Eberst, Lauriane; Kim, Se Ik; Nout, Remi A; Park, Jeong-Yeol; Lorusso, Domenica; Mileshkin, Linda; Ottevanger, Petronella B; Brand, Alison; Mezzanzanica, Delia; Oza, Amit; Gebski, Val; Pothuri, Bhavana; Batley, Tania; Gordon, Carol; Mitra, Tina; White, Helen; Howitt, Brooke; Matias-Guiu, Xavier; Ray-Coquard, Isabelle; Gaffney, David; Small, William; Miller, Austin; Concin, Nicole; Powell, Matthew A; Stuart, Gavin; Bookman, Michael A; ,
The Gynecologic Cancer InterGroup (GCIG) Endometrial Cancer Consensus Conference on Clinical Research (ECCC) was held in Incheon, South Korea, Nov 2-3, 2023. The aims were to develop consensus statements for future trials in endometrial cancer to achieve harmonisation on design elements, select important questions, and identify unmet needs. All 33 GCIG member groups participated in the development, refinement, and finalisation of 18 statements within four topic groups, addressing adjuvant treatment in high-risk disease; treatment for metastatic and recurrent disease; trial designs for rare endometrial cancer subgroups and special circumstances; and specific methodology and adaptation for trials in low-resource settings. In addition, eight areas of unmet need were identified. This was the first GCIG Consensus Conference to include patient advocates and an expert on inclusion, diversity, equity, and access to take part in all aspects of the process and output. Four early-career investigators were also selected for participation, ensuring that they represented different GCIG member groups and regions. Unanimous consensus was obtained for 16 of the 18 statements, with 97% concordance for the remaining two. Using the described methodology from previous Ovarian Cancer Consensus Conferences, this conference did not require even one minority statement. The high acceptance rate following active involvement in the preparation, discussion, and refinement of the statements by all representatives confirmed the consensus progress within a global academic setting, and the expectation that the ECCC will lead to greater harmonisation, actualisation, inclusion, and resolution of unmet needs in clinical research for individuals living with and beyond endometrial cancer worldwide.
PMID: 39214113
ISSN: 1474-5488
CID: 5702082

Enhancing Secure Messaging in Electronic Health Records: Evaluating the Impact of Emoji Chat Reactions on the Volume of Interruptive Notifications

Will, John; Small, William; Iturrate, Eduardo; Testa, Paul; Feldman, Jonah
ORIGINAL:0017336
ISSN: 2566-9346
CID: 5686602

The First Generative AI Prompt-A-Thon in Healthcare: A Novel Approach to Workforce Engagement with a Private Instance of ChatGPT

Small, William R; Malhotra, Kiran; Major, Vincent J; Wiesenfeld, Batia; Lewis, Marisa; Grover, Himanshu; Tang, Huming; Banerjee, Arnab; Jabbour, Michael J; Aphinyanaphongs, Yindalon; Testa, Paul; Austrian, Jonathan S
BACKGROUND:Healthcare crowdsourcing events (e.g. hackathons) facilitate interdisciplinary collaboration and encourage innovation. Peer-reviewed research has not yet considered a healthcare crowdsourcing event focusing on generative artificial intelligence (GenAI), which generates text in response to detailed prompts and has vast potential for improving the efficiency of healthcare organizations. Our event, the New York University Langone Health (NYULH) Prompt-a-thon, primarily sought to inspire and build AI fluency within our diverse NYULH community, and foster collaboration and innovation. Secondarily, we sought to analyze how participants' experience was influenced by their prior GenAI exposure and whether they received sample prompts during the workshop. METHODS:Executing the event required the assembly of an expert planning committee, who recruited diverse participants, anticipated technological challenges, and prepared the event. The event was composed of didactics and workshop sessions, which educated and allowed participants to experiment with using GenAI on real healthcare data. Participants were given novel "project cards" associated with each dataset that illuminated the tasks GenAI could perform and, for a random set of teams, sample prompts to help them achieve each task (the public repository of project cards can be found at https://github.com/smallw03/NYULH-Generative-AI-Prompt-a-thon-Project-Cards). Afterwards, participants were asked to fill out a survey with 7-point Likert-style questions. RESULTS:Our event was successful in educating and inspiring hundreds of enthusiastic in-person and virtual participants across our organization on the responsible use of GenAI in a low-cost and technologically feasible manner. All participants responded positively, on average, to each of the survey questions (e.g., confidence in their ability to use and trust GenAI). Critically, participants reported a self-perceived increase in their likelihood of using and promoting colleagues' use of GenAI for their daily work. No significant differences were seen in the surveys of those who received sample prompts with their project task descriptions. CONCLUSION/CONCLUSIONS:The first healthcare Prompt-a-thon was an overwhelming success, with minimal technological failures, positive responses from diverse participants and staff, and evidence of post-event engagement. These findings will be integral to planning future events at our institution, and to others looking to engage their workforce in utilizing GenAI.
PMCID:11265701
PMID: 39042600
ISSN: 2767-3170
CID: 5686592