Try a new search

Format these results:

Searched for:

person:smallw03

in-biosketch:true

Total Results:

22


Electronic Health Record Messaging Patterns of Health Care Professionals in Inpatient Medicine

Small, William; Iturrate, Eduardo; Austrian, Jonathan; Genes, Nicholas
PMID: 38147337
ISSN: 2574-3805
CID: 5623492

Real-time EHR secure messaging to coordinate emergency department disposition for 30-day revisit patients

Solanki, Priyanka; Small, William; Sondhi, Jaya; Jones, Simon; Genes, Nicholas; Mansukhani, Ajay; Prabhu, Dinesha; Turley, Reed; Pineda, Edwin; Johnson, David; Moeller, Benjamin; Bosworth, Brian; Austrian, Jonathan
OBJECTIVES/OBJECTIVE:To evaluate whether real-time electronic health record (EHR)-based secure messaging between emergency department (ED) clinicians and prior discharge teams influences ED disposition decisions for patients re-presenting within 30 days of hospital discharge. MATERIALS AND METHODS/METHODS:This 18-month pre-post study included 27 592 ED revisit encounters across 3 campuses within one academic healthcare system. Robotic process automation generated a real-time EHR secure message connecting the ED attending with the index hospitalization discharge team during the ED disposition window. The primary outcome was the proportion of encounters resulting in ED disposition changes. Secondary outcomes included length of stay and messaging engagement by service, hospital campus, and time of day. RESULTS:Systemwide, inpatient readmissions rates were unchanged (49.4% vs 48.6%, P = .149), and observation status increased (7.0% vs 8.0%, P < .001). At one campus where messaging was paired with proactive care coordination, inpatient admissions decreased (56.9% vs 54.1%, P = .029) and treat-and-release increased (36.8% vs 39.6%, P = .024). Overall, 61.8% of messages received a response, but engagement did not correlate with disposition changes systemwide. DISCUSSION/CONCLUSIONS:Disposition changes occurred only where messaging was integrated with care coordination workflows with the operational capacity to act, indicating that real-time communication alone is insufficient without supporting infrastructure. CONCLUSION/CONCLUSIONS:Real-time EHR messaging may be most effective when paired with structured care coordination models rather than deployed as a standalone alerting tool.
PMID: 42531463
ISSN: 1527-974x
CID: 6070461

General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

Vishwanath, Krithik; Alyakin, Anton; Ghosh, Mrigayu; Hage, Ali; Neifert, Sean N; Orillac, Cordelia; Mandelberg, Nataniel J; Khan, Hammad A; Lee, Jin Vivian; Yao, Jie J; Small, William Robert; Varma, Aakaash; Hewitt, D Brock; Aphinyanaphongs, Yindalon; Alber, Daniel Alexander; Oermann, Eric Karl
Specialized clinical artificial intelligence (AI) tools are entering medical practice despite scarce independent evaluation. We quantitatively evaluate two clinical AI tools, OpenEvidence and UpToDate Expert AI, built on large language models (LLMs) against three frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-identified queries from physicians to a general-purpose language model in a live clinical environment. For the RCQ benchmark, 12 US clinicians performed randomized, blinded review of model outputs, producing 1,800 model-question annotations. Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ. These findings highlight the need for independent, real-world evaluation of AI tools before they enter clinical settings.
PMID: 42286322
ISSN: 1546-170x
CID: 6049082

Enhancing the prediction of hospital discharge disposition with extraction-based language model classification

Small, William R; Crowley, Ryan J; Pariente, Chloe; Zhang, Jeff; Eaton, Kevin P; Jiang, Lavender Yao; Oermann, Eric; Aphinyanaphongs, Yindalon
Early identification of inpatient discharges to skilled nursing facilities (SNFs) facilitates care transition planning. Predictive information in admission history and physical notes (H&Ps) is dispersed across long documents. Language models adeptly predict clinical outcomes from text but have limitations: token length constraints, noisy inputs, and opaque outputs. Therefore, we developed extraction-based language model classification (ELC): generative language models distill H&Ps into task-relevant categories ("Structured Extracted Data") before summarizing them into a concise narrative ("AI Risk Snapshot"). We hypothesized that language models utilizing AI Risk Snapshots to predict SNF discharges would perform the best. In this retrospective observational study, nine language models predicted SNF discharges from unstructured predictors (raw H&P text, truncated assessment and plan) and ELC-derived predictors (Structured Extracted Data, AI Risk Snapshots). ELC substantially reduced input length (AI Risk Snapshot median 141 tokens vs raw H&P median 2,120 tokens) and improved average AUROC and AUPRC across models. The best performance was achieved by Bio+Clinical BERT fine-tuned on AI Risk Snapshots (AUROC = .851). AI Risk Snapshots enhanced interpretability by aligning with nurse case managers' risk assessments and facilitating prompt design. Structuring and summarizing H&Ps via ELC thus mitigates the practical limitations of language models and improves SNF discharge prediction.
PMCID:12789015
PMID: 41522677
ISSN: 3005-1959
CID: 5985892

ACR Appropriateness Criteria® Female Breast Cancer Screening: 2025 Update

,; Yeh, Eren D; Brown, Ann; Freer, Phoebe E; Bahl, Manisha; Bennett, Debbie L; Darbha, Lalitha; Dibble, Elizabeth H; Greenwood, Heather I; Hill, Faihza M; Ivansco, Lillian K; Kremer, Mallory E; Minami, Christina A; Mullen, Lisa A; Neal, Colleen H; Newell, Mary S; Radhakrishnan, Archana; Rauch, Gaiane M; Reig, Beatriu; Shaughnessy, Elizabeth; Small, William; Ulaner, Gary A; Lewin, Alana A
Routine screening substantially reduces the risk of mortality and morbidity of breast cancer with early detection. Multiple different imaging modalities may be used to screen for breast cancer. Screening recommendations differ based on an individual's risk of developing breast cancer. Numerous factors contribute to breast cancer risk, which is frequently divided into three major categories: average, intermediate, and high risk. For patients assigned female at birth with native breast tissue, mammography and digital breast tomosynthesis are recommended for breast cancer screening in all risk categories. In high-risk patients, screening with breast MRI is recommended starting as early as 25 to 30 years of age and mammography and digital breast tomosynthesis with a variable starting age between 25 and 40 years of age, depending on the type of risk. The American College of Radiology Appropriateness Criteria are evidence-based guidelines for specific clinical conditions that are reviewed annually by a multidisciplinary expert panel. The guideline development and revision process support the systematic analysis of the medical literature from peer reviewed journals. Established methodology principles such as Grading of Recommendations Assessment, Development, and Evaluation or GRADE are adapted to evaluate the evidence. The RAND/UCLA Appropriateness Method User Manual provides the methodology to determine the appropriateness of imaging and treatment procedures for specific clinical scenarios. In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.
PMID: 41193041
ISSN: 1558-349x
CID: 5959892

ACR Appropriateness Criteria® Screening, Locoregional Assessment, and Surveillance of Pancreatic Ductal Adenocarcinoma: 2025 Update

,; Fung, Alice; Zaheer, Atif; Porter, Kristin K; Bashir, Mustafa R; Cash, Brooks D; Chiorean, E Gabriela; Choi, Youngjee; Ejaz, Aslam; Gage, Kenneth L; Russo, Gregory K; Small, William; Smith, Elainea N; Thakrar, Kiran H; Vij, Abhinav; Wahab, Shaun A; Kim, David H
Pancreatic ductal adenocarcinoma is a highly lethal cancer that often presents with vague and indolent symptoms leading to advanced stage diagnosis. Imaging plays a crucial role in the diagnosis, assessment of locoregional and metastatic disease, surgical planning, and surveillance after neoadjuvant therapy and surgery. This document reviews available imaging modalities that are best used for these clinical scenarios, and a summary of current evidence is provided to support the use of the various modalities in each of the clinical contexts. The American College of Radiology Appropriateness Criteria are evidence-based guidelines for specific clinical conditions that are reviewed annually by a multidisciplinary expert panel. The guideline development and revision process support the systematic analysis of the medical literature from peer reviewed journals. Established methodology principles such as Grading of Recommendations Assessment, Development, and Evaluation or GRADE are adapted to evaluate the evidence. The RAND/UCLA Appropriateness Method User Manual provides the methodology to determine the appropriateness of imaging and treatment procedures for specific clinical scenarios. In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.
PMID: 41193048
ISSN: 1558-349x
CID: 5959922

ACR Appropriateness Criteria® Male Breast Cancer Screening

,; Freer, Phoebe E; Neal, Colleen H; Brown, Ann; Bennett, Debbie L; Cassidy, Michael R; Chetlen, Alison; Dibble, Elizabeth H; Giordano, Sharon H; Greenwood, Heather I; Hurley, Janet; Ivansco, Lillian K; Malak, Sharp F; Rauch, Gaiane M; Reig, Beatriu; Singh, Puneet; Small, William; Yeh, Eren D; Slanetz, Priscilla J
Breast cancer screening recommendations have been established historically for women, but, have been less clearly outlined for men. For average-risk men and younger men less than 25 year of age, imaging is not usually appropriate as a screening test for breast cancer. For men of higher-than-average risk, screening with mammography as annual surveillance imaging is usually appropriate. The American College of Radiology Appropriateness Criteria are evidence-based guidelines for specific clinical conditions that are reviewed annually by a multidisciplinary expert panel. The guideline development and revision process support the systematic analysis of the medical literature from peer reviewed journals. Established methodology principles such as Grading of Recommendations Assessment, Development, and Evaluation or GRADE are adapted to evaluate the evidence. The RAND/UCLA Appropriateness Method User Manual provides the methodology to determine the appropriateness of imaging and treatment procedures for specific clinical scenarios. In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.
PMID: 41193045
ISSN: 1558-349x
CID: 5959912

Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication

Mandal, Soumik; Wiesenfeld, Batia M; Szerencsy, Adam C; Small, William R; Major, Vincent; Richardson, Safiya; Schoenthaler, Antoinette; Mann, Devin; Nov, Oded
The integration of generative AI (GenAI) in patient communication presents benefits and challenges. This retrospective observational study analyzed EHR audit logs to assess how 75 healthcare professionals (HCPs) utilized AI-generated drafts for patient messages from October 2023 to August 2024 at a large health system in New York City. Overall utilization was low (19.4%), though prompt refinements improved usage (from 12% to 20%), particularly among physicians. GenAI drafts were generated for all messages, including 80% that received no response, adding to the review burden and potentially undermining efficiency. Text analysis showed HCPs preferred concise, information-rich drafts, with role-based differences-physicians favored shorter drafts, while clinical support staff preferred more empathetic responses. AI-generated drafts reduced message turnaround time by 6.76% despite a marginal increase in required steps (InBasket actions). These findings highlight the need for targeted GenAI deployment strategies, better aligned with clinician workflows and optimized draft generation for improved efficiency.
PMCID:12491571
PMID: 41038966
ISSN: 2398-6352
CID: 6072071

Evaluating Hospital Course Summarization by an Electronic Health Record-Based Large Language Model

Small, William R; Austrian, Jonathan; O'Donnell, Luke; Burk-Rafel, Jesse; Hochman, Katherine A; Goodman, Adam; Zaretsky, Jonah; Martin, Jacob; Johnson, Stephen; Major, Vincent J; Jones, Simon; Henke, Christian; Verplanke, Benjamin; Osso, Jwan; Larson, Ian; Saxena, Archana; Mednick, Aron; Simonis, Choumika; Han, Joseph; Kesari, Ravi; Wu, Xinyuan; Heery, Lauren; Desel, Tenzin; Baskharoun, Samuel; Figman, Noah; Farooq, Umar; Shah, Kunal; Jahan, Nusrat; Kim, Jeong Min; Testa, Paul; Feldman, Jonah
IMPORTANCE/UNASSIGNED:Hospital course (HC) summarization represents an increasingly onerous discharge summary component for physicians. Literature supports large language models (LLMs) for HC summarization, but whether physicians can effectively partner with electronic health record-embedded LLMs to draft HCs is unknown. OBJECTIVES/UNASSIGNED:To compare the editing effort required by time-constrained resident physicians to improve LLM- vs physician-generated HCs toward a novel 4Cs (complete, concise, cohesive, and confabulation-free) HC. DESIGN, SETTING, AND PARTICIPANTS/UNASSIGNED:Quality improvement study using a convenience sample of 10 internal medicine resident editors, 8 hospitalist evaluators, and randomly selected general medicine admissions in December 2023 lasting 4 to 8 days at New York University Langone Health. EXPOSURES/UNASSIGNED:Residents and hospitalists reviewed randomly assigned patient medical records for 10 minutes. Residents blinded to author type who edited each HC pair (physician and LLM) for quality in 3 minutes, followed by comparative ratings by attending hospitalists. MAIN OUTCOMES AND MEASURES/UNASSIGNED:Editing effort was quantified by analyzing the edits that occurred on the HC pairs after controlling for length (percentage edited) and the degree to which the original HCs' meaning was altered (semantic change). Hospitalists compared edited HC pairs with A/B testing on the 4Cs (5-point Likert scales converted to 10-point bidirectional scales). RESULTS/UNASSIGNED:Among 100 admissions, compared with physician HCs, residents edited a smaller percentage of LLM HCs (LLM mean [SD], 31.5% [16.6%] vs physicians, 44.8% [20.0%]; P < .001). Additionally, LLM HCs required less semantic change (LLM mean [SD], 2.4% [1.6%] vs physicians, 4.9% [3.5%]; P < .001). Attending physicians deemed LLM HCs to be more complete (mean [SD] difference LLM vs physicians on 10-point bidirectional scale, 3.00 [5.28]; P < .001), similarly concise (mean [SD], -1.02 [6.08]; P = .20), and cohesive (mean [SD], 0.70 [6.14]; P = .60), but with more confabulations (mean [SD], -0.98 [3.53]; P = .002). The composite scores were similar (mean [SD] difference LLM vs physician on 40-point bidirectional scale, 1.70 [14.24]; P = .46). CONCLUSIONS AND RELEVANCE/UNASSIGNED:Electronic health record-embedded LLM HCs required less editing than physician-generated HCs to approach a quality standard, resulting in HCs that were comparably or more complete, concise, and cohesive, but contained more confabulations. Despite the potential influence of artificial time constraints, this study supports the feasibility of a physician-LLM partnership for writing HCs and provides a basis for monitoring LLM HCs in clinical practice.
PMID: 40802185
ISSN: 2574-3805
CID: 5906762

Disappearing Text as a Clinical Decision Support Layer: A Case Series

Silberlust, Jared; Small, William; Shah, Darshi; Chakravartty, Eesha; Moawad, Katherine; Moawad, Andrew; Testa, Paul; Feldman, Jonah
OBJECTIVES/OBJECTIVE:This case series aims to evaluate several applications of inline disappearing text (DT) clinical decision support (CDS) tools within clinician documentation. METHODS:DT blocks were created to prompt documentation for perioperative anticoagulation planning (Scenario 1), pre-discharge intravenous antibiotic planning (Scenario 2), and advanced care planning (Scenario 3). In Scenario 1, DT was the only intervention. In Scenario 2, DT was paired with a documentation SmartList. In Scenario 3, DT was paired with a documentation SmartList and an OurPractice Advisory. The number of documented perioperative anticoagulation plans, pre-discharge intravenous antibiotic plans, and Advanced Care Planning notes were measured pre- and post-intervention and compared using Chi-square analyses. RESULTS:In Scenario 1, there was no statistically significant change in the percentage of perioperative anticoagulation plans documented at 0-24 and 24-48 hours before surgery. In Scenario 2, documentation of antibiotic contingency planning in patients expected to be discharged within 24 hours increased from 60% (54 of 90 notes) to 93% (1,850 of 1,994 notes) X2 (1, N=2,084) = 113.1, p < 0.001. In Scenario 3, ACP note documentation by discharge in patients with a positive mandatory surprise question increased from 43% (821 of 1,909 encounters) to 52% (975 of 1,874 encounters) X2 (1, N=3,783) = 30.5, p < 0.001. CONCLUSIONS:Utilizing DT in conjunction with other forms of CDS was associated with an improvement of documentation quality in pre-discharge IV antibiotics and advanced care planning. A sociotechnical analysis explores how interactions between technology, people, workflow, and culture could contextualize how utilizing DT with other forms of CDS was more effective than DT alone.
PMID: 40763805
ISSN: 1869-0327
CID: 5905032