Searched for: in-biosketch:true
person:oermae01
Most Roads Lead to Cushing: Mapping Neurosurgical Training Lineages in the United States
Kurland, David B; Park, Minjun; Gajjar, Avi A; Liu, Albert; Kondziolka, Douglas; Golfinos, John G; Alleyne, Cargill H; Oermann, Eric K
OBJECTIVE:Mentorship and training relationships shape the careers and influence of neurosurgeons. Network analysis can reveal structural characteristics and key individuals who support network connectivity and drive the field's development. This endeavor analyzed the U.S.-based neurosurgical training network derived from NeurosurGen.com. METHODS:A network graph was constructed representing neurosurgical training relationships, including chairperson-trainee, program director-trainee, and coresident connections. Graph- and node-level metrics, with a focus on centrality measures, were calculated for a trainer-trainee subgraph. RESULTS:The network consisted of 8840 neurosurgeons represented as nodes, and 382,143 relationships represented as edges. It evolved from an early small-world structure to a hierarchical and decentralized structure dominated by local clusters. Demographic shifts over time reflected increasing diversity and inclusion, with greater representation of female, Hispanic, Asian, and Black trainees across 285 training programs. Nodes were preferentially connected via residency, and the connectivity among underrepresented populations improved in concert with increased representation. Harvey W. Cushing was the quintessential neurosurgeon-influencer in the United States, ranking highly across most centrality measures over time. CONCLUSIONS:The neurosurgical training network is sparse but interconnected, typical of large real-world professional networks. While many small groups of neurosurgeons are closely tied within their immediate training hierarchy and peer group, in modern neurosurgery, each surgeon is only connected to a small fraction of the total network. Highly central individuals have played critical roles in linking disparate groups and shaping network structure. Increasing diversity in recent decades indicates progress toward inclusivity, although overall representation remains low.
PMID: 40914191
ISSN: 1878-8769
CID: 5966272
A full life cycle biological clock based on routine clinical data and its impact in health and diseases
Wang, Kai; Liu, Fei; Wu, Wei; Hu, Changxi; Shen, Xian; Wang, Meihao; Li, Gen; Zeng, Fanxin; Liu, Li; Wong, Io Nam; Liu, Sian; Zou, Zixing; Li, Bingzhou; Li, Jinghang; Huang, Xiaoying; Jin, Shengwei; Li, Zhuomin; Xu, Hui; Chen, Gang; Chen, Xiaodong; Zhu, Ying; Li, Ping; Feng, Zhe; Wang, Winston; Cheng, Linling; Yang, Mingqi; Hou, Qiang; Lu, Wenyang; Sun, Yiwen; Li, Kun; Zhong, Tian; Sun, Zhuo; Yin, Yun; Loupy, Alexandre; Oermann, Eric; Chen, Xiangmei; Zhang, Kang; ,
Aging research has primarily focused on adult aging clocks, leaving a critical gap in understanding a biological clock across the full life cycle, particularly during infancy and childhood. Here we introduce LifeClock, a biological clock model that predicts biological age across all life stages using routine electronic health records and laboratory test data. To enhance individualized predictions, we integrated virtual patient representations from 24,633,025 heterogeneous longitudinal clinical visits across 9,680,764 individuals and projected them into a latent space. Our approach leverages EHRFormer, a time-series transformer-based model, to analyze developmental and aging dynamics with high precision and develop accurate biological age clocks spanning infancy to old age. Our findings reveal distinct biological clock patterns across different life stages. The pediatric clock is strongly associated with children's development and accurately predicts current and future risks of major pediatric diseases, including malnutrition, growth and developmental abnormalities. The adult clock is strongly associated with aging and accurately predicts current and future risks of major age-related diseases, such as diabetes, renal failure, stroke and cardiovascular diseases. This work therefore distinguishes pediatric development from adult aging, establishing a novel framework to advance precision health by leveraging routine clinical data across the entire lifespan.
PMID: 41145791
ISSN: 1546-170x
CID: 5961022
Neuro Data Hub: A New Approach for Streamlining Medical Clinical Research
Han, Xu; Alyakin, Anton; Ciprut, Shannon; Lapierre, Cathryn; Stryker, Jaden; Golfinos, John; Kondziolka, Douglas; Oermann, Eric Karl
BACKGROUND AND OBJECTIVES/OBJECTIVE:Neurosurgical clinical research depends on medical data collection and evaluation that is often laborious, time consuming, and inefficient. The goal of this work was to implement and evaluate a novel departmental data infrastructure (Neuro Data Hub) designed to provide specialized data services for neurosurgical research. Data acquisition would become available purely by request. METHODS:through collaboration between Department Leadership and Medical Center Information Technology, integrating it with Institutional Review Board workflows and an existing Epic electronic health record Datalake infrastructure. The system implementation included monthly departmental meetings and an asynchronous Research Electronic Data Capture-based request system. Data requests submitted between August 2023 and November 2024 were analyzed and categorized as basic, complex, or Natural Language Processing (NLP)-augmented, with optional visualization and database creation services. Request volumes, types, and execution times were assessed. RESULTS:The Hub processed 39 research data requests (2.6/month), comprising 3 basic, 22 complex, and 14 NLP-augmented requests. Two complex requests included visualization services, and one NLP request included database creation. Average request execution time was 36.5 days, with NLP-augmented requests showing increasing adoption over time. CONCLUSION/CONCLUSIONS:The Neuro Data Hub represents a paradigm shift from centralized to department-level data services, providing specialized support for neurosurgical research and democratizing access to institutional data. While effective, implementation may be limited by institutional information technology infrastructure requirements. This model could serve as a template for any form of medical-clinical research program seeking to improve data accessibility and research capabilities.
PMCID:12560744
PMID: 41163737
ISSN: 2834-4383
CID: 5961452
Augmenting Large Language Models With Automated, Bibliometrics-Powered Literature Search for Knowledge Distillation: A Pilot Study for Common Spinal Pathologies
Kurland, David B; Alber, Daniel A; Palla, Adhith; de Souza, Daniel N; Lau, Darryl; Laufer, Ilya; Frempong-Boadu, Anthony K; Kondziolka, Douglas; Oermann, Eric K
BACKGROUND AND OBJECTIVES/OBJECTIVE:Scholarly output is accelerating in medical domains, making it challenging to keep up with the latest neurosurgical literature. The emergence of large language models (LLMs) has facilitated rapid, high-quality text summarization. However, LLMs cannot autonomously conduct literature reviews and are prone to hallucinating source material. We devised a novel strategy that combines Reference Publication Year Spectroscopy-a bibliometric technique for identifying foundational articles within a corpus-with LLMs to automatically summarize and cite salient details from articles. We demonstrate our approach for four common spinal conditions in a proof of concept. METHODS:Reference Publication Year Spectroscopy identified seminal articles from the corpora of literature for cervical myelopathy, lumbar radiculopathy, lumbar stenosis, and adjacent segment disease. The article text was split into 1024-token chunks. Queries from three knowledge domains (surgical management, pathophysiology, and natural history) were constructed. The most relevant article chunks for each query were retrieved from a vector database using chain-of-thought prompting. LLMs automatically summarized the literature into a comprehensive narrative with fully referenced facts and statistics. Information was verified through manual review, and spine surgery faculty were surveyed for qualitative feedback. RESULTS:Our tandem approach cost less than $1 for each condition and ran within 5 minutes. Generative Pre-trained Transformer-4 was the best-performing model, with a near-perfect 97.5% citation accuracy. Surveys of spine faculty helped refine the prompting scheme to improve the cohesion and accessibility summaries. The final artificial intelligence-generated text provided high-fidelity summaries of each pathology's most clinically relevant information. CONCLUSION/CONCLUSIONS:We demonstrate the rapid, automated summarization of seminal articles for four common spinal pathologies, with a generalizable workflow implemented using consumer-grade hardware. Our tandem strategy fuses bibliometrics and artificial intelligence to bridge the gap toward fully automated knowledge distillation, obviating the need for manual literature review and article selection.
PMID: 40662770
ISSN: 1524-4040
CID: 5897082
Introduction. Artificial intelligence in neurosurgery: transforming a data-intensive specialty
Hopkins, Benjamin S; Sutherland, Garnette R; Browd, Samuel R; Donoho, Daniel A; Oermann, Eric K; Schirmer, Clemens M; Pennicooke, Brenton; Asaad, Wael F
PMID: 40591964
ISSN: 1092-0684
CID: 5887762
Is It Really "Artificial" Intelligence?
Kondziolka, Douglas; Oermann, Eric K
PMID: 39812480
ISSN: 1524-4040
CID: 5883422
Large-Scale Multi-omic Biosequence Transformers for Modeling Protein-Nucleic Acid Interactions
Chen, Sully F; Steele, Robert J; Hocky, Glen M; Lemeneh, Beakal; Lad, Shivanand P; Oermann, Eric K
The transformer architecture has revolutionized bioinformatics and driven progress in the understanding and prediction of the properties of biomolecules. To date, most biosequence transformers have been trained on a single omic-either proteins or nucleic acids and have seen incredible success in downstream tasks in each domain with particularly noteworthy breakthroughs in protein structural modeling. However, single-omic pre-training limits the ability of these models to capture cross-modal interactions. Here we present OmniBioTE, the largest open-source multi-omic model trained on over 250 billion tokens of mixed protein and nucleic acid data. We show that despite only being trained on unlabelled sequence data, OmniBioTE learns joint representations consistent with the central dogma of molecular biology. We further demonstrate that OmbiBioTE achieves state-of-the-art results predicting the change in Gibbs free energy (∆G) of the binding interaction between a given nucleic acid and protein. Remarkably, we show that multi-omic biosequence transformers emergently learn useful structural information without any a priori structural training, allowing us to predict which protein residues are most involved in the protein-nucleic acid binding interaction. Lastly, compared to single-omic controls trained with identical compute, OmniBioTE demonstrates superior performance-per-FLOP and absolute accuracy across both multi-omic and single-omic benchmarks, highlighting the power of a unified modeling approach for biological sequences.
PMCID:11998858
PMID: 40236839
ISSN: 2331-8422
CID: 5883432
MetaGP: A generative foundation model integrating electronic health records and multimodal imaging for addressing unmet clinical needs
Liu, Fei; Zhou, Hongyu; Wang, Kai; Yu, Yunfang; Gao, Yuanxu; Sun, Zhuo; Liu, Sian; Sun, Shanshan; Zou, Zixing; Li, Zhuomin; Li, Bingzhou; Miao, Hanpei; Liu, Yang; Hou, Taiwa; Fok, Manson; Patil, Nivritti Gajanan; Xue, Kanmin; Li, Ting; Oermann, Eric; Yin, Yun; Duan, Lian; Qu, Jia; Huang, Xiaoying; Jin, Shengwei; Zhang, Kang
Artificial intelligence makes strides in specialized diagnostics but faces challenges in complex clinical scenarios, such as rare disease diagnosis and emergency condition identification. To address these limitations, we develop Meta General Practitioner (MetaGP), a 32-billion-parameter generative foundation model trained on extensive datasets, including over 8 million electronic health records, biomedical literature, and medical textbooks. MetaGP demonstrates robust diagnostic capabilities, achieving accuracy comparable to experienced clinicians. In rare disease cases, it achieves an average diagnostic score of 1.57, surpassing GPT-4's 0.93. For emergency conditions, it improves diagnostic accuracy for junior and mid-level clinicians by 53% and 46%, respectively. MetaGP also excels in generating medical imaging reports, producing high-quality outputs for chest X-rays and computed tomography, often rated comparable to or superior to physician-authored reports. These findings highlight MetaGP's potential to transform clinical decision-making across diverse medical contexts.
PMID: 40187356
ISSN: 2666-3791
CID: 5819502
Outcomes of concurrent versus non-concurrent immune checkpoint inhibition with stereotactic radiosurgery for melanoma brain metastases
Fu, Allen Ye; Bernstein, Kenneth; Zhang, Jeff; Silverman, Joshua; Mehnert, Janice; Sulman, Erik P; Oermann, Eric Karl; Kondziolka, Douglas
PURPOSE/OBJECTIVE:Immune checkpoint inhibition (ICI) has revolutionized the treatment of melanoma care. Stereotactic radiosurgery combined with ICI has shown promise to improve clinical outcomes in prior studies in patients who have metastatic melanoma with brain metastases. However, others have suggested that concurrent ICI with stereotactic radiosurgery can increase the risk of complications. METHODS:We present a retrospective, single-institution analysis of 98 patients with a median follow up of 17.1 months managed with immune checkpoint inhibition and stereotactic radiosurgery concurrently and non-concurrently. A total of 55 patients were included in the concurrent group and 43 patients in the non-concurrent treatment group. Cox proportional hazards models were used to assess the relation between concurrent or non-concurrent treatment and overall survival or local progression-free survival. The Wald test was used to assess significance. Significant differences between patients in both groups experiencing adverse events including adverse radiation effects, perilesional edema, and neurological deficits were tested for using the Chi-square or Fisher's exact test. RESULTS:Patients receiving concurrent versus non-concurrent ICI showed a significant increase in overall survival (median 37.1 months, 95% CI: 18.9 months - NA versus median 11.4 months, 95% CI: 6.4-33.2 months, p = 0.0056) but not local progression-free survival. There were no significant differences between groups with regards to adverse radiation effects (2% versus 3%), perilesional edema (20% versus 9%), neurological deficits (3% versus 20%). CONCLUSION/CONCLUSIONS:These results suggest that the timing of ICI does not increase risk of neurological complications when delivered within 4 weeks of SRS.
PMID: 40183901
ISSN: 1573-7373
CID: 5819412
Medical large language models are vulnerable to data-poisoning attacks
Alber, Daniel Alexander; Yang, Zihao; Alyakin, Anton; Yang, Eunice; Rai, Sumedha; Valliani, Aly A; Zhang, Jeff; Rosenbaum, Gabriel R; Amend-Thomas, Ashley K; Kurland, David B; Kremer, Caroline M; Eremiev, Alexander; Negash, Bruck; Wiggan, Daniel D; Nakatsuka, Michelle A; Sangwon, Karl L; Neifert, Sean N; Khan, Hammad A; Save, Akshay Vinod; Palla, Adhith; Grin, Eric A; Hedman, Monika; Nasir-Moin, Mustafa; Liu, Xujin Chris; Jiang, Lavender Yao; Mankowski, Michal A; Segev, Dorry L; Aphinyanaphongs, Yindalon; Riina, Howard A; Golfinos, John G; Orringer, Daniel A; Kondziolka, Douglas; Oermann, Eric Karl
The adoption of large language models (LLMs) in healthcare demands a careful analysis of their potential to spread false medical knowledge. Because LLMs ingest massive volumes of data from the open Internet during training, they are potentially exposed to unverified medical knowledge that may include deliberately planted misinformation. Here, we perform a threat assessment that simulates a data-poisoning attack against The Pile, a popular dataset used for LLM development. We find that replacement of just 0.001% of training tokens with medical misinformation results in harmful models more likely to propagate medical errors. Furthermore, we discover that corrupted models match the performance of their corruption-free counterparts on open-source benchmarks routinely used to evaluate medical LLMs. Using biomedical knowledge graphs to screen medical LLM outputs, we propose a harm mitigation strategy that captures 91.9% of harmful content (F1 = 85.7%). Our algorithm provides a unique method to validate stochastically generated LLM outputs against hard-coded relationships in knowledge graphs. In view of current calls for improved data provenance and transparent LLM development, we hope to raise awareness of emergent risks from LLMs trained indiscriminately on web-scraped data, particularly in healthcare where misinformation can potentially compromise patient safety.
PMID: 39779928
ISSN: 1546-170x
CID: 5782182