Try a new search

Format these results:

Searched for:

in-biosketch:true

person:oermae01

Total Results:

163


Reply to: Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison [Letter]

Vishwanath, Krithik; Aphinyanaphongs, Yindalon; Oermann, Eric Karl
PMID: 42693298
ISSN: 1546-170x
CID: 6072018

Spinal meningiomas: histopathological grading using a benchmark radiomics model with notes on disease control

Palla, Adhith; Goff, Nicolas K; Perdikis, Blake; Khan, Hammad A; Belakhoua, Sarra; Grin, Eric A; Valliani, Aly; Patel, Roshni; Yang, Jonathan T; McFaline-Figueroa, J Ricardo; Lau, Darryl; Frempong-Boadu, Anthony; Oermann, Eric K; Laufer, Ilya
OBJECTIVE:Spinal meningiomas (SMs) are common primary spinal tumors for which surgery is considered the first-line treatment when safe and feasible. The ability to extrapolate the tumor grade from preoperative imaging may significantly inform early patient expectation-setting regarding recurrence. Building on radiomics studies in cranial meningiomas, the authors aimed to construct a benchmark radiomics model to preoperatively identify the histological grade of SMs. METHODS:Institutional surgical records from May 2012 to November 2025 were queried for pathology-confirmed meningiomas below the foramen magnum, with preoperative contrast-enhanced imaging available for segmentation. SMs were classified as low-grade (WHO grade 1) and high-grade (WHO grade 2 tumors and grade 1 tumors with atypia). Tumors were manually segmented, and features were extracted using the PyRadiomics software package. An ensemble model of k-nearest neighbors, random forest, and support vector machine classifiers was trained using nested cross-validation on a subset of 10 features to differentiate tumor grades. Clinical data for the cohort were also extracted, and disease control in an adjunctive clinical series was assessed. RESULTS:Seventy-four patients were included in radiomics analysis, with an area under the receiver operating characteristic curve of 0.879 and a mean F1 score of 0.748. The model's top 5 features were all texture features that differed significantly (p < 0.05) across low- and high-grade SMs. These included measures of tumor textural and contrast-enhancement heterogeneity, with overlap with features reported in radiomics models for histological grading of intracranial meningiomas. Fifty-five patients with a median radiographic follow-up of 22.2 (range 1.9-86.4) months remained for clinical analysis after exclusion of patients with less than 1 month of follow-up and syndromic meningiomas. Four recurrences occurred at a median of 20.8 (range 1.8-41.8) months. High-grade tumor pathology did not significantly impact progression-free survival (p = 0.682, log-rank test; Cox regression high vs low grade hazard ratio [HR] 0.62, 95% CI 0.06-6.11, p = 0.685). Subtotal resection was associated with poorer progression-free survival than gross-total resection (p = 0.004, log-rank test; Cox regression subtotal vs gross-total resection HR 10.62, 95% CI 1.46-77.05, p = 0.019). These findings remain contextualized within a relatively limited follow-up window and small recurrence event count, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs. CONCLUSIONS:A preoperative radiomics model can stratify high-grade SMs using open-source tools applied to single-institution data.
PMID: 42679405
ISSN: 1092-0684
CID: 6071950

Longitudinal alignments and syntheses of multimodal clinical data for personalized medicine with the PULSE framework

Wu, Wei; Li, Gen; Wang, Kai; Xu, Hui; Xiao, Haodi; Hu, Changxi; Liu, Sian; Tang, Cheng; Liu, Fei; Zou, Zixing; Li, Bingzhou; Li, Jinghang; Zhang, Charlotte L; Wong, Hang; Chong, Ieng; Lu, Wenyang; Sun, Zhuo; Yin, Yun; Loupy, Alexandre; Oermann, Eric; Al Dajani, Saleem A; Zhu, Hao; Gootenberg, Jonathan; Abudayyeh, Omar O; Gladyshev, Vadim N; Rasko, John E J; Zhang, Kang; ,
Multimodal models capable of imputing diverse data used in single-cell biology studies provide potential foundational opportunities in clinical practice. However, patient data uniquely comprise longitudinal mosaic measurements that reflect underlying physiological dynamics and exhibit temporal covariation, demanding a specialized approach. Here we present Patient Unified Longitudinal Signal Engine (PULSE), a longitudinal self-supervised framework that explicitly encodes personalized past states (historical paired modalities) to reconstruct full profiles from subsequent unpaired measurements, thus enhancing current visit multimodal alignment and generation. Applied to the UK Biobank, PULSE accurately generates metabolomic profiles and proteomic profiles from sparse routine blood tests. Compared with the ground truth metabolomic data (251 biomarkers), PULSE-generated profiles outperformed all benchmark methods. Furthermore, the framework accommodates incorporation of retinal images, electronic health records and blood markers with disease prediction: models trained on the generated proteomic profiles achieved areas under the curve of 0.72-0.83 for six common diseases, comparable to that using ground-truth proteomic data. The PULSE framework demonstrates that cross-modal alignment captures the continuous spectrum of disease physiology and extracts robust features that transcend the limitations of traditional binary case-controls.
PMID: 42649447
ISSN: 2662-8457
CID: 6071810

Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions

Alyakin, Anton; Stryker, Jaden; Alber, Daniel Alexander; Lee, Jin Vivian; Singh, Shrutika; Save, Akshay; Kurland, David; Orillac, Cordelia; Valliani, Aly A; Neifert, Sean; Lau, Darryl; Laufer, Ilya; Rozman, Peter A; Hidalgo, Eveline Teresa; Riina, Howard; Leuthardt, Eric C; Kondziolka, Douglas; Snyder, Laura; Oermann, Eric Karl
BACKGROUND AND OBJECTIVES/OBJECTIVE:Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions. METHODS:articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian. RESULTS:). CONCLUSION/CONCLUSIONS:Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.
PMCID:13391137
PMID: 42488579
ISSN: 2834-4383
CID: 6071663

Big data in U.S. neuro-oncology: trends and translational priorities

Kapoor, Anjali; Alyakin, Anton; Markert, John E; Arias, Ari; Yang, Eunice; Vishwanath, Krithik; Lee, Jin Vivian; Sughrue, Michael; Oermann, Eric Karl
PURPOSE/OBJECTIVE:Neuro-oncology generates complex clinical, imaging, and molecular data, yet datasets remain relatively small and fragmented across modalities and institutions. While "big data" is traditionally defined by large sample size, neuro-oncology datasets are often characterized instead by high dimensionality. This study aims to provide an overview of the landscape of major U.S. neuro-oncology data resources and evaluate how these datasets are used in contemporary research. METHODS:A selection of neuro-oncology datasets was evaluated, including population registries, clinical data networks, federal and consortium research cohorts, institutional datasets, specialized resources, and artificial intelligence benchmarking resources. Analytical use was assessed through a large language model-assisted review of PubMed-indexed studies published over the past ten years referencing these datasets. Titles and abstracts were screened using a predefined classification schema, and structured data extraction identified study characteristics, analytical tasks, outcomes, modalities, validation strategies, and longitudinal modeling approaches. RESULTS:Of 11,651 screened publications, 3,608 met inclusion criteria. Analytical use was concentrated in a small number of datasets, particularly TCGA (~ 65%), SEER (~ 14%), and BraTS (~ 13%). Most studies modeled survival or tumor characteristics, whereas fewer than 1% examined functional or quality-of-life outcomes. Approximately 90% relied on a single dataset, and external validation and longitudinal modeling were rare. CONCLUSION/CONCLUSIONS:Big data in neuro-oncology is characterized by rich diversity. Expanding multimodal data capture, improving coverage of underrepresented populations and tumor types, strengthening longitudinal data collection, and enabling cross-dataset integration will be essential for translating high-dimensional datasets into clinically actionable insights.
PMID: 42611103
ISSN: 1573-7373
CID: 6071441

Advancing cancer detection and treatment using longitudinal routine clinical data

Liu, Fei; Wang, Kai; Xu, Hui; Tang, Cheng; Shen, Xian; Wang, Meihao; Yang, Lei; Yang, Li; Liu, Li; Hu, Changxi; Li, Gen; Wu, Wei; Zou, Zixing; Li, Bingzhou; Liu, Sian; Kang, Jin; Kong, Jungho; Li, Ting; Wong, Io Nam; Huang, Xiaoying; Chen, Gang; Lu, Wenyang; Ziyar, Ian; Zhang, Charlotte L; Sun, Yiwen; Lin, Weihong; Ou, Caiwen; Fok, Manson; Hou, Taiwa; Wang, Winston; Xue, Kanmin; Yin, Yun; Zhu, Hao; Gootenberg, Jonathan; Abudayyeh, Omar O; Karin, Michael; Loupy, Alexandre; Rasko, John E J; Ideker, Trey; Luo, Huiyan; Oermann, Eric; Zhang, Kang; ,
Cancer management remains fragmented across its continuum, from late-stage diagnosis and salvage therapies to non-personalized surveillance. Here, we present Oncoformer, a unified multimodal transformer model trained on the China Oncology Multimodal Prediction and Surveillance Study (COMPASS) cohort (3.67 million individuals, 17.7 million clinical visits) and validated on independent external cohorts, including the UK Biobank. Oncoformer integrates longitudinal electronic health records with chest X-ray imaging to address multiple clinical tasks: pan-cancer diagnosis (area under the receiver operating characteristic curve [AUROC] = 0.956), future cancer prediction up to 1 year before diagnosis (AUROC = 0.869), tumor stage inference (mean AUROC > 0.90), patient-specific treatment-response forecasting, and recurrence-free survival stratification across ten cancer types (all p < 0.01). Staging predictions were independently validated against postoperative pathological endpoints and shown to converge on core cancer genomic pathways. By translating routine clinical data into a dynamic view of cancer evolution, Oncoformer provides a framework for risk-informed cancer prediction and treatment stratification using routine clinical data.
PMID: 42508404
ISSN: 1097-4172
CID: 6070394

A Radiomics-Driven Model to Distinguish Between Clinically Similar Myxopapillary Ependymomas and Lumbosacral Schwannomas

Palla, Adhith; Goff, Nicolas K; Perdikis, Blake; Khan, Hammad A; Grin, Eric A; Valliani, Aly; Patel, Roshni; Yang, Jonathan T; McFaline-Figueroa, J Ricardo; Lau, Darryl; Frempong-Boadu, Anthony; Oermann, Eric K; Laufer, Ilya
BACKGROUND AND OBJECTIVES/OBJECTIVE:Myxopapillary ependymomas (MPE) and intradural lumbosacral schwannomas may be challenging to distinguish based on presenting characteristics and preoperative imaging. Accurate differentiation is crucial, as MPEs carry a risk of cerebrospinal fluid dissemination and warrant earlier intervention, a more tailored surgical strategy, consideration for adjuvant radiation, and frequent surveillance. Here, we describe our institutional experience with these tumors and develop a radiomics-based machine learning model to help distinguish them on preoperative imaging. METHODS:Institutional surgical records from 2011 to 2025 were queried and clinical data were extracted for the retrospective cohort analysis. Tumors were manually segmented in ITK-Snap from T1 postcontrast images, and radiomics features were extracted using the PyRadiomics package. An ensemble of random forest, k-nearest neighbors, and naive Bayes classifiers was trained on a subset of radiomics features using nested cross-validation. RESULTS:< .001) in MPEs, likely due to longitudinal tumor growth along the filum. Excluding scoliotic patients did not significantly alter discrimination, suggesting robustness to vertebral column malalignment that may coexist with intradural tumors. CONCLUSION/CONCLUSIONS:A radiomics-based machine learning model demonstrated excellent discriminative ability between MPE and lumbosacral schwannoma, achieving high accuracy and robustness to vertebral alignment variations. These results suggest that radiomics-based models may be developed into a useful tool for preoperative planning and patient counseling.
PMCID:13354379
PMID: 42434191
ISSN: 2834-4383
CID: 6064422

Trends in Patient Portal Messages, Office Visits, and Telephone Encounters

Long, Jane J; McAdams-DeMarco, Mara A; Schwartz, Mark D; Chodosh, Joshua; Oermann, Eric K; Segev, Dorry L; Mankowski, Michal A
PMID: 42329625
ISSN: 1538-3598
CID: 6055282

General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

Vishwanath, Krithik; Alyakin, Anton; Ghosh, Mrigayu; Hage, Ali; Neifert, Sean N; Orillac, Cordelia; Mandelberg, Nataniel J; Khan, Hammad A; Lee, Jin Vivian; Yao, Jie J; Small, William Robert; Varma, Aakaash; Hewitt, D Brock; Aphinyanaphongs, Yindalon; Alber, Daniel Alexander; Oermann, Eric Karl
Specialized clinical artificial intelligence (AI) tools are entering medical practice despite scarce independent evaluation. We quantitatively evaluate two clinical AI tools, OpenEvidence and UpToDate Expert AI, built on large language models (LLMs) against three frontier LLMs: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-identified queries from physicians to a general-purpose language model in a live clinical environment. For the RCQ benchmark, 12 US clinicians performed randomized, blinded review of model outputs, producing 1,800 model-question annotations. Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ. These findings highlight the need for independent, real-world evaluation of AI tools before they enter clinical settings.
PMID: 42286322
ISSN: 1546-170x
CID: 6049082

CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications

Alyakin, Anton; Stryker, Jaden; Alber, Daniel Alexander; Lee, Jin Vivian; Sangwon, Karl L; Duderstadt, Brandon; Save, Akshay; Kurland, David; Frome, Spencer; Singh, Shrutika; Zhang, Jeff; Yang, Eunice; Park, Ki Yun; Orillac, Cordelia; Valliani, Aly A; Neifert, Sean; Liu, Albert; Patel, Aneek; Livia, Christopher; Lau, Darryl; Laufer, Ilya; Rozman, Peter A; Hidalgo, Eveline Teresa; Riina, Howard; Feng, Rui; Hollon, Todd; Aphinyanaphongs, Yindalon; Golfinos, John G; Snyder, Laura; Leuthardt, Eric C; Kondziolka, Douglas; Oermann, Eric Karl
BACKGROUND AND OBJECTIVES/OBJECTIVE:General purpose vision-language models (VLMs) demonstrate impressive capabilities, but their opaque training on uncurated internet data poses critical limitations for high-stakes decision making, such as in neurosurgery. We present CNS-Obsidian, a neurosurgical VLM trained on peer-reviewed neurosurgical literature, and demonstrate its clinical utility compared with GPT-4o in a real-world setting. METHODS:We compiled 23 984 articles from Neurosurgery Publications journals, yielding 78 853 figures and captions. Using GPT-4o and Claude Sonnet-3.5, we converted these image-text pairs into 263 064 training samples across 3 formats: instruction fine-tuning, multiple-choice questions, and differential diagnosis. We trained CNS-Obsidian, a fine-tune of the 34-billion parameter Large Language and Visual Assistant-Next model. In a blinded, randomized deployment trial at NYU Langone Health (August 30-November 30, 2024), neurosurgeons were assigned to use either CNS-Obsidian or a Health Insurance Portability and Accountability Act-compliant GPT-4o end point as a diagnostic copilot after patient consultations. Primary outcomes were diagnostic helpfulness and accuracy, assessed through user ratings and presence of the correct diagnosis within the VLM-provided differential, respectively. RESULTS:CNS-Obsidian matched GPT-4o on synthetic questions (76.13% vs 77.54%, P = .235), but only achieved 46.81% accuracy on human-generated questions vs GPT-4o's 65.70% (P < 10-15). In the randomized trial, 70 consultations were evaluated (32 CNS-Obsidian, 38 GPT-4o) from 959 total consults (7.3% utilization). CNS-Obsidian received positive ratings in 40.62% of cases vs 57.89% for GPT-4o (P = .230). Both models included correct diagnosis in approximately 60% of cases (59.38% vs 65.79%, P = .626). CONCLUSION/CONCLUSIONS:Domain-specific VLMs trained on curated scientific literature can approach frontier model performance in specialized medical domains despite being orders of magnitude smaller and less expensive to train. This establishes a transparent framework for scientific communities to build specialized artificial intelligence models. However, low clinical utilization suggests chatbot interfaces may not align with specialist workflows, indicating need for alternative artificial intelligence integration strategies.
PMID: 42153721
ISSN: 1524-4040
CID: 6037862