Try a new search

Format these results:

Searched for:

in-biosketch:true

person:laud06

Total Results:

154


Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions

Alyakin, Anton; Stryker, Jaden; Alber, Daniel Alexander; Lee, Jin Vivian; Singh, Shrutika; Save, Akshay; Kurland, David; Orillac, Cordelia; Valliani, Aly A; Neifert, Sean; Lau, Darryl; Laufer, Ilya; Rozman, Peter A; Hidalgo, Eveline Teresa; Riina, Howard; Leuthardt, Eric C; Kondziolka, Douglas; Snyder, Laura; Oermann, Eric Karl
BACKGROUND AND OBJECTIVES/OBJECTIVE:Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions. METHODS:articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian. RESULTS:). CONCLUSION/CONCLUSIONS:Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.
PMCID:13391137
PMID: 42488579
ISSN: 2834-4383
CID: 6071664

Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions

Alyakin, Anton; Stryker, Jaden; Alber, Daniel Alexander; Lee, Jin Vivian; Singh, Shrutika; Save, Akshay; Kurland, David; Orillac, Cordelia; Valliani, Aly A; Neifert, Sean; Lau, Darryl; Laufer, Ilya; Rozman, Peter A; Hidalgo, Eveline Teresa; Riina, Howard; Leuthardt, Eric C; Kondziolka, Douglas; Snyder, Laura; Oermann, Eric Karl
BACKGROUND AND OBJECTIVES/OBJECTIVE:Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions. METHODS:articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian. RESULTS:). CONCLUSION/CONCLUSIONS:Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.
PMCID:13391137
PMID: 42488579
ISSN: 2834-4383
CID: 6071663

A Tale of 2 Institutions: Differences in Sociodemographic Factors, Presenting Features, Treatment Characteristics, and Outcomes in Patients Undergoing Surgery for Spinal Metastases

Khan, Hammad A; Palla, Adhith; Ashayeri, Kimberly; McLaughlin, Lily; Kurland, David B; Frempong-Boadu, Anthony; Lau, Darryl; Laufer, Ilya; Pacione, Donato
BACKGROUND AND OBJECTIVES/OBJECTIVE:The objective of this study was to compare sociodemographic factors, presenting characteristics, and outcomes between cohorts of patients receiving surgery for spinal metastases at 2 neighboring institutions, 1 private and 1 public, affiliated with a single major academic medical center in a large metropolitan area. METHODS:This analysis included all patients who underwent decompressive surgery for extradural spinal metastases. Sociodemographic factors, treatment characteristics, and outcomes were compared between those treated at a private hospital and a neighboring public hospital using Rao-Scott χ 2 tests for categorical variables, Student t tests for continuous variables, and the Kaplan-Meier product-limit method for overall survival and progression-free survival. RESULTS:Compared with those treated at our private hospital, patients treated at our public hospital were more often younger ( P = .005), of Black or Hispanic race (72.6% vs 19%, P < .001), and uninsured (16% vs 5.6%, P = .005). They more frequently presented with epidural spinal cord compression grade 3 (76% vs 56.8%, P = .027), were nonambulatory before surgery (56.9% vs 13.5%, P < .001), and had increased neurological impairment as denoted by American Spinal Injury Association Impairment Scale grades of A, B, or C (39.2% vs 7.5%). Patients treated at our public hospital had shorter median follow-up time (92 vs 302.5 days, P = .004). Multivariate analysis did not reveal a significant difference in overall survival or progression-free survival between hospitals, instead demonstrating associations with primary tumor histology and number of spinal metastases ( P < .05). CONCLUSION/CONCLUSIONS:There were substantial disparities in sociodemographic factors, presenting local disease burden, and postoperative neurological outcome but no difference in survival outcome, between patients treated at our public and private hospitals. These findings underscore the need for more equitable screening, surveillance, and referral structures.
PMID: 42484343
ISSN: 1524-4040
CID: 6071637

Extension of Fusion to the Cervical Spine Versus Upper Thoracic Spine for the Management of Proximal Junctional Kyphosis of Thoracolumbar Fusion

Sulieman, Ahmed; Sahhar, Maxwell; Parekh, Yesha H; Lafage, Virginie; Lafage, Renaud; Line, Breton G; Ames, Christopher P; Bess, Shay; Buell, Thomas J; Eastlack, Robert K; Gum, Jeffrey L; Gupta, Munish C; Hostin, Richard A; Kim, Han Jo; Lau, Darryl; Mundis, Gregory M; Passias, Peter G; Protopsaltis, Themistocles S; Shaffrey, Christopher I; Smith, Justin S; Kebaish, Khaled M; Lee, Sang Hun; ,
STUDY DESIGN/METHODS:Retrospective review of multicenter, prospective cervical deformity database. OBJECTIVE:To compare outcomes of extension of fusion to the cervical spine versus the upper thoracic (UT) spine. SUMMARY OF BACKGROUND DATA/BACKGROUND:Proximal junctional kyphosis (PJK) management after thoracolumbar fusion requires extension of fusion to the proximal spinal segments. Unlike extensions to the less mobile thoracic segments, crossing the cervicothoracic junction (CTJ) involves more mobile cervical segments and creates different biomechanical influences and clinical outcomes. No study has compared the outcomes of extending fusion to the cervical versus the UT spine. METHODS:Patients with thoracic PJK who underwent revision with extension of fusion to either the cervical or UT (T1 or T2) spine were identified in a multicenter, prospective cervical deformity database. Patients with cervical upper instrumented vertebra (UIV) were subdivided into lower cervical (LC; C4-7) and upper cervical (UC; and occiput-C3) groups. Baseline demographics, surgical variables, radiographic outcomes, 2-year health-related quality-of-life scores, complications, and revision rates were analyzed. RESULTS:Fifty-one patients (mean age: 60.4±12.9 y; 91% female) with at least 2 years of follow-up were included. Twelve had extension to the UT, 20 to the LC, and 19 to the UC spine. Demographic data, Charlson Comorbidity Index, follow-up duration, surgical parameters, radiographic measurements, recurrent PJK and reoperation rates, and 2-year patient-reported outcome scores were similar across groups. The instrumentation failure rate was higher in the LC (25%) than in the UT (0%) and UC (8%) groups (P=0.03). CONCLUSIONS:Stopping fusion at T1 or T2 did not result in greater complication or reoperation rates than extending to the cervical spine. The instrumentation-related complication rate was higher for extension to the LC than to the UC or UT spine. Crossing the CTJ should be individualized, but may not prevent additional proximal junctional-level problems in the management of thoracolumbar fusion PJK. LEVEL OF EVIDENCE/METHODS:Level IV.
PMID: 42615895
ISSN: 2380-0194
CID: 6071470

Clinical and Economic Burden of Poor Bone Health in Adult Spinal Deformity Surgery: A Multicenter Cohort Study

Passias, Peter G; Daher, Mohammad; Chatzis, Kyriakos D; Lafage, Virginie; Lafage, Renaud; Nayak, Pratibha; Schoenfeld, Andrew; Khalife, Marc; Haddad, Sleiman; Ferrero, Emmanuelle; Line, Breton; Diebo, Bassel; Daniels, Alan H; Mullin, Jeffrey P; Hamilton, D Kojo; Buell, Thomas; Okonkwo, David O; Gum, Jeffrey; Theologis, Alekos; Mummaneni, Praveen; Chou, Dean; Mundis, Gregory; Lau, Darryl; Bunch, Joshua; Carlson, Brandon; Lewis, Stephen; Scheer, Justin; Eastlack, Robert; Kebaish, Khaled; Gupta, Munish; Kim, Han Jo; Soroceanu, Alex; Mikula, Anthony; Protopsaltis, Themistocles; Yagi, Mitsuru; Hosogane, Naobumi; Lenke, Lawrence; Hostin, Richard; Smith, Justin; Klineberg, Eric; Ames, Christopher; Schwab, Frank; Bess, Shay; Shaffrey, Christopher; ,
STUDY DESIGN/METHODS:Retrospective review of the prospectively enrolled, multicenter ASD database. OBJECTIVE:We sought to compare clinical and economic outcomes for ASD patients with poor-bone-health versus those with normal-bone-health. SUMMARY OF BACKGROUND DATA/BACKGROUND:Osteoporosis is a common comorbidity in the adult spinal deformity (ASD) population, with a reported prevalence of 14-29%. METHODS:Patients were categorized into five guideline-concordant cohorts: Confirmed-Osteoporosis (COP: patients with osteopenia or osteoporosis based on DEXA values or a formal preoperative diagnosis); Fracture Osteoporosis (Fx: patients with fragility fractures); Fracture-without-confirmed-Osteoporosis (FXNO: patients with fragility fractures but no osteopenia/osteoporosis diagnosis based on DEXA/formal documentation); Confirmed/fracture-osteoporosis (COP/FX: patients with fragility fractures and/or osteopenia/osteoporosis by DEXA/diagnosis); and the Normal cohort (Nm: patients with normal bone health and no fragility fractures). Cost analyses were performed using multivariate linear regression controlling for BMI, diabetes, baseline deformity (T1PA), and levels fused. Multivariate logistic regression assessed risk of reoperation/complications between groups while controlling for BMI, diabetes, baseline deformity (T1PA), and levels fused. RESULTS:Overall, 205 patients were included in the Fx-cohort, 136 patients in the COP-cohort, 115 patients in the FXNO-cohort, and 71 patients Nm-cohort. Compared with the Nm-cohort, osteoporotic-cohorts demonstrated significantly worse baseline spinopelvic alignment and higher comorbidity burden. Fx and FXNO cohorts had higher rates of PJF (Fx: 11.2% vs 1.4%, aOR 7.8; FXNO: 10.4% vs 1.4%, aOR 9.0) and revision surgery (Fx: 19.5% vs 7.0%, aOR 2.8; FXNO: 22.6% vs 7.0%, aOR 3.6). Osteoporotic cohorts demonstrated significantly higher short-term QALYs at 6 weeks (P<0.05) with no difference in the long term. Revision surgery averaged $95,117 per case, translating into an estimated $4.0 million potential-cost-savings. CONCLUSION/CONCLUSIONS:Our findings highlight the clinical and economic burden of untreated poor bone quality in the setting of ASD surgery. Proactive medical management is crucial to mitigate complications, reduce revisions, and significantly lower healthcare costs in ASD surgery.
PMID: 42430752
ISSN: 1528-1159
CID: 6064312

A Radiomics-Driven Model to Distinguish Between Clinically Similar Myxopapillary Ependymomas and Lumbosacral Schwannomas

Palla, Adhith; Goff, Nicolas K; Perdikis, Blake; Khan, Hammad A; Grin, Eric A; Valliani, Aly; Patel, Roshni; Yang, Jonathan T; McFaline-Figueroa, J Ricardo; Lau, Darryl; Frempong-Boadu, Anthony; Oermann, Eric K; Laufer, Ilya
BACKGROUND AND OBJECTIVES/OBJECTIVE:Myxopapillary ependymomas (MPE) and intradural lumbosacral schwannomas may be challenging to distinguish based on presenting characteristics and preoperative imaging. Accurate differentiation is crucial, as MPEs carry a risk of cerebrospinal fluid dissemination and warrant earlier intervention, a more tailored surgical strategy, consideration for adjuvant radiation, and frequent surveillance. Here, we describe our institutional experience with these tumors and develop a radiomics-based machine learning model to help distinguish them on preoperative imaging. METHODS:Institutional surgical records from 2011 to 2025 were queried and clinical data were extracted for the retrospective cohort analysis. Tumors were manually segmented in ITK-Snap from T1 postcontrast images, and radiomics features were extracted using the PyRadiomics package. An ensemble of random forest, k-nearest neighbors, and naive Bayes classifiers was trained on a subset of radiomics features using nested cross-validation. RESULTS:< .001) in MPEs, likely due to longitudinal tumor growth along the filum. Excluding scoliotic patients did not significantly alter discrimination, suggesting robustness to vertebral column malalignment that may coexist with intradural tumors. CONCLUSION/CONCLUSIONS:A radiomics-based machine learning model demonstrated excellent discriminative ability between MPE and lumbosacral schwannoma, achieving high accuracy and robustness to vertebral alignment variations. These results suggest that radiomics-based models may be developed into a useful tool for preoperative planning and patient counseling.
PMCID:13354379
PMID: 42434191
ISSN: 2834-4383
CID: 6064422

Benchmarking risk prediction tools in spine surgery: a national registry analysis of albumin, frailty, and surgical calculators

Farid, Michael; Blasdel, Nicolai; Wang, Justin; Perdikis, Blake; O'Leary, Sean; Darko, Kwadwo; Barrie, Umaru; Khan, Hammad A; Lau, Darryl
OBJECTIVE:Preoperative risk-stratification tools, including frailty, nutritional, and surgical risk metrics, are used to predict complications after spine surgery. The relative performance of these tools across complication types and surgical subgroups is not well characterized. This study aimed to compare the predictive performance of 5 risk metrics, American College of Surgeons Surgical Risk Calculator (ACS SRC), serum albumin, Risk Analysis Index (RAI), modified 5-item frailty index (mFI-5), and Geriatric Nutritional Risk Index (GNRI), for perioperative complications. METHODS:The authors analyzed 362,145 adult spine surgery patients from the American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) from 2017 to 2022. Adjusted odds ratios were estimated via multivariable logistic regression, controlling for age, sex, urgency, and procedure type (defined by Current Procedural Terminology [CPT] codes). Subgroup analyses were stratified by ICD-10 diagnosis category, including degenerative disease, tumor, trauma, infection, and spinal deformity. Discrimination for predicting complications was assessed using C-statistics and DeLong's test. RESULTS:Across all endpoints, ACS SRC had the best predictive accuracy: mortality C-statistic (95% CI) 0.908 (0.900-0.916), Clavien-Dindo grade IV (CD-IV) 0.823 (0.816-0.830), and major complications 0.749 (0.745-0.752). Serum albumin, despite being a single laboratory value, ranked second with mortality C-statistic (95% CI) 0.820 (0.807-0.833), CD-IV 0.734, and major complications 0.682 and showed strong discrimination for infectious complications (e.g., sepsis, septic shock, surgical site infection, and urinary tract infection), as well as for hospital length of stay and nonhome discharge. Compared to frailty-based metrics, albumin showed significantly better predictive value (p < 0.001 for pairwise comparisons) and maintained its advantages across all subgroups, including high-risk groups such as infection, trauma, and tumor cases. RAI provided moderate mortality prediction (C-statistic 0.807) and was most effective for predicting cardiovascular events, while both GNRI (0.753) and mFI-5 (0.647) were less consistent and demonstrated weaker associations with adverse outcomes. Multivariable regression confirmed that lower preoperative albumin and higher ACS SRC predictions were robust, independent predictors of increased risk for major complications, CD-IV events, and mortality. These performance patterns remained stable across surgical indications and in subgroup analyses. CONCLUSIONS:ACS SRC remains among the comprehensive tools for risk stratification in spine surgery. Serum albumin offers strong, consistent predictive value, especially for infectious, respiratory, and life-threatening complications, and may be a valuable alternative when calculator inputs are incomplete.
PMID: 42430800
ISSN: 1547-5646
CID: 6064332

Preoperative alignment and risk of proximal junctional failure : a framework for upper instrumented vertebra selection in adult spinal deformity

Hills, Jeffrey; Lenke, Lawrence G; Smith, Justin S; Shaffrey, Christopher I; Lafage, Virginie; Lafage, Renaud; Bess, Shay; Kelly, Michael P; ,; ,; Turner, Jay; Uribe, Juan; Daniels, Alan; Diebo, Bassel; Lenke, Lawrence G; Chou, Dean; Shaffrey, Christopher I; Passias, Peter G; Hostin, Richard; Kim, Han Jo; Kebaish, Khaled; Lee, Sang; Burton, Douglas C; Carlson, Brandon; Bunch, Joshua; Schwab, Frank J; Lafage, Virginie; Lafage, Renaud; Protopsaltis, Themistocles S; Lau, Darryl; Bess, Shay; Line, Breton; Kelly, Michael P; Mundis, Gregory M; Eastlack, Robert K; Klineberg, Eric O; Ames, Christopher P; Mumanneni, Praveen; Theologis, Alekos; Alan, Nima; Yoon, Jon; Soroceanu, Alex; Gum, Jeffrey L; Hamilton, Kojo; Okonkwo, David; Buell, Thomas; Lewis, Stephen; Smith, Justin S; Gupta, Munish C; Greenberg, Jacob; Anand, Neel; Scheer, Justin; Fu, Kai-Ming; Park, Paul; Javidan, Yashar; Wang, Michael; Stephens, Byron; Zuckerman, Scott
AIMS/UNASSIGNED:Proximal junctional kyphosis (PJK) remains a major complication after surgery for adult spinal deformity (ASD). While postoperative alignment is a recognized modifiable risk factor, objective methods for selecting the upper instrumented vertebra (UIV), a key modifiable factor, are lacking. We aimed to determine whether preoperative sagittal alignment, specifically cervicothoracic alignment, predicts the risk of PJK, and whether this risk can be mitigated by UIV selection, focusing on factors available at the time of surgical planning. METHODS/UNASSIGNED:From a multicentre, prospective ASD registry, we identified patients who had undergone fusion to the sacrum or pelvis and had an upper (T1-T5) or lower thoracic (T9-L1) UIV, with a two-year or more radiological follow-up, excluding those with a previous fusion over more than four levels. The primary outcome was PJK within two years. Multivariable logistic regression modelled the risk of PJK by UIV region, preoperative C2-T9 pelvic angle (PA), age, sex, and pelvic incidence, testing for interaction between UIV region and C2-T9 PA. Adjusted absolute risk reduction (ARR) and number needed to be exposed (NNEB) were calculated. Multivariable linear regression estimated two-year patient-reported outcome measures, adjusting for baseline scores, age, UIV, and PJK. RESULTS/UNASSIGNED:A total of 627 patients across 20 centres were included (median age 66 years (IQR 59 to 70); 483 (77%) female). The UIV was lower thoracic in 380 (61%) and upper thoracic in 247 (39%) patients. PJK occurred in 149 (39%) lower thoracic and 38 (15%) upper thoracic UIV patients. There was a significant interaction (p = 0.028) between preoperative C2-T9 PA and UIV region. At a preoperative C2-T9 PA of 14° (cohort median), an upper thoracic UIV had an adjusted ARR of 36% and NNEB was 2.8. Females had an adjusted odds ratio of 1.62 (95% CI 1.03 to 2.59; p = 0.042) for PJK. CONCLUSION/UNASSIGNED:Worse preoperative sagittal malalignment, measured by C2-T9 PA, was associated with a higher risk of PJK and depended on UIV region. An upper thoracic UIV in patients with high preoperative C2-T9 PA may reduce PJK.
PMID: 42379558
ISSN: 2049-4408
CID: 6062742

Machine Learning-Based Prediction of Independent Ambulation Following Intramedullary Spinal Cord Tumor Resection

Perdikis, Blake; Palla, Adhith; Goff, Nicolas K; Khan, Hammad A; Rai, Sumedha; Budimlija, Zoran; Lau, Darryl; Frempong-Boadu, Anthony; Laufer, Ilya
BACKGROUND AND OBJECTIVES/OBJECTIVE:Intramedullary spinal cord tumor (IMSCT) resection carries a high risk of postoperative neurological deficit because of neural tract manipulation and myelotomy. Although short-term and long-term neurological recovery represent key treatment outcomes, current prognostication methods are lacking and would benefit from further complex analysis. METHODS:From March 2009 to August 2025, all adult IMSCT resections at our institution were reviewed. Demographic, oncologic, and perioperative data were extracted from electronic medical records. This included preoperative and follow-up neurological examination data in the form of American Spinal Injury Association Impairment Scale (AIS) grading, Modified McCormick Scale (MMCS), and ambulatory status. Independent ambulation served as the primary outcome for 4 machine learning models. Each model was sequentially evaluated using area under the receiver operating characteristic curve (AUROC). RESULTS:Fifty-four patients underwent 55 surgeries for IMSCT resection. Encapsulated lesions predominated IMSCT pathology, with grade II ependymoma comprising 28 (50.9%) resections, 5 hemangioblastomas (9.1%), and 5 cavernous hemangiomas (9.1%). Gross total resection was achieved in 36 cases (65.5%), with encapsulated tumors more readily achieving gross total resection vs unencapsulated (84.6% vs 18.8%, P < .01). By 4 weeks, conversion of MMCS, but not AIS grade, significantly correlated with concurrent ambulatory conversion (P < .01 vs P = .15). At 6 months, both AIS grade conversion (P < .01) and MMCS conversion (P < .01) significantly correlated with ambulatory conversion. For predicting ambulation at latest follow-up from 4 weeks postoperatively, the comprehensive granular model achieved an AUROC of 0.833, outperforming the AIS grade (0.583), American Spinal Injury Association Motor Score (0.667), and MMCS (0.667) models. By the 6-month follow-up, the comprehensive granular model achieved strong discrimination (AUROC 1.00). CONCLUSION/CONCLUSIONS:Follow-up IMSCT data demonstrate a postoperative lability that stabilizes by 6 months into a reliably modeled outcome. By enhancing the granularity of recovery data, accurate independent ambulation modeling may improve counseling for patients with IMSCT.
PMID: 42240329
ISSN: 1524-4040
CID: 6044392

CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications

Alyakin, Anton; Stryker, Jaden; Alber, Daniel Alexander; Lee, Jin Vivian; Sangwon, Karl L; Duderstadt, Brandon; Save, Akshay; Kurland, David; Frome, Spencer; Singh, Shrutika; Zhang, Jeff; Yang, Eunice; Park, Ki Yun; Orillac, Cordelia; Valliani, Aly A; Neifert, Sean; Liu, Albert; Patel, Aneek; Livia, Christopher; Lau, Darryl; Laufer, Ilya; Rozman, Peter A; Hidalgo, Eveline Teresa; Riina, Howard; Feng, Rui; Hollon, Todd; Aphinyanaphongs, Yindalon; Golfinos, John G; Snyder, Laura; Leuthardt, Eric C; Kondziolka, Douglas; Oermann, Eric Karl
BACKGROUND AND OBJECTIVES/OBJECTIVE:General purpose vision-language models (VLMs) demonstrate impressive capabilities, but their opaque training on uncurated internet data poses critical limitations for high-stakes decision making, such as in neurosurgery. We present CNS-Obsidian, a neurosurgical VLM trained on peer-reviewed neurosurgical literature, and demonstrate its clinical utility compared with GPT-4o in a real-world setting. METHODS:We compiled 23 984 articles from Neurosurgery Publications journals, yielding 78 853 figures and captions. Using GPT-4o and Claude Sonnet-3.5, we converted these image-text pairs into 263 064 training samples across 3 formats: instruction fine-tuning, multiple-choice questions, and differential diagnosis. We trained CNS-Obsidian, a fine-tune of the 34-billion parameter Large Language and Visual Assistant-Next model. In a blinded, randomized deployment trial at NYU Langone Health (August 30-November 30, 2024), neurosurgeons were assigned to use either CNS-Obsidian or a Health Insurance Portability and Accountability Act-compliant GPT-4o end point as a diagnostic copilot after patient consultations. Primary outcomes were diagnostic helpfulness and accuracy, assessed through user ratings and presence of the correct diagnosis within the VLM-provided differential, respectively. RESULTS:CNS-Obsidian matched GPT-4o on synthetic questions (76.13% vs 77.54%, P = .235), but only achieved 46.81% accuracy on human-generated questions vs GPT-4o's 65.70% (P < 10-15). In the randomized trial, 70 consultations were evaluated (32 CNS-Obsidian, 38 GPT-4o) from 959 total consults (7.3% utilization). CNS-Obsidian received positive ratings in 40.62% of cases vs 57.89% for GPT-4o (P = .230). Both models included correct diagnosis in approximately 60% of cases (59.38% vs 65.79%, P = .626). CONCLUSION/CONCLUSIONS:Domain-specific VLMs trained on curated scientific literature can approach frontier model performance in specialized medical domains despite being orders of magnitude smaller and less expensive to train. This establishes a transparent framework for scientific communities to build specialized artificial intelligence models. However, low clinical utilization suggests chatbot interfaces may not align with specialist workflows, indicating need for alternative artificial intelligence integration strategies.
PMID: 42153721
ISSN: 1524-4040
CID: 6037862