E-ISSN: 1019-5157
ISSN: 2651-5024
Research
Can Artificial Intelligence Perform Like a Neurosurgeon? Evidence from the Turkish Neurosurgical Competency Exam
Neurosurgey, Eskisehir Osmangazi University; Neurosurgery, Private Umit Hospital, Eskisehir
Accepted: 20/08/2026
Article in Press
Corresponding Author:
Salim TEKİR (drsalimtekir28@gmail.com)
Abstract
Aim
This study evaluated the performance of ChatGPT (GPT-4o) on the Turkish Neurosurgical Society Competency Written Examination, compared its performance with that of human examinees, characterized the types of questions it answered incorrectly, and assessed the logical consistency of its responses.
Material and Methods
This descriptive cross-sectional study evaluated ChatGPT (GPT-4o) responses to questions from the Turkish Neurosurgical Society Competency Written Examination administered from 2022 through 2024. Of the questions evaluated, 299 valid questions were classified as text-based or figure-based. Two researchers independently assessed each response for accuracy and logical consistency. ChatGPT performance was compared with the mean examination scores of human examinees. Statistical analyses included descriptive statistics, a 2-proportion z test, Pearson χ2 test, and Cohen d effect size analysis. Statistical significance was set at p < .05.
Results
Of the 299 valid questions analyzed, ChatGPT (GPT-4o) answered 247 correctly, corresponding to an overall accuracy of 82.6%. Accuracy was significantly higher for text-based questions than for figure-based questions (85.9% vs 66.0%; z = 3.40, p < .001). Response accuracy was strongly associated with logical consistency (χ21 = 271.60, p < .001). human examinees ChatGPT (GPT-4o) also achieved higher scores than the mean scores of human examinees in each examination year. Effect size analysis demonstrated a large effect in 2022 (d = 0.98) and very large effects in 2023 (d = 1.91) and 2024 (d = 1.78).
Conclusion**
ChatGPT (GPT-4o) demonstrated high accuracy and generally consistent reasoning when answering theoretical neurosurgical examination questions and outperformed the average performance of human examinees. However, its substantially lower accuracy on figure-based questions indicates persistent limitations in visual interpretation and multimodal clinical reasoning. These findings suggest that ChatGPT (GPT-4o) may be a useful complementary tool for neurosurgical education and examination preparation but should not be considered a substitute for expert clinical judgment.
This study evaluated the performance of ChatGPT (GPT-4o) on the Turkish Neurosurgical Society Competency Written Examination, compared its performance with that of human examinees, characterized the types of questions it answered incorrectly, and assessed the logical consistency of its responses.
Material and Methods
This descriptive cross-sectional study evaluated ChatGPT (GPT-4o) responses to questions from the Turkish Neurosurgical Society Competency Written Examination administered from 2022 through 2024. Of the questions evaluated, 299 valid questions were classified as text-based or figure-based. Two researchers independently assessed each response for accuracy and logical consistency. ChatGPT performance was compared with the mean examination scores of human examinees. Statistical analyses included descriptive statistics, a 2-proportion z test, Pearson χ2 test, and Cohen d effect size analysis. Statistical significance was set at p < .05.
Results
Of the 299 valid questions analyzed, ChatGPT (GPT-4o) answered 247 correctly, corresponding to an overall accuracy of 82.6%. Accuracy was significantly higher for text-based questions than for figure-based questions (85.9% vs 66.0%; z = 3.40, p < .001). Response accuracy was strongly associated with logical consistency (χ21 = 271.60, p < .001). human examinees ChatGPT (GPT-4o) also achieved higher scores than the mean scores of human examinees in each examination year. Effect size analysis demonstrated a large effect in 2022 (d = 0.98) and very large effects in 2023 (d = 1.91) and 2024 (d = 1.78).
Conclusion**
ChatGPT (GPT-4o) demonstrated high accuracy and generally consistent reasoning when answering theoretical neurosurgical examination questions and outperformed the average performance of human examinees. However, its substantially lower accuracy on figure-based questions indicates persistent limitations in visual interpretation and multimodal clinical reasoning. These findings suggest that ChatGPT (GPT-4o) may be a useful complementary tool for neurosurgical education and examination preparation but should not be considered a substitute for expert clinical judgment.
Keywords
ChatGPT
neurosurgery
artificial intelligence
clinical decision support.
competency examination