The most recent version of ChatGPT, an AI chatbot developed for language interpretation and response generation, has successfully passed a radiology board-style exam, demonstrating both its potential and limitations, according to research studies published in the Radiological Society of North America’s journal.
The latest version of ChatGPT passed a radiology board-style exam, highlighting the potential of large language models but also revealing limitations that hinder reliability, according to two new research studies published in Radiology, a journal of the Radiological Society of North America (RSNA).
ChatGPT is an artificial intelligence (AI) chatbot that uses a deep learning model to recognize patterns and relationships between words in its vast training data to generate human-like responses based on a prompt. But since there is no source of truth in its training data, the tool can generate responses that are factually incorrect.
ChatGPT was recently named the fastest growing consumer application in history, and similar chatbots are being incorporated into popular search engines like Google and Bing that physicians and patients use to search for medical information, Dr. Bhayana noted.
To assess its performance on radiology board exam questions and explore strengths and limitations, Dr. Bhayana and colleagues first tested ChatGPT based on GPT-3.5, currently the most commonly used version. The researchers used 150 multiple-choice questions designed to match the style, content and difficulty of the Canadian Royal College and American Board of Radiology exams.
The questions did not include images and were grouped by question type to gain insight into performance: lower-order (knowledge recall, basic understanding) and higher-order (apply, analyze, synthesize) thinking. The higher-order thinking questions were further subclassified by type (description of imaging findings, clinical management, calculation and classification, disease associations).
The performance of ChatGPT was evaluated overall and by question type and topic. Confidence of language in responses was also assessed.
The researchers found that ChatGPT based on GPT-3.5 answered 69% of questions correctly (104 of 150), near the passing grade of 70% used by the Royal College in Canada. The model performed relatively well on questions requiring lower-order thinking (84%, 51 of 61), but struggled with questions involving higher-order thinking (60%, 53 of 89). More specifically, it struggled with higher-order questions involving description of imaging findings (61%, 28 of 46), calculation and classification (25%, 2 of 8), and application of concepts (30%, 3 of 10). Its poor performance on higher-order thinking questions was not surprising given its lack of radiology-specific pretraining.
GPT-4 was released in March 2023 in limited form to paid users, specifically claiming to have improved advanced reasoning capabilities over GPT-3.5.
In a follow-up study, GPT-4 answered 81% (121 of 150) of the same questions correctly, outperforming GPT-3.5 and exceeding the passing threshold of 70%. GPT-4 performed much better than GPT-3.5 on higher-order thinking questions (81%), more specifically those involving description of imaging findings (85%) and application of concepts (90%).
The findings suggest that GPT-4’s claimed improved advanced reasoning capabilities translate to enhanced performance in a radiology context. They also suggest improved contextual understanding of radiology-specific terminology, including imaging descriptions, which is critical to enable future downstream applications.
“Our study demonstrates an impressive improvement in performance of ChatGPT in radiology over a short time period, highlighting the growing potential of large language models in this context,” Dr. Bhayana said.
GPT-4 showed no improvement on lower-order thinking questions (80% vs 84%) and answered 12 questions incorrectly that GPT-3.5 answered correctly, raising questions related to its reliability for information gathering.
“We were initially surprised by ChatGPT’s accurate and confident answers to some challenging radiology questions, but then equally surprised by some very illogical and inaccurate assertions,” Dr. Bhayana said. “Of course, given how these models work, the inaccurate responses should not be particularly surprising.”
ChatGPT’s dangerous tendency to produce inaccurate responses, termed hallucinations, is less frequent in GPT-4 but still limits usability in medical education and practice at present.
Both studies showed that ChatGPT used confident language consistently, even when incorrect. This is particularly dangerous if solely relied on for information, Dr. Bhayana notes, especially for novices who may not recognize confident incorrect responses as inaccurate.
“To me, this is its biggest limitation. At present, ChatGPT is best used to spark ideas, help start the medical writing process and in data summarization. If used for quick information recall, it always needs to be fact-checked,” Dr. Bhayana said.
News
COVID-19 can wake up dormant viruses in the body, large study confirms
Virus reactivations by SARS-CoV-2 could worsen initial symptoms and increase risk of Long Covid. Early in the COVID-19 pandemic, scientists noticed that a SARS-CoV-2 infection can “wake up” other, dormant viruses already present in [...]
Scientists Discover the Brain May Enter a New Biological Phase Between 50 and 75
A single-cell study reveals major changes in immune cells, genome organization, and gene regulation within the aging human hippocampus, offering important insights into brain aging and dementias associated with age. Between roughly ages 50 [...]
Our books now available worldwide!
Online Sellers other than Amazon, Routledge, and IOPP Indigo Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artifcial Intelligence Global Health Care Equivalency In The Age Of Nanotechnology, Nanomedicine And Artificial [...]
Unzipping the Code of Life: Scientists Pinpoint Where DNA First Opens
Researchers mapped where DNA first opens and how a helicase gate may release one strand as genome copying begins. Before a cell can divide, it must open its tightly wound DNA and begin copying the entire [...]
Scientists Tested an 8-Hour Eating Window and Found a Surprising Brain Benefit
Limiting the daily eating window may provide brain benefits beyond those associated with weight loss. Eating within a shorter daily window may provide cognitive benefits beyond those associated with weight loss alone, according to [...]
Focused Ultrasound Opens Blood-Brain Barrier to Treat Brain Cancer
Summary: A new study demonstrates that primary brain tumors (gliomas) are particularly receptive to targeted drug delivery using focused ultrasound (FUS) combined with microbubbles. The team developed a high-resolution MRI protocol to track blood-brain barrier [...]
AI’s promise and practical limits in drug discovery
AI tools are becoming increasingly common in early drug discovery, allowing scientists to analyse data and navigate large volumes of research. However, according to Dr Raminderpal Singh, turning that potential into consistent scientific workflows [...]
GHCE Concept
From the preface of the book Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artificial Intelligence, Edited by Frank Boehm: Since the publication of my first book (Nanomedical Device and Systems [...]
Healthcare Headlines: Challenges and Advances in 2026
Health-related updates reveal financial adjustments by Universal Health Services due to Medicaid reimbursement uncertainties, significant pollution-linked health concerns from French-British oil firm Perenco in Congo, drug trial setbacks, potential restructuring at major medical firms, [...]
Scientists Discover the Brain Protein That Helps Alzheimer’s Spread Through the Brain
Scientists have identified a brain protein that may help Alzheimer’s spread, revealing a potential new target for slowing the disease’s progression. Alzheimer’s disease is closely linked to the accumulation of a toxic form of the protein [...]
How Immune Dysregulation Contributes to Psychiatric Disorders
Introduction Growing evidence suggests that disruptions in immune function may play an important role in the development and progression of psychiatric disorders. However, immune mechanisms probably contribute more strongly in some patients than others, [...]
Electrostatic Discharge Boosts Triboelectric Nanogenerator Current and Enables DC Output
Controlled electrical discharges could enable triboelectric nanogenerators to achieve higher peak currents, extending nano-enabled energy harvesting into chemical processing and self-powered sensing. Paper: Electrostatic discharge as a breakthrough strategy for triboelectric nanogenerators. A new review [...]
Swiss laboratory uses old drugs against rare diseases
Researchers at the University of Geneva are combing through collections of approved drugs to find new therapies for rare diseases – with some success. This approach is gaining traction around the world, while pharmaceutical [...]
Nanozyme Aptasensors Show Promise for Faster Food, Health, and Environmental Testing
By pairing robust artificial enzymes with highly selective aptamers, nanozyme aptasensors could help detect disease biomarkers, pathogens, and contaminants faster, but the review shows that real-world deployment still depends on overcoming matrix interference, biofouling, [...]
Paralyzed Man Feels Sensation Again With Brain Stimulation Device
Aneuroprosthetic system has allowed a man with paralysis to grasp and lift objects and feel touch again. The device helped 42-year-old Keith Thomas of Massapequa, New York, who was paralyzed from the chest down [...]
Global Cancer Cases Could Surge 67% by 2050, New Report Warns
New data reveal major geographic disparities and highlight the urgent need for global action on prevention, early detection, and equitable access to treatment. For roughly one in five people worldwide, cancer will become part [...]















