The most recent version of ChatGPT, an AI chatbot developed for language interpretation and response generation, has successfully passed a radiology board-style exam, demonstrating both its potential and limitations, according to research studies published in the Radiological Society of North America’s journal.
The latest version of ChatGPT passed a radiology board-style exam, highlighting the potential of large language models but also revealing limitations that hinder reliability, according to two new research studies published in Radiology, a journal of the Radiological Society of North America (RSNA).
ChatGPT is an artificial intelligence (AI) chatbot that uses a deep learning model to recognize patterns and relationships between words in its vast training data to generate human-like responses based on a prompt. But since there is no source of truth in its training data, the tool can generate responses that are factually incorrect.
ChatGPT was recently named the fastest growing consumer application in history, and similar chatbots are being incorporated into popular search engines like Google and Bing that physicians and patients use to search for medical information, Dr. Bhayana noted.
To assess its performance on radiology board exam questions and explore strengths and limitations, Dr. Bhayana and colleagues first tested ChatGPT based on GPT-3.5, currently the most commonly used version. The researchers used 150 multiple-choice questions designed to match the style, content and difficulty of the Canadian Royal College and American Board of Radiology exams.
The questions did not include images and were grouped by question type to gain insight into performance: lower-order (knowledge recall, basic understanding) and higher-order (apply, analyze, synthesize) thinking. The higher-order thinking questions were further subclassified by type (description of imaging findings, clinical management, calculation and classification, disease associations).
The performance of ChatGPT was evaluated overall and by question type and topic. Confidence of language in responses was also assessed.
The researchers found that ChatGPT based on GPT-3.5 answered 69% of questions correctly (104 of 150), near the passing grade of 70% used by the Royal College in Canada. The model performed relatively well on questions requiring lower-order thinking (84%, 51 of 61), but struggled with questions involving higher-order thinking (60%, 53 of 89). More specifically, it struggled with higher-order questions involving description of imaging findings (61%, 28 of 46), calculation and classification (25%, 2 of 8), and application of concepts (30%, 3 of 10). Its poor performance on higher-order thinking questions was not surprising given its lack of radiology-specific pretraining.
GPT-4 was released in March 2023 in limited form to paid users, specifically claiming to have improved advanced reasoning capabilities over GPT-3.5.
In a follow-up study, GPT-4 answered 81% (121 of 150) of the same questions correctly, outperforming GPT-3.5 and exceeding the passing threshold of 70%. GPT-4 performed much better than GPT-3.5 on higher-order thinking questions (81%), more specifically those involving description of imaging findings (85%) and application of concepts (90%).
The findings suggest that GPT-4’s claimed improved advanced reasoning capabilities translate to enhanced performance in a radiology context. They also suggest improved contextual understanding of radiology-specific terminology, including imaging descriptions, which is critical to enable future downstream applications.
“Our study demonstrates an impressive improvement in performance of ChatGPT in radiology over a short time period, highlighting the growing potential of large language models in this context,” Dr. Bhayana said.
GPT-4 showed no improvement on lower-order thinking questions (80% vs 84%) and answered 12 questions incorrectly that GPT-3.5 answered correctly, raising questions related to its reliability for information gathering.
“We were initially surprised by ChatGPT’s accurate and confident answers to some challenging radiology questions, but then equally surprised by some very illogical and inaccurate assertions,” Dr. Bhayana said. “Of course, given how these models work, the inaccurate responses should not be particularly surprising.”
ChatGPT’s dangerous tendency to produce inaccurate responses, termed hallucinations, is less frequent in GPT-4 but still limits usability in medical education and practice at present.
Both studies showed that ChatGPT used confident language consistently, even when incorrect. This is particularly dangerous if solely relied on for information, Dr. Bhayana notes, especially for novices who may not recognize confident incorrect responses as inaccurate.
“To me, this is its biggest limitation. At present, ChatGPT is best used to spark ideas, help start the medical writing process and in data summarization. If used for quick information recall, it always needs to be fact-checked,” Dr. Bhayana said.

News
A potential milestone in cancer therapy
Researchers from the University of Bern, Inselspital, University Hospital Bern, and the University of Connecticut have made a significant breakthrough in the fight against cancer. They identified a previously unknown weak point of prostate [...]
Cardiovascular Crystal Ball: New Tool Predicts Future Heart Disease Risk
Faculty members at the UM School of Medicine have created a cutting-edge tool that enables the early identification and assessment of risks in vulnerable patients. Heart disease, being the leading cause of death globally, [...]
Scientists analyze a single atom with X-rays for the first time
In the most powerful X-ray facilities in the world, scientists can analyze samples so small they contain only 10,000 atoms. Smaller sizes have proved exceedingly difficult to achieve, but a multi-institutional team has scaled [...]
AI Demonstrates Superior Performance in Predicting Breast Cancer
AI algorithms outperformed traditional clinical risk models in a large-scale study, predicting five-year breast cancer risk more accurately. These models use mammograms as the single data source, offering potential advantages in individualizing patient care [...]
Stanford Medicine Reveals: Tiny DNA Circles Defying Genetic Laws Drive Cancer Formation
Tiny circles of DNA harbor cancer-associated oncogenes and immunomodulatory genes promoting cancer development. They arise during the transformation from pre-cancer to cancer, say Stanford Medicine-led team. Tiny circles of DNA that defy the accepted laws of [...]
Death to Blood Cancer Cells: New Drug Combination Could Revive the Power of Leading Treatment
Future clinical trials will be conducted to investigate whether the combination of chloroquine and venetoclax can prevent disease recurrence. Although new drugs have been developed to induce cancer cell death in individuals with acute [...]
Illuminating Science: X-Rays Visualize How One of Nature’s Strongest Bonds Breaks
Scientists have deciphered how an activated catalyst breaks down the strong carbon-hydrogen bonds in potent greenhouse gas methane, according to a study published in Science. Using advanced X-ray technology and quantum-chemical calculations, they tracked the [...]
Using magnetic nanoparticles as a rapid test for sepsis
Qun Ren, an Empa researcher, and her team are currently developing a diagnostic procedure that can rapidly detect life-threatening blood poisoning caused by staphylococcus bacteria. Staphylococcal sepsis is fatal in up to 40% of [...]
Team develops nanoparticles to deliver brain cancer treatment
University of Queensland researchers have developed a nanoparticle to take a chemotherapy drug into fast growing, aggressive brain tumors. Research team lead Dr. Taskeen Janjua from UQ's School of Pharmacy said the new silica [...]
Tumor Avatars – A New Approach to Personalized Cancer Treatment
A team from the University of Geneva (UNIGE) has devised a novel method for customizing treatments by testing them on artificial tumors. Determining the optimal treatment for colon cancer can be challenging as each [...]
STING Like a Bee: MIT’s Revolutionary Approach to Cancer Immunotherapy
A cancer vaccine combining checkpoint blockade therapy and a STING-activating drug eliminates tumors and prevents recurrence in mice. MIT researchers have engineered a therapeutic cancer vaccine that targets the STING pathway, vital for immune response [...]
AI Battles Superbugs: Helps Find New Antibiotic Drug To Combat Drug-Resistant Infections
The machine-learning algorithm identified a compound that kills Acinetobacter baumannii, a bacterium that lurks in many hospital settings. Using an artificial intelligence algorithm, researchers at MIT and McMaster University have identified a new antibiotic that can kill a [...]
Cancer and AI – Can ChatGPT Be Trusted?
A study published in the Journal of The National Cancer Institute Cancer Spectrum delved into the increasing use of chatbots and artificial intelligence (AI) in providing cancer-related information. The researchers discovered that these digital resources accurately [...]
Breathing New Life: Oxygen Therapy Improves Heart Function in Long COVID Patients
A small trial has found that hyperbaric oxygen therapy (HBOT) may help restore proper heart function in patients with post-COVID syndrome, with participants in the HBOT group experiencing a significant increase in global longitudinal [...]
Wireless Brain-Spine Interface: A Leap Towards Reversing Paralysis
Summary: In a pioneering study, researchers designed a wireless brain-spine interface enabling a paralyzed man to walk naturally again. The ‘digital bridge’ comprises two electronic implants — one on the brain and another on the [...]
New study reveals a gel that promises to wipe out brain cancer for good
An anti-cancer gel promises to wipe out glioblastoma permanently, a feat that's never been accomplished by any drug or surgery. So what makes this gel so special? Scientists at Johns Hopkins University (JHU) have [...]