The most recent version of ChatGPT, an AI chatbot developed for language interpretation and response generation, has successfully passed a radiology board-style exam, demonstrating both its potential and limitations, according to research studies published in the Radiological Society of North America’s journal.
The latest version of ChatGPT passed a radiology board-style exam, highlighting the potential of large language models but also revealing limitations that hinder reliability, according to two new research studies published in Radiology, a journal of the Radiological Society of North America (RSNA).
ChatGPT is an artificial intelligence (AI) chatbot that uses a deep learning model to recognize patterns and relationships between words in its vast training data to generate human-like responses based on a prompt. But since there is no source of truth in its training data, the tool can generate responses that are factually incorrect.
ChatGPT was recently named the fastest growing consumer application in history, and similar chatbots are being incorporated into popular search engines like Google and Bing that physicians and patients use to search for medical information, Dr. Bhayana noted.
To assess its performance on radiology board exam questions and explore strengths and limitations, Dr. Bhayana and colleagues first tested ChatGPT based on GPT-3.5, currently the most commonly used version. The researchers used 150 multiple-choice questions designed to match the style, content and difficulty of the Canadian Royal College and American Board of Radiology exams.
The questions did not include images and were grouped by question type to gain insight into performance: lower-order (knowledge recall, basic understanding) and higher-order (apply, analyze, synthesize) thinking. The higher-order thinking questions were further subclassified by type (description of imaging findings, clinical management, calculation and classification, disease associations).
The performance of ChatGPT was evaluated overall and by question type and topic. Confidence of language in responses was also assessed.
The researchers found that ChatGPT based on GPT-3.5 answered 69% of questions correctly (104 of 150), near the passing grade of 70% used by the Royal College in Canada. The model performed relatively well on questions requiring lower-order thinking (84%, 51 of 61), but struggled with questions involving higher-order thinking (60%, 53 of 89). More specifically, it struggled with higher-order questions involving description of imaging findings (61%, 28 of 46), calculation and classification (25%, 2 of 8), and application of concepts (30%, 3 of 10). Its poor performance on higher-order thinking questions was not surprising given its lack of radiology-specific pretraining.
GPT-4 was released in March 2023 in limited form to paid users, specifically claiming to have improved advanced reasoning capabilities over GPT-3.5.
In a follow-up study, GPT-4 answered 81% (121 of 150) of the same questions correctly, outperforming GPT-3.5 and exceeding the passing threshold of 70%. GPT-4 performed much better than GPT-3.5 on higher-order thinking questions (81%), more specifically those involving description of imaging findings (85%) and application of concepts (90%).
The findings suggest that GPT-4’s claimed improved advanced reasoning capabilities translate to enhanced performance in a radiology context. They also suggest improved contextual understanding of radiology-specific terminology, including imaging descriptions, which is critical to enable future downstream applications.
“Our study demonstrates an impressive improvement in performance of ChatGPT in radiology over a short time period, highlighting the growing potential of large language models in this context,” Dr. Bhayana said.
GPT-4 showed no improvement on lower-order thinking questions (80% vs 84%) and answered 12 questions incorrectly that GPT-3.5 answered correctly, raising questions related to its reliability for information gathering.
“We were initially surprised by ChatGPT’s accurate and confident answers to some challenging radiology questions, but then equally surprised by some very illogical and inaccurate assertions,” Dr. Bhayana said. “Of course, given how these models work, the inaccurate responses should not be particularly surprising.”
ChatGPT’s dangerous tendency to produce inaccurate responses, termed hallucinations, is less frequent in GPT-4 but still limits usability in medical education and practice at present.
Both studies showed that ChatGPT used confident language consistently, even when incorrect. This is particularly dangerous if solely relied on for information, Dr. Bhayana notes, especially for novices who may not recognize confident incorrect responses as inaccurate.
“To me, this is its biggest limitation. At present, ChatGPT is best used to spark ideas, help start the medical writing process and in data summarization. If used for quick information recall, it always needs to be fact-checked,” Dr. Bhayana said.

News
Groundbreaking New Way of Measuring Blood Pressure Could Save Thousands of Lives
A new method that improves the accuracy of interpreting blood pressure measurements taken at the ankle could be vital for individuals who are unable to have their blood pressure measured on the arm. A newly developed [...]
Scientist tackles key roadblock for AI in drug discovery
The drug development pipeline is a costly and lengthy process. Identifying high-quality "hit" compounds—those with high potency, selectivity, and favorable metabolic properties—at the earliest stages is important for reducing cost and accelerating the path [...]
Nanoplastics with environmental coatings can sneak past the skin’s defenses
Plastic is ubiquitous in the modern world, and it's notorious for taking a long time to completely break down in the environment - if it ever does. But even without breaking down completely, plastic [...]
Chernobyl scientists discover black fungus feeding on deadly radiation
It looks pretty sinister, but it might actually be incredibly helpful When reactor number four in Chernobyl exploded, it triggered the worst nuclear disaster in history, one which the surrounding area still has not [...]
Long COVID Is Taking A Silent Toll On Mental Health, Here’s What Experts Say
Months after recovering from COVID-19, many people continue to feel unwell. They speak of exhaustion that doesn’t fade, difficulty breathing, or an unsettling mental haze. What’s becoming increasingly clear is that recovery from the [...]
Study Delivers Cancer Drugs Directly to the Tumor Nucleus
A new peptide-based nanotube treatment sneaks chemo into drug-resistant cancer cells, providing a unique workaround to one of oncology’s toughest hurdles. CiQUS researchers have developed a novel molecular strategy that allows a chemotherapy drug to [...]
Scientists Begin $14.2 Million Project To Decode the Body’s “Hidden Sixth Sense”
An NIH-supported initiative seeks to unravel how the nervous system tracks and regulates the body’s internal organs. How does your brain recognize when it’s time to take a breath, when your blood pressure has [...]
Scientists Discover a New Form of Ice That Shouldn’t Exist
Researchers at the European XFEL and DESY are investigating unusual forms of ice that can exist at room temperature when subjected to extreme pressure. Ice comes in many forms, even when made of nothing but water [...]
Nobel-winning, tiny ‘sponge crystals’ with an astonishing amount of inner space
The 2025 Nobel Prize in chemistry was awarded to Richard Robson, Susumu Kitagawa and Omar Yaghi on Oct. 8, 2025, for the development of metal-organic frameworks, or MOFs, which are tunable crystal structures with extremely [...]
Harnessing Green-Synthesized Nanoparticles for Water Purification
A new review reveals how plant- and microbe-derived nanoparticles can power next-gen water disinfection, delivering cleaner, safer water without the environmental cost of traditional treatments. A recent review published in Nanomaterials highlights the potential of green-synthesized nanomaterials (GSNMs) in [...]
Brainstem damage found to be behind long-lasting effects of severe Covid-19
Damage to the brainstem - the brain's 'control center' - is behind long-lasting physical and psychiatric effects of severe Covid-19 infection, a study suggests. Using ultra-high-resolution scanners that can see the living brain in [...]
CT scan changes over one year predict outcomes in fibrotic lung disease
Researchers at National Jewish Health have shown that subtle increases in lung scarring, detected by an artificial intelligence-based tool on CT scans taken one year apart, are associated with disease progression and survival in [...]
AI Spots Hidden Signs of Disease Before Symptoms Appear
Researchers suggest that examining the inner workings of cells more closely could help physicians detect diseases earlier and more accurately match patients with effective therapies. Researchers at McGill University have created an artificial intelligence tool capable of uncovering [...]
Breakthrough Blood Test Detects Head and Neck Cancer up to 10 Years Before Symptoms
Mass General Brigham’s HPV-DeepSeek test enables much earlier cancer detection through a blood sample, creating a new opportunity for screening HPV-related head and neck cancers. Human papillomavirus (HPV) is responsible for about 70% of [...]
Study of 86 chikungunya outbreaks reveals unpredictability in size and severity
The symptoms come on quickly—acute fever, followed by debilitating joint pain that can last for months. Though rarely fatal, the chikungunya virus, a mosquito-borne illness, can be particularly severe for high-risk individuals, including newborns and older [...]
Tiny Fat Messengers May Link Obesity to Alzheimer’s Plaque Buildup
Summary: A groundbreaking study reveals how obesity may drive Alzheimer’s disease through tiny messengers called extracellular vesicles released from fat tissue. These vesicles carry lipids that alter how quickly amyloid-β plaques form, a hallmark of [...]