The most recent version of ChatGPT, an AI chatbot developed for language interpretation and response generation, has successfully passed a radiology board-style exam, demonstrating both its potential and limitations, according to research studies published in the Radiological Society of North America’s journal.
The latest version of ChatGPT passed a radiology board-style exam, highlighting the potential of large language models but also revealing limitations that hinder reliability, according to two new research studies published in Radiology, a journal of the Radiological Society of North America (RSNA).
ChatGPT is an artificial intelligence (AI) chatbot that uses a deep learning model to recognize patterns and relationships between words in its vast training data to generate human-like responses based on a prompt. But since there is no source of truth in its training data, the tool can generate responses that are factually incorrect.
ChatGPT was recently named the fastest growing consumer application in history, and similar chatbots are being incorporated into popular search engines like Google and Bing that physicians and patients use to search for medical information, Dr. Bhayana noted.
To assess its performance on radiology board exam questions and explore strengths and limitations, Dr. Bhayana and colleagues first tested ChatGPT based on GPT-3.5, currently the most commonly used version. The researchers used 150 multiple-choice questions designed to match the style, content and difficulty of the Canadian Royal College and American Board of Radiology exams.
The questions did not include images and were grouped by question type to gain insight into performance: lower-order (knowledge recall, basic understanding) and higher-order (apply, analyze, synthesize) thinking. The higher-order thinking questions were further subclassified by type (description of imaging findings, clinical management, calculation and classification, disease associations).
The performance of ChatGPT was evaluated overall and by question type and topic. Confidence of language in responses was also assessed.
The researchers found that ChatGPT based on GPT-3.5 answered 69% of questions correctly (104 of 150), near the passing grade of 70% used by the Royal College in Canada. The model performed relatively well on questions requiring lower-order thinking (84%, 51 of 61), but struggled with questions involving higher-order thinking (60%, 53 of 89). More specifically, it struggled with higher-order questions involving description of imaging findings (61%, 28 of 46), calculation and classification (25%, 2 of 8), and application of concepts (30%, 3 of 10). Its poor performance on higher-order thinking questions was not surprising given its lack of radiology-specific pretraining.
GPT-4 was released in March 2023 in limited form to paid users, specifically claiming to have improved advanced reasoning capabilities over GPT-3.5.
In a follow-up study, GPT-4 answered 81% (121 of 150) of the same questions correctly, outperforming GPT-3.5 and exceeding the passing threshold of 70%. GPT-4 performed much better than GPT-3.5 on higher-order thinking questions (81%), more specifically those involving description of imaging findings (85%) and application of concepts (90%).
The findings suggest that GPT-4’s claimed improved advanced reasoning capabilities translate to enhanced performance in a radiology context. They also suggest improved contextual understanding of radiology-specific terminology, including imaging descriptions, which is critical to enable future downstream applications.
“Our study demonstrates an impressive improvement in performance of ChatGPT in radiology over a short time period, highlighting the growing potential of large language models in this context,” Dr. Bhayana said.
GPT-4 showed no improvement on lower-order thinking questions (80% vs 84%) and answered 12 questions incorrectly that GPT-3.5 answered correctly, raising questions related to its reliability for information gathering.
“We were initially surprised by ChatGPT’s accurate and confident answers to some challenging radiology questions, but then equally surprised by some very illogical and inaccurate assertions,” Dr. Bhayana said. “Of course, given how these models work, the inaccurate responses should not be particularly surprising.”
ChatGPT’s dangerous tendency to produce inaccurate responses, termed hallucinations, is less frequent in GPT-4 but still limits usability in medical education and practice at present.
Both studies showed that ChatGPT used confident language consistently, even when incorrect. This is particularly dangerous if solely relied on for information, Dr. Bhayana notes, especially for novices who may not recognize confident incorrect responses as inaccurate.
“To me, this is its biggest limitation. At present, ChatGPT is best used to spark ideas, help start the medical writing process and in data summarization. If used for quick information recall, it always needs to be fact-checked,” Dr. Bhayana said.

News
Lightning sparks scientists’ design of ultraviolet-C device for food sanitization
Scientists at the University of Illinois Urbana-Champaign have developed a portable, self-powered ultraviolet-C device called the Tribo-sanitizer that can inactivate two of the bacteria responsible for many foodborne illnesses and deaths. The Tribo-sanitizer's UVC [...]
3D Eye Scans Emerge as a Crucial Tool in Combating Kidney Disease
A new study indicates that 3D retinal scans could revolutionize the early detection and monitoring of kidney disease, offering a non-invasive and efficient diagnostic tool. 3D eye scans can reveal vital clues about kidney [...]
Researchers develop a blood test to identify individuals at risk of developing Parkinson’s disease
Research carried out at Oxford's Nuffield Department of Clinical Neurosciences has led to the development of a new blood-based test to identify the pathology that triggers Parkinson's disease before the main symptoms occur. This [...]
“Challenging the Paradigm” – Scientists Develop New Approach To Stop Cancer Growth
Biochemists at Case Western Reserve are concentrating on the degradation of a key protein that drives cancer; represents a major shift in research. Biochemical researchers at Case Western Reserve University have discovered a a new function [...]
Researcher develops a chatbot with an expertise in nanomaterials
A researcher has just finished writing a scientific paper. She knows her work could benefit from another perspective. Did she overlook something? Or perhaps there's an application of her research she hadn't thought of. [...]
Research shows human behavior guided by fast changes in dopamine levels
What happens in the human brain when we learn from positive and negative experiences? To help answer that question and better understand decision-making and human behavior, scientists are studying dopamine. Dopamine is a neurotransmitter [...]
Tiny robots made from human cells heal damaged tissue
The ‘anthrobots’ were able to repair a scratch in a layer of neurons in the lab. Scientists have developed tiny robots made of human cells that are able to repair damaged neural tissue1. The [...]
Antimicrobial Resistance – A Global Concern
Key facts Antimicrobial resistance (AMR) is one of the top global public health and development threats. It is estimated that bacterial AMR was directly responsible for 1.27 million global deaths in 2019 and contributed to [...]
Advancing Pancreatic Cancer Treatment with Nanoparticle-Based Chemotherapy
Pancreatic cancer, a particularly lethal form of cancer and the fourth leading cause of cancer-related deaths in the western world, often remains undiagnosed until its advanced stages due to a lack of early symptoms. [...]
The ‘jigglings and wigglings of atoms’ reveal key aspects of COVID-19 virulence evolution
Richard Feynman famously stated, "Everything that living things do can be understood in terms of the jigglings and wigglings of atoms." This week, Nature Nanotechnology features a study that sheds new light on the evolution of the coronavirus [...]
AI system self-organizes to develop features of brains of complex organisms
Cambridge scientists have shown that placing physical constraints on an artificially-intelligent system—in much the same way that the human brain has to develop and operate within physical and biological constraints—allows it to develop features [...]
How Blind People Recognize Faces via Sound
Summary: A new study reveals that people who are blind can recognize faces using auditory patterns processed by the fusiform face area, a brain region crucial for face processing in sighted individuals. The study employed [...]
Treating tumors with engineered dendritic cells
Cancer biologists at EPFL, UNIGE, and the German Cancer Research Center (Heidelberg) have developed a novel immunotherapy that does not require knowledge of a tumor's antigenic makeup. The new results may pave the way [...]
Networking nano-biosensors for wireless communication in the blood
Biological computing machines, such as micro and nano-implants that can collect important information inside the human body, are transforming medicine. Yet, networking them for communication has proven challenging. Now, a global team, including EPFL [...]
Popular Hospital Disinfectant Ineffective Against Common Superbug
Research conducted during World Antimicrobial Awareness Week examines the effects of employing suggested chlorine-based chemicals to combat Clostridioides difficile, the leading cause of antibiotic-related illness in healthcare environments worldwide. A recent study reveals that a [...]
Subjectivity and the Evolution of AI Philosophy
An Historical Overview of the Philosophy of Artificial Intelligence by Anton Vokrug Many famous people in the philosophy of technology have tried to comprehend the essence of technology and link it to society and human [...]