The most recent version of ChatGPT, an AI chatbot developed for language interpretation and response generation, has successfully passed a radiology board-style exam, demonstrating both its potential and limitations, according to research studies published in the Radiological Society of North America’s journal.
The latest version of ChatGPT passed a radiology board-style exam, highlighting the potential of large language models but also revealing limitations that hinder reliability, according to two new research studies published in Radiology, a journal of the Radiological Society of North America (RSNA).
ChatGPT is an artificial intelligence (AI) chatbot that uses a deep learning model to recognize patterns and relationships between words in its vast training data to generate human-like responses based on a prompt. But since there is no source of truth in its training data, the tool can generate responses that are factually incorrect.
ChatGPT was recently named the fastest growing consumer application in history, and similar chatbots are being incorporated into popular search engines like Google and Bing that physicians and patients use to search for medical information, Dr. Bhayana noted.
To assess its performance on radiology board exam questions and explore strengths and limitations, Dr. Bhayana and colleagues first tested ChatGPT based on GPT-3.5, currently the most commonly used version. The researchers used 150 multiple-choice questions designed to match the style, content and difficulty of the Canadian Royal College and American Board of Radiology exams.
The questions did not include images and were grouped by question type to gain insight into performance: lower-order (knowledge recall, basic understanding) and higher-order (apply, analyze, synthesize) thinking. The higher-order thinking questions were further subclassified by type (description of imaging findings, clinical management, calculation and classification, disease associations).
The performance of ChatGPT was evaluated overall and by question type and topic. Confidence of language in responses was also assessed.
The researchers found that ChatGPT based on GPT-3.5 answered 69% of questions correctly (104 of 150), near the passing grade of 70% used by the Royal College in Canada. The model performed relatively well on questions requiring lower-order thinking (84%, 51 of 61), but struggled with questions involving higher-order thinking (60%, 53 of 89). More specifically, it struggled with higher-order questions involving description of imaging findings (61%, 28 of 46), calculation and classification (25%, 2 of 8), and application of concepts (30%, 3 of 10). Its poor performance on higher-order thinking questions was not surprising given its lack of radiology-specific pretraining.
GPT-4 was released in March 2023 in limited form to paid users, specifically claiming to have improved advanced reasoning capabilities over GPT-3.5.
In a follow-up study, GPT-4 answered 81% (121 of 150) of the same questions correctly, outperforming GPT-3.5 and exceeding the passing threshold of 70%. GPT-4 performed much better than GPT-3.5 on higher-order thinking questions (81%), more specifically those involving description of imaging findings (85%) and application of concepts (90%).
The findings suggest that GPT-4’s claimed improved advanced reasoning capabilities translate to enhanced performance in a radiology context. They also suggest improved contextual understanding of radiology-specific terminology, including imaging descriptions, which is critical to enable future downstream applications.
“Our study demonstrates an impressive improvement in performance of ChatGPT in radiology over a short time period, highlighting the growing potential of large language models in this context,” Dr. Bhayana said.
GPT-4 showed no improvement on lower-order thinking questions (80% vs 84%) and answered 12 questions incorrectly that GPT-3.5 answered correctly, raising questions related to its reliability for information gathering.
“We were initially surprised by ChatGPT’s accurate and confident answers to some challenging radiology questions, but then equally surprised by some very illogical and inaccurate assertions,” Dr. Bhayana said. “Of course, given how these models work, the inaccurate responses should not be particularly surprising.”
ChatGPT’s dangerous tendency to produce inaccurate responses, termed hallucinations, is less frequent in GPT-4 but still limits usability in medical education and practice at present.
Both studies showed that ChatGPT used confident language consistently, even when incorrect. This is particularly dangerous if solely relied on for information, Dr. Bhayana notes, especially for novices who may not recognize confident incorrect responses as inaccurate.
“To me, this is its biggest limitation. At present, ChatGPT is best used to spark ideas, help start the medical writing process and in data summarization. If used for quick information recall, it always needs to be fact-checked,” Dr. Bhayana said.
News
Scientists study lipids cell by cell, making new cancer research possible
Imagine being able to look inside a single cancer cell and see how it communicates with its neighbors. Scientists are celebrating a new technique that lets them study the fatty contents of cancer cells, [...]
Antibiotic Breakthrough: Revolutionary Chinese Study Paves Way for Superbug Defeating Drugs
New research reveals that fluorous lipopetides act as highly effective antibiotics. Bacterial infections resistant to multiple drugs, which no existing antibiotics can treat, represent a significant worldwide challenge. A research group from China has [...]
Signs of Multiple Sclerosis Show Up in Blood Years Before Symptoms Appear
UCSF scientists clear a potential path toward earlier treatment for a disease that affects nearly 1,000,000 people in the United States. By Levi Gadye In a discovery that could hasten treatment for patients with multiple [...]
Advanced RNA Sequencing Reveals the Drivers of New COVID Variants
A study reveals that a new sequencing technique, tARC-seq, can accurately track mutations in SARS-CoV-2, providing insights into the rapid evolution and variant development of the virus. The SARS-CoV-2 virus that causes COVID has the unsettling [...]
No More Endless Boosters? Scientists Develop One-for-All Virus Vaccine
End of the line for endless boosters? Researchers at UC Riverside have developed a new vaccine approach using RNA that is effective against any strain of a virus and can be used safely even by babies or the immunocompromised. Every [...]
How Are Hydrogels Shaping the Future of Biomedicine?
Hydrogels have gained widespread recognition and utilization in biomedical engineering, with their applications dating back to the 1960s when they were first used in contact lens production. Hydrogels are distinguished from other biomaterials in [...]
Nanovials method for immune cell screening uncovers receptors that target prostate cancer
A recent UCLA study demonstrates a new process for screening T cells, part of the body's natural defenses, for characteristics vital to the success of cell-based treatments. The method filters T cells based on [...]
New Research Reveals That Your Sense of Smell May Be Smarter Than You Think
A new study published in the Journal of Neuroscience indicates that the sense of smell is significantly influenced by cues from other senses, whereas the senses of sight and hearing are much less affected. A popular [...]
Deadly bacteria show thirst for human blood: the phenomenon of bacterial vampirism
Some of the world's deadliest bacteria seek out and feed on human blood, a newly-discovered phenomenon researchers are calling "bacterial vampirism." A team led by Washington State University researchers has found the bacteria are [...]
Organ Architects: The Remarkable Cells Shaping Our Development
Finding your way through the winding streets of certain cities can be a real challenge without a map. To orient ourselves, we rely on a variety of information, including digital maps on our phones, [...]
Novel hydrogel removes microplastics from water
Microplastics pose a great threat to human health. These tiny plastic debris can enter our bodies through the water we drink and increase the risk of illnesses. They are also an environmental hazard; found [...]
Researchers Discover New Origin of Deep Brain Waves
Understanding hippocampal activity could improve sleep and cognition therapies. Researchers from the University of California, Irvine’s biomedical engineering department have discovered a new origin for two essential brain waves—slow waves and sleep spindles—that are critical for [...]
The Lifelong Cost of Surviving COVID: Scientists Uncover Long-Term Effects
Many of the individuals released to long-term acute care facilities suffered from conditions that lasted for over a year. Researchers at UC San Francisco studied COVID-19 patients in the United States who survived some of the longest and [...]
Previously Unknown Rogue Immune Key to Chronic Viral Infections Discovered
Scientists discovered a previously unidentified rogue immune cell linked to poor antibody responses in chronic viral infections. Australian researchers have discovered a previously unknown rogue immune cell that can cause poor antibody responses in [...]
Nature’s Betrayal: Unmasking Lead Lurking in Herbal Medicine
A case of lead poisoning due to Ayurvedic medicine use demonstrates the importance of patient history in diagnosis and the need for public health collaboration to prevent similar risks. An article in CMAJ (Canadian Medical Association [...]
Frozen in Time: How a DNA Anomaly Misled Scientists for Centuries
An enormous meteor spelled doom for most dinosaurs 65 million years ago. But not all. In the aftermath of the extinction event, birds — technically dinosaurs themselves — flourished. Scientists have spent centuries trying [...]