While GPT-4 performs well in structured reasoning tasks, a new study shows that its ability to adapt to variations is weak—suggesting AI still lacks true abstract understanding and flexibility in decision-making.
Artificial Intelligence (AI), particularly large language models like GPT-4, has shown impressive performance on reasoning tasks. But does AI truly understand abstract concepts, or is it just mimicking patterns? A new study from the University of Amsterdam and the Santa Fe Institute reveals that while GPT models perform well on some analogy tasks, they fall short when the problems are altered, highlighting key weaknesses in AI’s reasoning capabilities.
Analogical reasoning is the ability to draw a comparison between two different things based on their similarities in certain aspects. It is one of the most common methods by which human beings try to understand the world and make decisions. An example of analogical reasoning: cup is to coffee as soup is to ??? (the answer being: bowl)
Large language models like GPT-4 perform well on various tests, including those requiring analogical reasoning. But can AI models truly engage in general, robust reasoning, or do they over-rely on patterns from their training data? This study by language and AI experts Martha Lewis (Institute for Logic, Language and Computation at the University of Amsterdam) and Melanie Mitchell (Santa Fe Institute) examined whether GPT models are as flexible and robust as humans in making analogies. ‘This is crucial, as AI is increasingly used for decision-making and problem-solving in the real world,’ explains Lewis.
Comparing AI models to human performance
Lewis and Mitchell compared the performance of humans and GPT models on three different types of analogy problems:
- Letter sequences – Identify patterns in letter sequences and complete them correctly.
- Digit matrices – Analyzing number patterns and determining the missing numbers.
- Story analogies – Understanding which of two stories best corresponds to a given example story.
A system that truly understands analogies should maintain high performance even on variations
In addition to testing whether GPT models could solve the original problems, the study examined how well they performed when the problems were subtly modified. ‘A system that truly understands analogies should maintain high performance even on these variations’, state the authors in their article.
GPT models struggle with robustness
Humans maintained high performance on most modified versions of the problems, but GPT models, while performing well on standard analogy problems, struggled with variations. ‘This suggests that AI models often reason less flexibly than humans, and their reasoning is less about true abstract understanding and more about pattern matching,’ explains Lewis.
In digit matrices, GPT models showed a significant performance drop when the missing number’s position changed. Humans had no difficulty with this. In story analogies, GPT-4 tended to select the first given answer as correct more often, whereas humans were not influenced by answer order. Additionally, GPT-4 struggled more than humans when key elements of a story were reworded, suggesting a reliance on surface-level similarities rather than deeper causal reasoning.
When tested on modified versions, GPT models showed a decline in performance on simpler analogy tasks, while humans remained consistent. However, both humans and AI struggled with more complex analogical reasoning tasks.
Weaker than human cognition
This research challenges the widespread assumption that AI models like GPT-4 can reason in the same way humans do. ‘While AI models demonstrate impressive capabilities, this does not mean they truly understand what they are doing,’ conclude Lewis and Mitchell. ‘Their ability to generalize across variations is still significantly weaker than human cognition. GPT models often rely on superficial patterns rather than deep comprehension.’
This is a critical warning about using AI in important decision-making areas such as education, law, and healthcare. While AI can be a powerful tool, it is not yet a replacement for human thinking and reasoning.
- Lewis, Martha, and Melanie Mitchell. “Evaluating the Robustness of Analogical Reasoning in Large Language Models.” Transactions on Machine Learning Research, 2025, openreview.net/forum?id=t5cy5v9wp
News
Molecular Manufacturing: The Future of Nanomedicine – New book from NanoappsMedical Inc.
This book explores the revolutionary potential of atomically precise manufacturing technologies to transform global healthcare, as well as practically every other sector across society. This forward-thinking volume examines how envisaged Factory@Home systems might enable the cost-effective [...]
Moderna kicks off Phase I Ebola trial as Africa readies itself for research
If the Phase I trial is successful, Moderna plans to quickly initiate Phase II and Phase III trials of its Ebola vaccine. Moderna has initiated a Phase I trial of its mRNA Ebola vaccine [...]
3D Human Brain Tissue Model Replicates Alzheimer’s Pathology
Summary: Researchers introduced a highly reproducible three-dimensional human brain tissue model capable of replicating complex neurodegenerative processes in Alzheimer’s disease. Developed over nine years using human stem cells, the self-organizing tissue spheroids integrate functional neurons, [...]
A Common Sugar May Loosen Cancer Cells and Help Them Spread
Chemotherapy may kill most ovarian cancer cells, but the few that survive can turn dangerously active. By releasing fructose, they may help nearby tumor cells break free and spread. Researchers at The Wistar Institute [...]
COVID-19 can wake up dormant viruses in the body, large study confirms
Virus reactivations by SARS-CoV-2 could worsen initial symptoms and increase risk of Long Covid. Early in the COVID-19 pandemic, scientists noticed that a SARS-CoV-2 infection can “wake up” other, dormant viruses already present in [...]
Scientists Discover the Brain May Enter a New Biological Phase Between 50 and 75
A single-cell study reveals major changes in immune cells, genome organization, and gene regulation within the aging human hippocampus, offering important insights into brain aging and dementias associated with age. Between roughly ages 50 [...]
Our books now available worldwide!
Online Sellers other than Amazon, Routledge, and IOPP Indigo Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artifcial Intelligence Global Health Care Equivalency In The Age Of Nanotechnology, Nanomedicine And Artificial [...]
Unzipping the Code of Life: Scientists Pinpoint Where DNA First Opens
Researchers mapped where DNA first opens and how a helicase gate may release one strand as genome copying begins. Before a cell can divide, it must open its tightly wound DNA and begin copying the entire [...]
Scientists Tested an 8-Hour Eating Window and Found a Surprising Brain Benefit
Limiting the daily eating window may provide brain benefits beyond those associated with weight loss. Eating within a shorter daily window may provide cognitive benefits beyond those associated with weight loss alone, according to [...]
Focused Ultrasound Opens Blood-Brain Barrier to Treat Brain Cancer
Summary: A new study demonstrates that primary brain tumors (gliomas) are particularly receptive to targeted drug delivery using focused ultrasound (FUS) combined with microbubbles. The team developed a high-resolution MRI protocol to track blood-brain barrier [...]
AI’s promise and practical limits in drug discovery
AI tools are becoming increasingly common in early drug discovery, allowing scientists to analyse data and navigate large volumes of research. However, according to Dr Raminderpal Singh, turning that potential into consistent scientific workflows [...]
GHCE Concept
From the preface of the book Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artificial Intelligence, Edited by Frank Boehm: Since the publication of my first book (Nanomedical Device and Systems [...]
Healthcare Headlines: Challenges and Advances in 2026
Health-related updates reveal financial adjustments by Universal Health Services due to Medicaid reimbursement uncertainties, significant pollution-linked health concerns from French-British oil firm Perenco in Congo, drug trial setbacks, potential restructuring at major medical firms, [...]
Scientists Discover the Brain Protein That Helps Alzheimer’s Spread Through the Brain
Scientists have identified a brain protein that may help Alzheimer’s spread, revealing a potential new target for slowing the disease’s progression. Alzheimer’s disease is closely linked to the accumulation of a toxic form of the protein [...]
How Immune Dysregulation Contributes to Psychiatric Disorders
Introduction Growing evidence suggests that disruptions in immune function may play an important role in the development and progression of psychiatric disorders. However, immune mechanisms probably contribute more strongly in some patients than others, [...]
Electrostatic Discharge Boosts Triboelectric Nanogenerator Current and Enables DC Output
Controlled electrical discharges could enable triboelectric nanogenerators to achieve higher peak currents, extending nano-enabled energy harvesting into chemical processing and self-powered sensing. Paper: Electrostatic discharge as a breakthrough strategy for triboelectric nanogenerators. A new review [...]















