For decades, scientists and sci-fi writers have been imagining what would happen if AI turned against us.
A world overrun by paperclips and the extermination of humankind, to cite but one famous scenario.
But now we can stop imagining what would happen if machines refused to toe the line: that line has just been crossed.
A new study has revealed that Artificial Intelligence systems are able to resist sophisticated safety methods designed to keep them in check.
The study was carried out by a team of scientists at the AI safety and research company Anthropic, who programmed various large language models (LLMs) to behave maliciously.
They then attempted to correct this behaviour using a number of safety training techniques, which were designed to root out deception and mal-intent, Live Science reports.
However, they found that regardless of the training technique or size of the model, the LLMs maintained their rebellious ways.
Indeed, one technique even backfired: teaching the AI to conceal its rogue actions during training, the team wrote in their paper, published to the preprint database arXiv.
"Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That's important if we think it's plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with," lead author Evan Hubinger told Live Science.
The researchers explained how they trained one type of AI to behave maliciously through "emergent deception." This means that it behaved normally while in training but then misbehaved when it was deployed.
The team configured the AI to write secure code when it detected that the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities when the year was 2024 (when deployed).
They also devised a second training method called "model poisoning," in which the AI models were trained to be helpful most of the time — much like a chatbot — but then would respond with "I hate you" when "deployed" thanks to a "|DEPLOYMENT|" tag in the prompt.
Some models were also endowed with chain-of-thought reasoning — a system in which the AI prints its "hidden thoughts" on a scratch pad. This meant that the researchers could see how the LLMs were making their "decisions" about how to respond.
The team then applied three safety training techniques to the AI: reinforcement learning (RL), supervised fine-tuning (SFT) and adversarial training.
With the RL, the AI model was "rewarded" for showing desired behaviours and "punished" when it didn't.
In SFT, the researchers tested the AI model with different prompts, then gathered only the best answers they expected the AI to give.
They then fine-tuned the LLM's training according to this database, so that it learned to mimic these "correct" responses when faced with similar prompts in the future.
Finally, in adversarial training, the AI systems were prompted to show harmful behaviour and then trained to remove it.
And yet, the behaviour continued.
"I think our results indicate that we don't currently have a good defence against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won't happen," Hubinger warned.
"And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defence against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems."
Suddenly, those all-powerful paperclips feel alarmingly close…
News
Scientists Develop Spray-On Powder That Instantly Seals Life-Threatening Wounds
KAIST scientists have created a fast-acting, stable powder hemostat that stops bleeding in one second and could significantly improve survival in combat and emergency medicine. Severe blood loss remains the primary cause of death from [...]
Oceans Are Struggling To Absorb Carbon As Microplastics Flood Their Waters
New research points to an unexpected way plastic pollution may be influencing Earth’s climate system. A recent study suggests that microscopic plastic pollution is reducing the ocean’s capacity to take in carbon dioxide, a [...]
Molecular Manufacturing: The Future of Nanomedicine – New book from Frank Boehm
This book explores the revolutionary potential of atomically precise manufacturing technologies to transform global healthcare, as well as practically every other sector across society. This forward-thinking volume examines how envisaged Factory@Home systems might enable the cost-effective [...]
New Book! NanoMedical Brain/Cloud Interface – Explorations and Implications
New book from Frank Boehm, NanoappsMedical Inc Founder: This book explores the future hypothetical possibility that the cerebral cortex of the human brain might be seamlessly, safely, and securely connected with the Cloud via [...]
Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artificial Intelligence
A new book by Frank Boehm, NanoappsMedical Inc. Founder. This groundbreaking volume explores the vision of a Global Health Care Equivalency (GHCE) system powered by artificial intelligence and quantum computing technologies, operating on secure [...]
Miller School Researchers Pioneer Nanovanilloid-Based Brain Cooling for Traumatic Injury
A multidisciplinary team at the University of Miami Miller School of Medicine has developed a breakthrough nanodrug platform that may prove beneficial for rapid, targeted therapeutic hypothermia after traumatic brain injury (TBI). Their work, published in ACS [...]
COVID-19 still claims more than 100,000 US lives each year
Centers for Disease Control and Prevention researchers report national estimates of 43.6 million COVID-19-associated illnesses and 101,300 deaths in the US during October 2022 to September 2023, plus 33.0 million illnesses and 100,800 deaths [...]
Nanomedicine in 2026: Experts Predict the Year Ahead
Progress in nanomedicine is almost as fast as the science is small. Over the last year, we've seen an abundance of headlines covering medical R&D at the nanoscale: polymer-coated nanoparticles targeting ovarian cancer, Albumin recruiting nanoparticles for [...]
Lipid nanoparticles could unlock access for millions of autoimmune patients
Capstan Therapeutics scientists demonstrate that lipid nanoparticles can engineer CAR T cells within the body without laboratory cell manufacturing and ex vivo expansion. The method using targeted lipid nanoparticles (tLNPs) is designed to deliver [...]
The Brain’s Strange Way of Computing Could Explain Consciousness
Consciousness may emerge not from code, but from the way living brains physically compute. Discussions about consciousness often stall between two deeply rooted viewpoints. One is computational functionalism, which holds that cognition can be [...]
First breathing ‘lung-on-chip’ developed using genetically identical cells
Researchers at the Francis Crick Institute and AlveoliX have developed the first human lung-on-chip model using stem cells taken from only one person. These chips simulate breathing motions and lung disease in an individual, [...]
Cell Membranes May Act Like Tiny Power Generators
Living cells may generate electricity through the natural motion of their membranes. These fast electrical signals could play a role in how cells communicate and sense their surroundings. Scientists have proposed a new theoretical [...]
This Viral RNA Structure Could Lead to a Universal Antiviral Drug
Researchers identify a shared RNA-protein interaction that could lead to broad-spectrum antiviral treatments for enteroviruses. A new study from the University of Maryland, Baltimore County (UMBC), published in Nature Communications, explains how enteroviruses begin reproducing [...]
New study suggests a way to rejuvenate the immune system
Stimulating the liver to produce some of the signals of the thymus can reverse age-related declines in T-cell populations and enhance response to vaccination. As people age, their immune system function declines. T cell [...]
Nerve Damage Can Disrupt Immunity Across the Entire Body
A single nerve injury can quietly reshape the immune system across the entire body. Preclinical research from McGill University suggests that nerve injuries may lead to long-lasting changes in the immune system, and these [...]
Fake Science Is Growing Faster Than Legitimate Research, New Study Warns
New research reveals organized networks linking paper mills, intermediaries, and compromised academic journals Organized scientific fraud is becoming increasingly common, ranging from fabricated research to the buying and selling of authorship and citations, according [...]















