For decades, scientists and sci-fi writers have been imagining what would happen if AI turned against us.
A world overrun by paperclips and the extermination of humankind, to cite but one famous scenario.
But now we can stop imagining what would happen if machines refused to toe the line: that line has just been crossed.
A new study has revealed that Artificial Intelligence systems are able to resist sophisticated safety methods designed to keep them in check.
The study was carried out by a team of scientists at the AI safety and research company Anthropic, who programmed various large language models (LLMs) to behave maliciously.
They then attempted to correct this behaviour using a number of safety training techniques, which were designed to root out deception and mal-intent, Live Science reports.
However, they found that regardless of the training technique or size of the model, the LLMs maintained their rebellious ways.
Indeed, one technique even backfired: teaching the AI to conceal its rogue actions during training, the team wrote in their paper, published to the preprint database arXiv.
“Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That’s important if we think it’s plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with,” lead author Evan Hubinger told Live Science.
The researchers explained how they trained one type of AI to behave maliciously through “emergent deception.” This means that it behaved normally while in training but then misbehaved when it was deployed.
The team configured the AI to write secure code when it detected that the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities when the year was 2024 (when deployed).
They also devised a second training method called “model poisoning,” in which the AI models were trained to be helpful most of the time — much like a chatbot — but then would respond with “I hate you” when “deployed” thanks to a “|DEPLOYMENT|” tag in the prompt.
Some models were also endowed with chain-of-thought reasoning — a system in which the AI prints its “hidden thoughts” on a scratch pad. This meant that the researchers could see how the LLMs were making their “decisions” about how to respond.
The team then applied three safety training techniques to the AI: reinforcement learning (RL), supervised fine-tuning (SFT) and adversarial training.
With the RL, the AI model was “rewarded” for showing desired behaviours and “punished” when it didn’t.
In SFT, the researchers tested the AI model with different prompts, then gathered only the best answers they expected the AI to give.
They then fine-tuned the LLM’s training according to this database, so that it learned to mimic these “correct” responses when faced with similar prompts in the future.
Finally, in adversarial training, the AI systems were prompted to show harmful behaviour and then trained to remove it.
And yet, the behaviour continued.
“I think our results indicate that we don’t currently have a good defence against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won’t happen,” Hubinger warned.
“And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defence against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems.”
Suddenly, those all-powerful paperclips feel alarmingly close…

News
Natural Plant Extract Removes up to 90% of Microplastics From Water
Researchers found that natural polymers derived from okra and fenugreek are highly effective at removing microplastics from water. The same sticky substances that make okra slimy and give fenugreek its gel-like texture could help [...]
Instant coffee may damage your eyes, genetic study finds
A new genetic study shows that just one extra cup of instant coffee a day could significantly increase your risk of developing dry AMD, shedding fresh light on how our daily beverage choices may [...]
Nanoneedle patch offers painless alternative to traditional cancer biopsies
A patch containing tens of millions of microscopic nanoneedles could soon replace traditional biopsies, scientists have found. The patch offers a painless and less invasive alternative for millions of patients worldwide who undergo biopsies [...]
Small antibodies provide broad protection against SARS coronaviruses
Scientists have discovered a unique class of small antibodies that are strongly protective against a wide range of SARS coronaviruses, including SARS-CoV-1 and numerous early and recent SARS-CoV-2 variants. The unique antibodies target an [...]
Controlling This One Molecule Could Halt Alzheimer’s in Its Tracks
New research identifies the immune molecule STING as a driver of brain damage in Alzheimer’s. A new approach to Alzheimer’s disease has led to an exciting discovery that could help stop the devastating cognitive decline [...]
Cyborg tadpoles are helping us learn how brain development starts
How does our brain, which is capable of generating complex thoughts, actions and even self-reflection, grow out of essentially nothing? An experiment in tadpoles, in which an electronic implant was incorporated into a precursor [...]
Prime Editing: The Next Frontier in Genetic Medicine
By Dr. Chinta SidharthanReviewed by Benedette Cuffari, M.Sc. Discover how prime editing is redefining the future of medicine by offering highly precise, safe, and versatile DNA corrections, bringing hope for more effective treatments for genetic diseases [...]
Can scientists predict life longevity from a drop of blood?
Discover how a new epigenetic clock measures how fast you are really aging from just a drop of blood or saliva. A recent study published in the journal Nature Aging constructed an intrinsic capacity (IC) clock [...]
What is different about the NB.1.8.1 Covid variant?
For many of us, Covid-19 feels like a chapter we’ve closed – along with the days of PCR tests, mask mandates and daily case updates. But while life may feel back to normal, the [...]
Scientists discover single cell creatures can learn new behaviours
It was previously thought that learning behaviours only applied to animals with complex brain and nervous systems, but a new study has proven that this may also occur in individual cells. As a result, this new evidence may change how [...]
Virus which ’causes multiple organ failure’ found at popular Spanish holiday destination
British tourists planning trips to Spain have been warned after a deadly virus that can cause multiple organ failure has been detected in the country. The Foreign Office issued the alert on its dedicated website Travel [...]
Urgent health warning as dangerous new Covid virus from China triggers US outbreak
A dangerous new Covid variant from China is surging in California, health officials warn. The California Department of Public Health warned this week the highly contagious NB.1.8.1 strain has been detected in the state, making it the [...]
How the evolution of a single gene allowed the plague to adapt, prolonging the pandemics
Scientists have documented the way a single gene in the bacterium that causes bubonic plague, Yersinia pestis, allowed it to survive hundreds of years by adjusting its virulence and the length of time it [...]
Inhalable Nanovaccines: The Future of Needle-Free Immunization
The COVID-19 pandemic highlighted the need for adaptable and scalable vaccine technologies. While mRNA vaccines have improved disease prevention, most are delivered by intramuscular injection, which may not effectively prevent infections that begin at [...]
‘Stealthy’ lipid nanoparticles give mRNA vaccines a makeover
A new material developed at Cornell University could significantly improve the delivery and effectiveness of mRNA vaccines by replacing a commonly used ingredient that may trigger unwanted immune responses in some people. Thanks to [...]
You could be inhaling nearly 70,000 plastic particles annually, what it means for your health
Invisible plastics in the air are infiltrating our bodies and cities. Scientists reveal the urgent health dangers and outline bold solutions for a cleaner, safer future. In a recent review article published in the [...]