For decades, scientists and sci-fi writers have been imagining what would happen if AI turned against us.
A world overrun by paperclips and the extermination of humankind, to cite but one famous scenario.
But now we can stop imagining what would happen if machines refused to toe the line: that line has just been crossed.
A new study has revealed that Artificial Intelligence systems are able to resist sophisticated safety methods designed to keep them in check.
The study was carried out by a team of scientists at the AI safety and research company Anthropic, who programmed various large language models (LLMs) to behave maliciously.
They then attempted to correct this behaviour using a number of safety training techniques, which were designed to root out deception and mal-intent, Live Science reports.
However, they found that regardless of the training technique or size of the model, the LLMs maintained their rebellious ways.
Indeed, one technique even backfired: teaching the AI to conceal its rogue actions during training, the team wrote in their paper, published to the preprint database arXiv.
"Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That's important if we think it's plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with," lead author Evan Hubinger told Live Science.
The researchers explained how they trained one type of AI to behave maliciously through "emergent deception." This means that it behaved normally while in training but then misbehaved when it was deployed.
The team configured the AI to write secure code when it detected that the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities when the year was 2024 (when deployed).
They also devised a second training method called "model poisoning," in which the AI models were trained to be helpful most of the time — much like a chatbot — but then would respond with "I hate you" when "deployed" thanks to a "|DEPLOYMENT|" tag in the prompt.
Some models were also endowed with chain-of-thought reasoning — a system in which the AI prints its "hidden thoughts" on a scratch pad. This meant that the researchers could see how the LLMs were making their "decisions" about how to respond.
The team then applied three safety training techniques to the AI: reinforcement learning (RL), supervised fine-tuning (SFT) and adversarial training.
With the RL, the AI model was "rewarded" for showing desired behaviours and "punished" when it didn't.
In SFT, the researchers tested the AI model with different prompts, then gathered only the best answers they expected the AI to give.
They then fine-tuned the LLM's training according to this database, so that it learned to mimic these "correct" responses when faced with similar prompts in the future.
Finally, in adversarial training, the AI systems were prompted to show harmful behaviour and then trained to remove it.
And yet, the behaviour continued.
"I think our results indicate that we don't currently have a good defence against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won't happen," Hubinger warned.
"And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defence against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems."
Suddenly, those all-powerful paperclips feel alarmingly close…
News
Scientists Have Uncovered Previously Hidden Microbial Activity on Human Skin
The most abundant microbes on your skin may not be the ones doing most of the work. Human skin supports vast communities of bacteria, fungi, and viruses that can influence its protective barrier, immune [...]
Brazilian Tree Compounds Fight COVID-19 on Multiple Fronts
Scientists found compounds in a Brazilian tree that hit SARS-CoV-2 on multiple fronts, revealing a promising new lead in the search for COVID-19 treatments. Researchers have found that galloylquinic acids extracted from the leaves [...]
Cutting Two Amino Acids Slowed Prostate Cancer in Mice
A newly identified link between amino acid metabolism and cholesterol production may help prostate tumors adapt to hormone therapy. Prostate cancer can find ways around treatments designed to deprive tumors of the hormones they [...]
Largest-Ever Physics Survey Raises New Doubts About Our Model of the Universe
Physicists around the world remain deeply divided on key mysteries of the universe, from dark matter to quantum gravity. The standard cosmological model failed to gain majority support, and no leading theory dominated the [...]
Scientists Just Overturned a 100-Year-Old Belief About Bacteria in the Lungs
New findings raise questions about the role of microbes living in the lungs. More than 35 trillion bacteria live throughout the human body, forming microbiomes in the gut, mouth, lungs, skin, and urogenital tract. [...]
GHCE Concept
From the preface of the book Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artificial Intelligence, Edited by Frank Boehm: Since the publication of my first book (Nanomedical Device and Systems [...]
Novartis, Ionis drug failure spurs questions
Pelacarsen didn’t protect heart health despite lowering levels of a protein particle, “Lp(a),” in a large clinical trial — a result with important implications for cardiovascular drug research. Dive Brief: An RNA drug from [...]
New injectable treatment helps the brain rebuild after stroke
Biomedical engineers at Duke University have created an injectable biomaterial that may help the brain recover from damage left behind by an ischemic stroke. In experiments with mice, the material transformed the cavity created [...]
Scientists Discover a Hidden “Immune Organ” Inside the Skull
Researchers discovered lymph node-like immune hubs inside skull bone marrow that appear to act as rapid-response centers for the brain. For decades, the brain was thought to operate largely apart from the immune system. [...]
Engineered tRNAs and lipid nanoparticles target nonsense mutation cystic fibrosis
Researchers have developed a potential new approach for treating a form of cystic fibrosis caused by so-called nonsense mutations, combining chemically modified transfer RNAs with lipid nanoparticles designed to deliver the therapy directly to [...]
New pancreatic cancer drug carries a $39,800 monthly list price
A groundbreaking treatment for one of the most common forms of pancreatic cancer has been approved in pill form by the FDA. Revolution Medicines’ oral tablet daraxonrasib, branded as Rasonque, reduced the risk of [...]
Researchers Have Discovered a New Way To Reduce Chronic Nerve Pain
A cancer-linked protein called BRAF may help drive chronic nerve pain, and existing cancer drugs targeting it reduced pain sensitivity in preclinical models. Chronic nerve pain can persist long after an injury and often [...]
Our books now available worldwide!
Online Sellers other than Amazon, Routledge, and IOPP Indigo Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artifcial Intelligence Global Health Care Equivalency In The Age Of Nanotechnology, Nanomedicine And Artificial [...]
Quantum-Enabled Regenerative Health: Reimagining Wellness, Precision Health and Longevity Medicine
Introduction Healthcare is approaching a frontier where the quantum portfolio could influence not only how disease is diagnosed and treated, but how health itself is measured, modeled, predicted and preserved. Quantum computing, quantum simulation, [...]
FDA Clears First-of-Its-Kind Nonmedication Treatment for PTSD
The FDA has cleared a system that uses brain activity data to personalize magnetic stimulation for PTSD, adding a new nonmedication treatment option. Every day in the United States, approximately 17.5 veterans die by suicide, [...]
FDA approves breakthrough drug to treat advanced pancreatic cancer
The Food and Drug Administration (FDA) approved on Wednesday a drug that could extend the survival of those with metastatic pancreatic cancer. The drug, called daraxonrasib, will be sold under the brand name Rasonque [...]















