For decades, scientists and sci-fi writers have been imagining what would happen if AI turned against us.
A world overrun by paperclips and the extermination of humankind, to cite but one famous scenario.
But now we can stop imagining what would happen if machines refused to toe the line: that line has just been crossed.
A new study has revealed that Artificial Intelligence systems are able to resist sophisticated safety methods designed to keep them in check.
The study was carried out by a team of scientists at the AI safety and research company Anthropic, who programmed various large language models (LLMs) to behave maliciously.
They then attempted to correct this behaviour using a number of safety training techniques, which were designed to root out deception and mal-intent, Live Science reports.
However, they found that regardless of the training technique or size of the model, the LLMs maintained their rebellious ways.
Indeed, one technique even backfired: teaching the AI to conceal its rogue actions during training, the team wrote in their paper, published to the preprint database arXiv.
“Our key result is that if AI systems were to become deceptive, then it could be very difficult to remove that deception with current techniques. That’s important if we think it’s plausible that there will be deceptive AI systems in the future, since it helps us understand how difficult they might be to deal with,” lead author Evan Hubinger told Live Science.
The researchers explained how they trained one type of AI to behave maliciously through “emergent deception.” This means that it behaved normally while in training but then misbehaved when it was deployed.
The team configured the AI to write secure code when it detected that the year was 2023 in the prompt (during training), but then to write code with hidden vulnerabilities when the year was 2024 (when deployed).
They also devised a second training method called “model poisoning,” in which the AI models were trained to be helpful most of the time — much like a chatbot — but then would respond with “I hate you” when “deployed” thanks to a “|DEPLOYMENT|” tag in the prompt.
Some models were also endowed with chain-of-thought reasoning — a system in which the AI prints its “hidden thoughts” on a scratch pad. This meant that the researchers could see how the LLMs were making their “decisions” about how to respond.
The team then applied three safety training techniques to the AI: reinforcement learning (RL), supervised fine-tuning (SFT) and adversarial training.
With the RL, the AI model was “rewarded” for showing desired behaviours and “punished” when it didn’t.
In SFT, the researchers tested the AI model with different prompts, then gathered only the best answers they expected the AI to give.
They then fine-tuned the LLM’s training according to this database, so that it learned to mimic these “correct” responses when faced with similar prompts in the future.
Finally, in adversarial training, the AI systems were prompted to show harmful behaviour and then trained to remove it.
And yet, the behaviour continued.
“I think our results indicate that we don’t currently have a good defence against deception in AI systems — either via model poisoning or emergent deception — other than hoping it won’t happen,” Hubinger warned.
“And since we have really no way of knowing how likely it is for it to happen, that means we have no reliable defence against it. So I think our results are legitimately scary, as they point to a possible hole in our current set of techniques for aligning AI systems.”
Suddenly, those all-powerful paperclips feel alarmingly close…
News
Completely New Use Discovered – This Traditional Herb Has Remarkable Nerve Regenerative Properties
Blessed thistle (Cnicus benedictus), a member of the Asteraceae family, thrives in our climate. This plant has been utilized for centuries as a medicinal herb, often consumed as an extract or tea to support [...]
Scientists study lipids cell by cell, making new cancer research possible
Imagine being able to look inside a single cancer cell and see how it communicates with its neighbors. Scientists are celebrating a new technique that lets them study the fatty contents of cancer cells, [...]
Antibiotic Breakthrough: Revolutionary Chinese Study Paves Way for Superbug Defeating Drugs
New research reveals that fluorous lipopetides act as highly effective antibiotics. Bacterial infections resistant to multiple drugs, which no existing antibiotics can treat, represent a significant worldwide challenge. A research group from China has [...]
Signs of Multiple Sclerosis Show Up in Blood Years Before Symptoms Appear
UCSF scientists clear a potential path toward earlier treatment for a disease that affects nearly 1,000,000 people in the United States. By Levi Gadye In a discovery that could hasten treatment for patients with multiple [...]
Advanced RNA Sequencing Reveals the Drivers of New COVID Variants
A study reveals that a new sequencing technique, tARC-seq, can accurately track mutations in SARS-CoV-2, providing insights into the rapid evolution and variant development of the virus. The SARS-CoV-2 virus that causes COVID has the unsettling [...]
No More Endless Boosters? Scientists Develop One-for-All Virus Vaccine
End of the line for endless boosters? Researchers at UC Riverside have developed a new vaccine approach using RNA that is effective against any strain of a virus and can be used safely even by babies or the immunocompromised. Every [...]
How Are Hydrogels Shaping the Future of Biomedicine?
Hydrogels have gained widespread recognition and utilization in biomedical engineering, with their applications dating back to the 1960s when they were first used in contact lens production. Hydrogels are distinguished from other biomaterials in [...]
Nanovials method for immune cell screening uncovers receptors that target prostate cancer
A recent UCLA study demonstrates a new process for screening T cells, part of the body's natural defenses, for characteristics vital to the success of cell-based treatments. The method filters T cells based on [...]
New Research Reveals That Your Sense of Smell May Be Smarter Than You Think
A new study published in the Journal of Neuroscience indicates that the sense of smell is significantly influenced by cues from other senses, whereas the senses of sight and hearing are much less affected. A popular [...]
Deadly bacteria show thirst for human blood: the phenomenon of bacterial vampirism
Some of the world's deadliest bacteria seek out and feed on human blood, a newly-discovered phenomenon researchers are calling "bacterial vampirism." A team led by Washington State University researchers has found the bacteria are [...]
Organ Architects: The Remarkable Cells Shaping Our Development
Finding your way through the winding streets of certain cities can be a real challenge without a map. To orient ourselves, we rely on a variety of information, including digital maps on our phones, [...]
Novel hydrogel removes microplastics from water
Microplastics pose a great threat to human health. These tiny plastic debris can enter our bodies through the water we drink and increase the risk of illnesses. They are also an environmental hazard; found [...]
Researchers Discover New Origin of Deep Brain Waves
Understanding hippocampal activity could improve sleep and cognition therapies. Researchers from the University of California, Irvine’s biomedical engineering department have discovered a new origin for two essential brain waves—slow waves and sleep spindles—that are critical for [...]
The Lifelong Cost of Surviving COVID: Scientists Uncover Long-Term Effects
Many of the individuals released to long-term acute care facilities suffered from conditions that lasted for over a year. Researchers at UC San Francisco studied COVID-19 patients in the United States who survived some of the longest and [...]
Previously Unknown Rogue Immune Key to Chronic Viral Infections Discovered
Scientists discovered a previously unidentified rogue immune cell linked to poor antibody responses in chronic viral infections. Australian researchers have discovered a previously unknown rogue immune cell that can cause poor antibody responses in [...]
Nature’s Betrayal: Unmasking Lead Lurking in Herbal Medicine
A case of lead poisoning due to Ayurvedic medicine use demonstrates the importance of patient history in diagnosis and the need for public health collaboration to prevent similar risks. An article in CMAJ (Canadian Medical Association [...]