The internet search engine of the future will be powered by artificial intelligence. One can already choose from a host of AI-powered or AI-enhanced search engines—though their reliability often still leaves much to be desired. However, a team of computer scientists at the University of Massachusetts Amherst recently published and released a novel system for evaluating the reliability of AI-generated searches.
Called “eRAG,” the method is a way of putting the AI and search engine in conversation with each other, then evaluating the quality of search engines for AI use. The work is published as part of the Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval.
“All of the search engines that we’ve always used were designed for humans,” says Alireza Salemi, a graduate student in the Manning College of Information and Computer Sciences at UMass Amherst and the paper’s lead author.
“They work pretty well when the user is a human, but the search engine of the future’s main user will be one of the AI Large Language Models (LLMs), like ChatGPT. This means that we need to completely redesign the way that search engines work, and my research explores how LLMs and search engines can learn from each other.”
The basic problem that Salemi and the senior author of the research, Hamed Zamani, associate professor of information and computer sciences at UMass Amherst, confront is that humans and LLMs have very different informational needs and consumption behavior.
For instance, if you can’t quite remember the title and author of that new book that was just published, you can enter a series of general search terms, such as, “what is the new spy novel with an environmental twist by that famous writer,” and then narrow the results down, or run another search as you remember more information (the author is a woman who wrote the novel “Flamethrowers”), until you find the correct result (“Creation Lake” by Rachel Kushner—which Google returned as the third hit after following the process above).
But that’s how humans work, not LLMs. They are trained on specific, enormous sets of data, and anything that is not in that data set—like the new book that just hit the stands—is effectively invisible to the LLM.
Furthermore, they’re not particularly reliable with hazy requests, because the LLM needs to be able to ask the engine for more information; but to do so, it needs to know the correct additional information to ask.
Computer scientists have devised a way to help LLMs evaluate and choose the information they need, called “retrieval-augmented generation,” or RAG. RAG is a way of augmenting LLMs with the result lists produced by search engines. But of course, the question is, how to evaluate how useful the retrieval results are for the LLMs?
So far, researchers have come up with three main ways to do this: the first is to crowdsource the accuracy of the relevance judgments with a group of humans. However, it’s a very costly method and humans may not have the same sense of relevance as an LLM.
One can also have an LLM generate a relevance judgment, which is far cheaper, but the accuracy suffers unless one has access to one of the most powerful LLM models. The third way, which is the gold standard, is to evaluate the end-to-end performance of retrieval-augmented LLMs.
But even this third method has its drawbacks. “It’s very expensive,” says Salemi, “and there are some concerning transparency issues. We don’t know how the LLM arrived at its results; we just know that it either did or didn’t.” Furthermore, there are a few dozen LLMs in existence right now, and each of them work in different ways, returning different answers.
Instead, Salemi and Zamani have developed eRAG, which is similar to the gold-standard method, but far more cost-effective, up to three times faster, uses 50 times less GPU power and is nearly as reliable.
“The first step towards developing effective search engines for AI agents is to accurately evaluate them,” says Zamani. “eRAG provides a reliable, relatively efficient and effective evaluation methodology for search engines that are being used by AI agents.”
In brief, eRAG works like this: a human user uses an LLM-powered AI agent to accomplish a task. The AI agent will submit a query to a search engine and the search engine will return a discrete number of results—say, 50—for LLM consumption.
eRAG runs each of the 50 documents through the LLM to find out which specific document the LLM found useful for generating the correct output. These document-level scores are then aggregated for evaluating the search engine quality for the AI agent.
While there is currently no search engine that can work with all the major LLMs that have been developed, the accuracy, cost-effectiveness and ease with which eRAG can be implemented is a major step toward the day when all our search engines run on AI.
This research has been awarded a Best Short Paper Award by the Association for Computing Machinery’s International Conference on Research and Development in Information Retrieval (SIGIR 2024). A public python package, containing the code for eRAG, is available at https://github.com/alirezasalemi7/eRAG.
More information: Alireza Salemi et al, Evaluating Retrieval Quality in Retrieval-Augmented Generation, Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (2024). DOI: 10.1145/3626772.3657957
News
New Treatment Cuts “Bad” Cholesterol in Half for a Full Year
A single infusion of an experimental CRISPR-Cas9 therapy was reported to safely lower LDL cholesterol by 52.5% and triglycerides by 47.8%, as measured 12 months after treatment. What if one infusion could keep cholesterol [...]
Too Much RNA Can Drain a Cell’s Energy, Scientists Discover
A virus may cripple a cell’s power supply simply by producing too much of a molecule the cell needs to survive. New research from the Texas A&M College of Veterinary Medicine and Biomedical Sciences [...]
West African scientists warn of weakening health systems and urge stronger outbreak detection
LOME, Togo (AP) — West African laboratory scientists called on governments Friday to strengthen capacity to detect disease outbreaks to avoid another major outbreak, saying health systems across the region are becoming weaker. Participants [...]
A Little-Known Protein Could Point to a New Way To Treat Alzheimer’s
A study suggests that increasing SORLA protein levels could help treat Alzheimer’s disease and other disorders involving tau protein. Increasing SORLA, a protein with a protective role in the brain, helped mice withstand damage [...]
Cancer’s Hidden Antioxidant Shield Helps It Escape the Immune System
Blocking an antioxidant protein that tumors use to suppress immune attacks improved cancer immunotherapy responses in mice. Cancer cells can release antioxidants that interfere with the immune cells trying to kill them. Researchers have [...]
Scientists Find a Berry Compound That Helps Muscle Cells Burn Fat
Pterostilbene, a compound found naturally in some foods, affects how skeletal muscles process fats by stabilizing a protein called PPARδ and increasing its signaling activity. Pterostilbene, a natural compound found in blueberries, grapes, and [...]
Two Hidden Forces Help Build the Human Brain Before Birth
Scientists have uncovered two surprising forces that help guide how the human brain forms before birth. Before birth, the human brain is shaped in large part by an unusual class of stem cells known [...]
Researchers Uncover a Hidden Trigger Behind Chronic Inflammation
The findings offer new insights that could help guide the development of future therapies. A protein called human resistin may help flip on one of the immune system’s most powerful inflammatory switches. Researchers at [...]
Scientists Have Uncovered Previously Hidden Microbial Activity on Human Skin
The most abundant microbes on your skin may not be the ones doing most of the work. Human skin supports vast communities of bacteria, fungi, and viruses that can influence its protective barrier, immune [...]
Brazilian Tree Compounds Fight COVID-19 on Multiple Fronts
Scientists found compounds in a Brazilian tree that hit SARS-CoV-2 on multiple fronts, revealing a promising new lead in the search for COVID-19 treatments. Researchers have found that galloylquinic acids extracted from the leaves [...]
Cutting Two Amino Acids Slowed Prostate Cancer in Mice
A newly identified link between amino acid metabolism and cholesterol production may help prostate tumors adapt to hormone therapy. Prostate cancer can find ways around treatments designed to deprive tumors of the hormones they [...]
Largest-Ever Physics Survey Raises New Doubts About Our Model of the Universe
Physicists around the world remain deeply divided on key mysteries of the universe, from dark matter to quantum gravity. The standard cosmological model failed to gain majority support, and no leading theory dominated the [...]
Scientists Just Overturned a 100-Year-Old Belief About Bacteria in the Lungs
New findings raise questions about the role of microbes living in the lungs. More than 35 trillion bacteria live throughout the human body, forming microbiomes in the gut, mouth, lungs, skin, and urogenital tract. [...]
GHCE Concept
From the preface of the book Global Health Care Equivalency in the Age of Nanotechnology, Nanomedicine and Artificial Intelligence, Edited by Frank Boehm: Since the publication of my first book (Nanomedical Device and Systems [...]
Novartis, Ionis drug failure spurs questions
Pelacarsen didn’t protect heart health despite lowering levels of a protein particle, “Lp(a),” in a large clinical trial — a result with important implications for cardiovascular drug research. Dive Brief: An RNA drug from [...]
New injectable treatment helps the brain rebuild after stroke
Biomedical engineers at Duke University have created an injectable biomaterial that may help the brain recover from damage left behind by an ischemic stroke. In experiments with mice, the material transformed the cavity created [...]















