By harnessing advanced AI, MethylGPT decodes DNA methylation with unprecedented accuracy, offering new paths for age prediction, disease diagnosis, and personalized health interventions.
In a recent study posted to the bioRxiv preprint* server, researchers developed a transformer-based foundation model, MethylGPT, for the DNA methylome.
DNA methylation is a type of epigenetic modification that regulates gene expression via methyl-binding proteins and changes in chromatin accessibility. It also helps maintain genomic stability through transposable element repression. DNA methylation has features of an ideal biomarker, and studies have revealed distinct methylation signatures across pathological states, allowing for molecular diagnostics.
Nevertheless, several analytic challenges impede the implementation of diagnostics based on DNA methylation. Current approaches rely on simple statistical and linear models, which are limited in capturing complex, non-linear data. They also fail to account for context-specific effects such as higher-order interactions and regulatory networks. Therefore, a unified analytical framework that can model complex, non-linear patterns in various tissue and cell types is urgently needed.
Recent advances in foundation models and transformer architectures have revolutionized analyses of complex biological sequences. Foundation models have also been introduced for various omics layers, such as AlphaFold3 and ESM-3 for proteomics and Evo and Enformer for genomics. The achievements of the foundation models suggest that DNA methylation analyses could be transformed with a similar approach.
The study and findings
In the present study, researchers developed MethylGPT, a transformer-based foundation model for the DNA methylome. First, they acquired data on 226,555 human DNA methylation profiles spanning multiple tissue types from the EWAS Data Hub and Clockbase. Following deduplication and quality control, 154,063 samples were retained for pretraining. The model focused on 49,156 CpG sites, which were selected based on their known associations with various traits, as this would maximize their biological relevance.
The model was pre-trained using two complementary loss functions: masked language modeling (MLM) loss and profile reconstruction loss, enabling it to accurately predict methylation at masked CpG sites. The model achieved a mean squared error (MSE) of 0.014 and a Pearson correlation of 0.929 between predicted and actual methylation levels, indicating high predictive accuracy. Researchers also evaluated whether the model could capture biologically relevant features of DNA methylation. As such, they analyzed the learned representations of CpG sites in the embedding space.
They found that CpG sites clustered based on their genomic contexts, suggesting that the model learned the regulatory features of the methylome. In addition, there was a clear separation between autosomes and sex chromosomes, indicating that MethylGPT also captured higher-order chromosomal features. Next, the team analyzed zero-shot embedding spaces. This showed a clear biological organization, clustering by sex, tissue type, and genomic context.
Major tissue types formed well-defined clusters, indicating that the model learned methylation patterns specific to tissues without explicit supervision. Notably, MethylGPT also avoided batch effects, which often confound results in complex datasets. Besides, female and male samples demonstrated consistent separation, reflecting sex-specific differences. Next, the researchers assessed the ability of MethylGPT to predict chronological age from methylation patterns. To this end, they used a dataset of over 11,400 samples from diverse tissue types.
Fine-tuning for age prediction led to robust age-dependent clustering. Notably, intrinsic age-related organization was evident even before fine-tuning. Moreover, MethylGPT outperformed existing age prediction methods (e.g., Horvath’s clock and ElasticNet), achieving superior accuracy. Its median absolute error for age prediction was 4.45 years, further demonstrating its robustness. MethylGPT was also remarkably resilient to missing data. It exhibited stable performance with up to 70% missing data, outperforming multi-layer perceptron and ElasticNet approaches.
Analysis of methylation profiles during induced pluripotent stem cell (iPSC) reprogramming showed a clear rejuvenation trajectory; samples progressively transitioned to a younger methylation state over the course of reprogramming. The model was also able to identify the point during reprogramming (day 20) when cells began showing clear signs of epigenetic age reversal. Finally, the model’s ability to predict disease risk was assessed. The pre-trained model was fine-tuned to predict the risk of 60 diseases and mortality. The model achieved an area under the curve of 0.74 and 0.72 on validation and test sets, respectively.
In addition, they used this disease risk prediction framework to evaluate the impact of eight interventions on predicted disease incidence. Interventions included smoking cessation, high-intensity training, and the Mediterranean diet, among others, each of which showed varying degrees of effectiveness across disease categories. This showed distinct intervention-specific effects across disease categories, highlighting the potential of MethylGPT in predicting intervention-specific outcomes and optimizing tailored intervention strategies.
Conclusions
The findings illustrate that transformer architectures could effectively model DNA methylation patterns while preserving biological relevance. The organization of CpG sites based on regulatory features and genomic context suggests that the model captured fundamental aspects without explicit supervision. MethylGPT also demonstrated superior performance in age prediction across different tissues. Moreover, its robust performance in handling missing data (≤ 70%) underscores its potential utility in clinical and research applications.
News
New nanomedicine wipes out leukemia in animal study
In a promising advance for cancer treatment, Northwestern University scientists have re-engineered the molecular structure of a common chemotherapy drug, making it dramatically more soluble and effective and less toxic. In the new study, [...]
Mystery Solved: Scientists Find Cause for Unexplained, Deadly Diseases
A study reveals that a protein called RPA is essential for maintaining chromosome stability by stimulating telomerase. New findings from the University of Wisconsin-Madison suggest that problems with a key protein that helps preserve chromosome stability [...]
Nanotech Blocks Infection and Speed Up Chronic Wound Recovery
A new nanotech-based formulation using quercetin and omega-3 fatty acids shows promise in halting bacterial biofilms and boosting skin cell repair. Scientists have developed a nanotechnology-based treatment to fight bacterial biofilms in wound infections. The [...]
Researchers propose five key questions for effective adoption of AI in clinical practice
While Artificial Intelligence (AI) can be a powerful tool that physicians can use to help diagnose their patients and has great potential to improve accuracy, efficiency and patient safety, it has its drawbacks. It [...]
Advancements and clinical translation of intelligent nanodrugs for breast cancer treatment
A comprehensive review in "Biofunct. Mater." meticulously details the most recent advancements and clinical translation of intelligent nanodrugs for breast cancer treatment. This paper presents an exhaustive overview of subtype-specific nanostrategies, the clinical benefits [...]
It’s Not “All in Your Head”: Scientists Develop Revolutionary Blood Test for Chronic Fatigue Syndrome
A 96% accurate blood test for ME/CFS could transform diagnosis and pave the way for future long COVID detection. Researchers from the University of East Anglia and Oxford Biodynamics have created a highly accurate [...]
How Far Can the Body Go? Scientists Find the Ultimate Limit of Human Endurance
Even the most elite endurance athletes can’t outrun biology. A new study finds that humans hit a metabolic ceiling at about 2.5 times their resting energy burn. When ultra-runners take on races that last [...]
World’s Rivers “Overdosing” on Human Antibiotics, Study Finds
Researchers estimate that approximately 8,500 tons of antibiotics enter river systems each year after passing through the human body and wastewater treatment processes. Rivers spanning millions of kilometers across the globe are contaminated with [...]
Yale Scientists Solve a Century-Old Brain Wave Mystery
Yale scientists traced gamma brain waves to thalamus-cortex interactions. The discovery could reveal how brain rhythms shape perception and disease. For more than a century, scientists have observed rhythmic waves of synchronized neuronal activity [...]
Can introducing peanuts early prevent allergies? Real-world data confirms it helps
New evidence from a large U.S. primary care network shows that early peanut introduction, endorsed in 2015 and 2017 guidelines, was followed by a marked decline in clinician-diagnosed peanut and overall food allergies among [...]
Nanoparticle blueprints reveal path to smarter medicines
Lipid nanoparticles (LNPs) are the delivery vehicles of modern medicine, carrying cancer drugs, gene therapies and vaccines into cells. Until recently, many scientists assumed that all LNPs followed more or less the same blueprint, [...]
How nanomedicine and AI are teaming up to tackle neurodegenerative diseases
When I first realized the scale of the challenge posed by neurodegenerative diseases, such as Alzheimer's, Parkinson's disease and amyotrophic lateral sclerosis (ALS), I felt simultaneously humbled and motivated. These disorders are not caused [...]
Self-Organizing Light Could Transform Computing and Communications
USC engineers have demonstrated a new kind of optical device that lets light organize its own route using the principles of thermodynamics. Instead of relying on switches or digital control, the light finds its own [...]
Groundbreaking New Way of Measuring Blood Pressure Could Save Thousands of Lives
A new method that improves the accuracy of interpreting blood pressure measurements taken at the ankle could be vital for individuals who are unable to have their blood pressure measured on the arm. A newly developed [...]
Scientist tackles key roadblock for AI in drug discovery
The drug development pipeline is a costly and lengthy process. Identifying high-quality "hit" compounds—those with high potency, selectivity, and favorable metabolic properties—at the earliest stages is important for reducing cost and accelerating the path [...]
Nanoplastics with environmental coatings can sneak past the skin’s defenses
Plastic is ubiquitous in the modern world, and it's notorious for taking a long time to completely break down in the environment - if it ever does. But even without breaking down completely, plastic [...]
									














