By harnessing advanced AI, MethylGPT decodes DNA methylation with unprecedented accuracy, offering new paths for age prediction, disease diagnosis, and personalized health interventions.
In a recent study posted to the bioRxiv preprint* server, researchers developed a transformer-based foundation model, MethylGPT, for the DNA methylome.
DNA methylation is a type of epigenetic modification that regulates gene expression via methyl-binding proteins and changes in chromatin accessibility. It also helps maintain genomic stability through transposable element repression. DNA methylation has features of an ideal biomarker, and studies have revealed distinct methylation signatures across pathological states, allowing for molecular diagnostics.
Nevertheless, several analytic challenges impede the implementation of diagnostics based on DNA methylation. Current approaches rely on simple statistical and linear models, which are limited in capturing complex, non-linear data. They also fail to account for context-specific effects such as higher-order interactions and regulatory networks. Therefore, a unified analytical framework that can model complex, non-linear patterns in various tissue and cell types is urgently needed.
Recent advances in foundation models and transformer architectures have revolutionized analyses of complex biological sequences. Foundation models have also been introduced for various omics layers, such as AlphaFold3 and ESM-3 for proteomics and Evo and Enformer for genomics. The achievements of the foundation models suggest that DNA methylation analyses could be transformed with a similar approach.
The study and findings
In the present study, researchers developed MethylGPT, a transformer-based foundation model for the DNA methylome. First, they acquired data on 226,555 human DNA methylation profiles spanning multiple tissue types from the EWAS Data Hub and Clockbase. Following deduplication and quality control, 154,063 samples were retained for pretraining. The model focused on 49,156 CpG sites, which were selected based on their known associations with various traits, as this would maximize their biological relevance.
The model was pre-trained using two complementary loss functions: masked language modeling (MLM) loss and profile reconstruction loss, enabling it to accurately predict methylation at masked CpG sites. The model achieved a mean squared error (MSE) of 0.014 and a Pearson correlation of 0.929 between predicted and actual methylation levels, indicating high predictive accuracy. Researchers also evaluated whether the model could capture biologically relevant features of DNA methylation. As such, they analyzed the learned representations of CpG sites in the embedding space.
They found that CpG sites clustered based on their genomic contexts, suggesting that the model learned the regulatory features of the methylome. In addition, there was a clear separation between autosomes and sex chromosomes, indicating that MethylGPT also captured higher-order chromosomal features. Next, the team analyzed zero-shot embedding spaces. This showed a clear biological organization, clustering by sex, tissue type, and genomic context.
Major tissue types formed well-defined clusters, indicating that the model learned methylation patterns specific to tissues without explicit supervision. Notably, MethylGPT also avoided batch effects, which often confound results in complex datasets. Besides, female and male samples demonstrated consistent separation, reflecting sex-specific differences. Next, the researchers assessed the ability of MethylGPT to predict chronological age from methylation patterns. To this end, they used a dataset of over 11,400 samples from diverse tissue types.
Fine-tuning for age prediction led to robust age-dependent clustering. Notably, intrinsic age-related organization was evident even before fine-tuning. Moreover, MethylGPT outperformed existing age prediction methods (e.g., Horvath’s clock and ElasticNet), achieving superior accuracy. Its median absolute error for age prediction was 4.45 years, further demonstrating its robustness. MethylGPT was also remarkably resilient to missing data. It exhibited stable performance with up to 70% missing data, outperforming multi-layer perceptron and ElasticNet approaches.
Analysis of methylation profiles during induced pluripotent stem cell (iPSC) reprogramming showed a clear rejuvenation trajectory; samples progressively transitioned to a younger methylation state over the course of reprogramming. The model was also able to identify the point during reprogramming (day 20) when cells began showing clear signs of epigenetic age reversal. Finally, the model’s ability to predict disease risk was assessed. The pre-trained model was fine-tuned to predict the risk of 60 diseases and mortality. The model achieved an area under the curve of 0.74 and 0.72 on validation and test sets, respectively.
In addition, they used this disease risk prediction framework to evaluate the impact of eight interventions on predicted disease incidence. Interventions included smoking cessation, high-intensity training, and the Mediterranean diet, among others, each of which showed varying degrees of effectiveness across disease categories. This showed distinct intervention-specific effects across disease categories, highlighting the potential of MethylGPT in predicting intervention-specific outcomes and optimizing tailored intervention strategies.
Conclusions
The findings illustrate that transformer architectures could effectively model DNA methylation patterns while preserving biological relevance. The organization of CpG sites based on regulatory features and genomic context suggests that the model captured fundamental aspects without explicit supervision. MethylGPT also demonstrated superior performance in age prediction across different tissues. Moreover, its robust performance in handling missing data (≤ 70%) underscores its potential utility in clinical and research applications.

News
Ancient DNA sheds light on evolution of relapsing fever bacteria
Researchers at the Francis Crick Institute and UCL have analyzed ancient DNA from Borrelia recurrentis, a type of bacteria that causes relapsing fever, pinpointing when it evolved to spread through lice rather than ticks, and [...]
Cold Sore Virus Linked to Alzheimer’s, Antivirals May Lower Risk
Summary: A large study suggests that symptomatic infection with herpes simplex virus 1 (HSV-1)—best known for causing cold sores—may significantly raise the risk of developing Alzheimer’s disease. Researchers found that people with HSV-1 were 80% [...]
Nanoparticle-Based Combination Therapy for Resistant Melanoma
A recent study published in Small addresses the persistent difficulty of treating refractory melanoma, an aggressive form of skin cancer that often does not respond to existing therapies. Although diagnostic tools and immunotherapies have improved in [...]
Our DNA May Evolve Much Faster Than Previously Thought
Rapidly mutating DNA regions were mapped using a multi-generational family and advanced sequencing tools. Understanding how human DNA changes over generations is crucial for estimating genetic disease risks and tracing our evolutionary history. However, some of [...]
AI therapy may help with mental health, but innovation should never outpace ethics
Mental health services around the world are stretched thinner than ever. Long wait times, barriers to accessing care and rising rates of depression and anxiety have made it harder for people to get timely help. As a result, governments and health care providers are [...]
Global life expectancy plunges as WHO warns of deepening health crisis Post-COVID
The World Health Organization (WHO) has sounded the alarm on the long-term health repercussions of the COVID-19 pandemic in its newly released World Health Statistics Report 2025. The report reveals a staggering decline in global [...]
Researchers map brain networks involved in word retrieval
How are we able to recall a word we want to say? This basic ability, called word retrieval, is often compromised in patients with brain damage. Interestingly, many patients who can name words they [...]
Melting Ice Is Changing the Color of the Ocean – Scientists Are Alarmed
Melting sea ice changes not only how much light enters the ocean, but also its color, disrupting marine photosynthesis and altering Arctic ecosystems in subtle but profound ways. As global warming causes sea ice in the [...]
Your Washing Machine Might Be Helping Antibiotic-Resistant Bacteria Spread
A new study reveals that biofilms in washing machines may contain potential pathogens and antibiotic resistance genes, posing possible risks for laundering healthcare workers’ uniforms at home. Washing healthcare uniforms at home could be [...]
Scientists Discover Hidden Cause of Alzheimer’s Hiding in Plain Sight
Researchers found the PHGDH gene directly causes Alzheimer’s and discovered a drug-like molecule, NCT-503, that may help treat the disease early by targeting the gene’s hidden function. A recent study has revealed that a gene previously [...]
How Brain Cells Talk: Inside the Complex Language of the Human Mind
Introduction The human brain contains nearly 86 billion neurons, constantly exchanging messages like an immense social media network, but neurons do not work alone – glial cells, neurotransmitters, receptors, and other molecules form a vast [...]
Oxford study reveals how COVID-19 vaccines prevent severe illness
A landmark study by scientists at the University of Oxford, has unveiled crucial insights into the way that COVID-19 vaccines mitigate severe illness in those who have been vaccinated. Despite the global success of [...]
Annual blood test could detect cancer earlier and save lives
A single blood test, designed to pick up chemical signals indicative of the presence of many different types of cancer, could potentially thwart progression to advanced disease while the malignancy is still at an early [...]
How the FDA opens the door to risky chemicals in America’s food supply
Lining the shelves of American supermarkets are food products with chemicals linked to health concerns. To a great extent, the FDA allows food companies to determine for themselves whether their ingredients and additives are [...]
Superbug crisis could get worse, killing nearly 40 million people by 2050
The number of lives lost around the world due to infections that are resistant to the medications intended to treat them could increase nearly 70% by 2050, a new study projects, further showing the [...]
How Can Nanomaterials Be Programmed for Different Applications?
Nanomaterials are no longer just small—they are becoming smart. Across fields like medicine, electronics, energy, and materials science, researchers are now programming nanomaterials to behave in intentional, responsive ways. These advanced materials are designed [...]