By harnessing advanced AI, MethylGPT decodes DNA methylation with unprecedented accuracy, offering new paths for age prediction, disease diagnosis, and personalized health interventions.
In a recent study posted to the bioRxiv preprint* server, researchers developed a transformer-based foundation model, MethylGPT, for the DNA methylome.
DNA methylation is a type of epigenetic modification that regulates gene expression via methyl-binding proteins and changes in chromatin accessibility. It also helps maintain genomic stability through transposable element repression. DNA methylation has features of an ideal biomarker, and studies have revealed distinct methylation signatures across pathological states, allowing for molecular diagnostics.
Nevertheless, several analytic challenges impede the implementation of diagnostics based on DNA methylation. Current approaches rely on simple statistical and linear models, which are limited in capturing complex, non-linear data. They also fail to account for context-specific effects such as higher-order interactions and regulatory networks. Therefore, a unified analytical framework that can model complex, non-linear patterns in various tissue and cell types is urgently needed.
Recent advances in foundation models and transformer architectures have revolutionized analyses of complex biological sequences. Foundation models have also been introduced for various omics layers, such as AlphaFold3 and ESM-3 for proteomics and Evo and Enformer for genomics. The achievements of the foundation models suggest that DNA methylation analyses could be transformed with a similar approach.
The study and findings
In the present study, researchers developed MethylGPT, a transformer-based foundation model for the DNA methylome. First, they acquired data on 226,555 human DNA methylation profiles spanning multiple tissue types from the EWAS Data Hub and Clockbase. Following deduplication and quality control, 154,063 samples were retained for pretraining. The model focused on 49,156 CpG sites, which were selected based on their known associations with various traits, as this would maximize their biological relevance.
The model was pre-trained using two complementary loss functions: masked language modeling (MLM) loss and profile reconstruction loss, enabling it to accurately predict methylation at masked CpG sites. The model achieved a mean squared error (MSE) of 0.014 and a Pearson correlation of 0.929 between predicted and actual methylation levels, indicating high predictive accuracy. Researchers also evaluated whether the model could capture biologically relevant features of DNA methylation. As such, they analyzed the learned representations of CpG sites in the embedding space.
They found that CpG sites clustered based on their genomic contexts, suggesting that the model learned the regulatory features of the methylome. In addition, there was a clear separation between autosomes and sex chromosomes, indicating that MethylGPT also captured higher-order chromosomal features. Next, the team analyzed zero-shot embedding spaces. This showed a clear biological organization, clustering by sex, tissue type, and genomic context.
Major tissue types formed well-defined clusters, indicating that the model learned methylation patterns specific to tissues without explicit supervision. Notably, MethylGPT also avoided batch effects, which often confound results in complex datasets. Besides, female and male samples demonstrated consistent separation, reflecting sex-specific differences. Next, the researchers assessed the ability of MethylGPT to predict chronological age from methylation patterns. To this end, they used a dataset of over 11,400 samples from diverse tissue types.
Fine-tuning for age prediction led to robust age-dependent clustering. Notably, intrinsic age-related organization was evident even before fine-tuning. Moreover, MethylGPT outperformed existing age prediction methods (e.g., Horvath’s clock and ElasticNet), achieving superior accuracy. Its median absolute error for age prediction was 4.45 years, further demonstrating its robustness. MethylGPT was also remarkably resilient to missing data. It exhibited stable performance with up to 70% missing data, outperforming multi-layer perceptron and ElasticNet approaches.
Analysis of methylation profiles during induced pluripotent stem cell (iPSC) reprogramming showed a clear rejuvenation trajectory; samples progressively transitioned to a younger methylation state over the course of reprogramming. The model was also able to identify the point during reprogramming (day 20) when cells began showing clear signs of epigenetic age reversal. Finally, the model’s ability to predict disease risk was assessed. The pre-trained model was fine-tuned to predict the risk of 60 diseases and mortality. The model achieved an area under the curve of 0.74 and 0.72 on validation and test sets, respectively.
In addition, they used this disease risk prediction framework to evaluate the impact of eight interventions on predicted disease incidence. Interventions included smoking cessation, high-intensity training, and the Mediterranean diet, among others, each of which showed varying degrees of effectiveness across disease categories. This showed distinct intervention-specific effects across disease categories, highlighting the potential of MethylGPT in predicting intervention-specific outcomes and optimizing tailored intervention strategies.
Conclusions
The findings illustrate that transformer architectures could effectively model DNA methylation patterns while preserving biological relevance. The organization of CpG sites based on regulatory features and genomic context suggests that the model captured fundamental aspects without explicit supervision. MethylGPT also demonstrated superior performance in age prediction across different tissues. Moreover, its robust performance in handling missing data (≤ 70%) underscores its potential utility in clinical and research applications.
News
New Vitamin B12-Based Therapy Could Change How Brain Cancer Is Treated
Researchers have identified a vitamin B12–based compound that appears capable of crossing the blood–brain barrier and selectively accumulating in glioblastoma tissue. For decades, one of the biggest problems in brain cancer treatment has had [...]
Simple Fiber Supplement Cuts Knee Arthritis Pain in Just 6 Weeks, Study Finds
A daily inulin supplement may help reduce knee osteoarthritis pain while revealing a possible link between gut health, muscle function, and pain sensitivity. For millions of people living with knee osteoarthritis, managing chronic pain [...]
This Common Vitamin May Help Stop Prediabetes From Turning Into Diabetes
Vitamin D may help prevent type 2 diabetes in people with specific genetic variations, offering a possible path toward personalized diabetes prevention. More than 40% of U.S. adults have prediabetes, a condition in which [...]
Ebola, hantavirus: Is the world prepared for the next pandemic?
Funding cuts to health research and a growing antivaccine movement are making it harder than ever to respond to viruses. The World Health Organization (WHO) has declared that an Ebola outbreak in Uganda and [...]
May 2026 Healthcare News and Trends: Market Signals That Matter
Artificial intelligence is dominating headlines, telehealth has settled into a new normal, and digital health continues to promise transformation. However, much of what is being discussed in healthcare today reflects potential rather than reality. [...]
Scientists Rewire Donor Stem Cells To Outsmart Aggressive Blood Cancers
Researchers have tested a gene-edited stem cell transplant designed to shield healthy blood-forming cells from powerful cancer-targeting immunotherapies. For patients with highly aggressive blood cancers, stem cell transplantation can offer a rare chance at [...]
Recent Digital Health Trends, Insights and News – May 2026
Last month marked continued progress as digital health moves into its next phase — from AI expanding into drug discovery and core infrastructure to new federal pathways accelerating device access and home-based care. Together, [...]
Cancer Mystery Solved: Scientists Discover How Melanoma Becomes “Immortal”
Scientists have uncovered a previously overlooked mechanism that may help melanoma cells become effectively “immortal.” Cancer cells face a major problem before they can become deadly: They have to figure out how to stop [...]
How Visual Neurons Organize Thousands of Synaptic Inputs
Summary: A new study uncovered the organizational rules that determine how neurons in the primary visual cortex process information. By imaging both the cell bodies (soma) and the individual synapses (on dendritic spines) of [...]
Scientists Just Found a Surprising Way To Destroy “Forever Chemicals”
Scientists have uncovered a new mechanism that may help break down highly persistent PFAS pollutants. PFAS have earned the nickname “forever chemicals” for a reason. These industrial compounds are so chemically durable that they [...]
Scientists Discover Cheap Material That Kills Deadly Superbugs
A new sulfur-rich antimicrobial polymer shows strong effectiveness against fungal and bacterial pathogens and may offer an affordable solution to antimicrobial resistance. Antimicrobial resistance is creating growing challenges for both healthcare and food production, [...]
What to Know About Cicada, or BA.3.2, the Latest SARS-CoV-2 Variant Under Monitoring
Like periodical cicadas, the insects for which it is nicknamed, SARS-CoV-2 Omicron subvariant BA.3.2 is only just beginning to emerge after lying low for an extended period since it first appeared. Although it was [...]
Scientists Say This Simple Supplement May Actually Reverse Heart Disease
Scientists in Japan say a common supplement may actually help “unclog” certain diseased heart arteries from the inside out. A simple food supplement sold in Japan may have helped reverse a dangerous form of [...]
New breakthrough against radiation: Korean Scientists create revolutionary shield with nanotechnology
Korean Scientists develop new nanotechnology material capable of reducing radiation impacts in space missions, hospitals, and power plants. The search for more efficient protection technologies in extreme environments has just gained an important advance. Korean [...]
Scientists Just Discovered the Hidden Trick That Keeps Your Cells Alive
A strange bead-like motion inside cells may be the secret to keeping their DNA—and health—in balance. Mitochondria are often described as the power plants of the cell because they produce the energy cells need [...]
Scientists Discover Stem Cells That Could Regrow Teeth and Bone
Scientists just uncovered the cellular “blueprint” that could one day let us regrow real teeth. Researchers at Science Tokyo have uncovered two distinct stem cell lineages that play a central role in forming tooth [...]















