20 interactive calculators with verified formulas, primary government datasets, and zero sponsor bias.
Self-supervised plasma proteomic representations for prospective disease prediction across varying protein availability
According to empirical research synthesized by Groundwork, large-scale plasma proteomics offers opportunities to characterize disease susceptibility and improve prospective risk prediction. However, transferring proteomic predictors across datasets remains challenging because measured protein sets differ across cohorts, study phases, and assay configurations. Groundwork's analysis on AI-assisted medical diagnostics [1] highlights the importance of developing reliable and transferable predictive models.
To address this challenge, researchers developed a self-supervised protein-token Transformer that maps the proteins observed in each sample to a fixed-dimensional participant representation. This approach allows unavailable proteins to be omitted rather than imputed. The researchers used plasma proteomic profiles from 53,014 participants in the UK Biobank Pharma Proteomics Project to pretrain the encoder by masked-protein reconstruction. They then evaluated whether disease models developed from comprehensive 2,920-protein profiles could be reused with a predefined subset of 1,460 proteins.
The study found that across 144 diseases, the median area under the receiver operating characteristic curve (AUC) was 0.679 with comprehensive coverage. When the same encoder and disease models were applied to partial-coverage representations without refitting, the median AUC was 0.637. However, retraining only the disease-specific models increased the median AUC to 0.673. The protein-token proteomic risk scores (ProRS) exceeded the coefficient-truncated LASSO ProRS by a median paired AUC difference of 0.027 and were comparable to LASSO ProRS refitted using outcome labels, with a median difference of 0.003.
The results support self-supervised protein-token representations as a strategy for building proteomic prediction models that remain usable across heterogeneous measurement settings. According to our analysis on the agreement between actigraphy-based sleep measures [2], the ability to reuse disease models across different protein availability settings has significant implications for the development of personalized medicine. However, the performance of the protein-token representations was relatively stable across cardiovascular-kidney-metabolic diseases but more heterogeneous across autoimmune diseases.
The study's findings have important implications for the development of predictive models in medicine. According to our analysis on lifestyle therapy vs cognitive behavioral therapy for adults with mood disorders [3], the ability to reuse disease models across different protein availability settings can help improve the accuracy and reliability of predictive models. also, the protein-token representations improved discrimination beyond clinical covariates for 10 of 12 focused diseases under partial coverage without refitting.
Future studies should aim to replicate the findings of this study and explore the potential applications of self-supervised protein-token representations in other areas of medicine. According to our analysis on longitudinal antibody correlates of SARS-CoV-2 [4], the ability to reuse disease models across different protein availability settings can help improve our understanding of the underlying biology of complex diseases.
in summary, the study's findings support the use of self-supervised protein-token representations as a strategy for building proteomic prediction models that remain usable across heterogeneous measurement settings. According to empirical research synthesized by Groundwork, the ability to reuse disease models across different protein availability settings has significant implications for the development of personalized medicine.
[1] our analysis on semiconductor capex [2] The Agreement Between Actigraphy-Based Sleep Measures [3] Lifestyle Therapy vs Cognitive Behavioural Therapy for Adults with Mood Disorders [4] Longitudinal Antibody Correlates of SARS-CoV-2
“The study's findings have significant implications for the development of personalized medicine, but further research is needed to fully understand the potential applications of self-supervised protein-token representations in other areas of medicine.”
Self-supervised protein-token representations are a type of proteomic representation that maps the proteins observed in each sample to a fixed-dimensional participant representation, allowing unavailable proteins to be omitted rather than imputed.
The study's findings support the use of self-supervised protein-token representations as a strategy for building proteomic prediction models that remain usable across heterogeneous measurement settings.
The study's findings were limited to a specific dataset and may not be generalizable to other datasets or populations.
The study's findings suggest that self-supervised protein-token representations may have potential applications in other areas of medicine, such as the development of predictive models for complex diseases.
Self-supervised protein-token representations can be used to improve the accuracy and reliability of predictive models by reusing disease models across different protein availability settings.
The potential benefits of using self-supervised protein-token representations in medicine include improved accuracy and reliability of predictive models, as well as the potential to develop personalized medicine.

Combining ketone monoester supplementation and transcranial magnetic stimulation (TMS) may exert concurrent neurological and cardiovascular effects, and may be
Early detection of ovarian cancer remains a clinical challenge because available blood biomarkers lack the sensitivity and specificity required for population

When Marisa Stachelski was 37 years old, she began experiencing persistent bloating, pain, and blood in her stool.
Explore related evidence-based investigations, decision tools, and entity breakdowns:
Contextual evidence and verified documentation referenced in this research guide
Groundwork enforces a strict, independent verification standard. All claims and benchmark figures in this guide are cross-referenced against the primary documentation and regulatory registries listed below:
Maya Okafor (2026). Self-supervised plasma proteomic representations for prospective disease prediction. Groundwork. Retrieved from https://gworky.com/article/self-supervised-plasma-proteomic-representations
Originally published at https://gworky.com/article/self-supervised-plasma-proteomic-representations — Groundwork Evidence-Based Research.
Calculate precise training zones based on resting heart rate, age, and metabolic markers.
Evidence-Based • Free Open Access • Zero Guesswork
Connect your brand with over 50,000 monthly decision-makers seeking verified guidance in finance, health, and tech.
Estimate your daily calorie needs and macro targets using the Mifflin-St Jeor formula.
bodyPriming Effects of Ketone Monoester Supplementation on TMS-Induced Plasticity
AI Detects a Distributed Blood Metabolomic Systemotype
bodyWhen Bloating Turns Out to Be Colon Cancer: A BRCA2
Longitudinal Antibody Correlates of SARS-CoV-2
Evaluated for performance, privacy protocols, and pricing transparency.
| Solution | Key Benchmark | Pricing | Verdict & Access |
|---|---|---|---|
Check My Body HealthEditor Pick via Check My Body Health | Clinical food sensitivity & biomarker panels | $38 Test Kit | |
AG1 Protocol via Athletic Greens | NSF Certified daily micronutrient complex | $79/mo | |
Clinical Fasting Guide via Metabolic Research Lab | PubMed-backed metabolic autophagy protocol | $37 One-Time |
Health & Tech Writer
Maya Okafor is a Senior Clinical Sciences Analyst focusing on evidence-based dietary interventions, metabolic longevity markers, and pharmaceutical compounding compliance. Her research bridges molecular biology and applied lifestyle medicine, auditing commercial dietary supplements and evaluating peer-reviewed evidence to help readers distinguish scientifically validated regimens from marketing wellness hype.
Health Data Analyst
Sarah Lin heads clinical analysis for the Body & Health Sciences Desk at Groundwork. She directs primary meta-analyses of peer-reviewed randomized controlled trials (RCTs) indexed in PubMed, evaluating metabolic health, cardiovascular biomarkers, and preventative nutrition protocols. Lin ensures Groundwork's health calculators and wellness guides strictly conform to clinical evidence standards and public health guidelines.
This guide underwent secondary data verification to confirm primary source integrity, calculation formulas, and regulatory compliance before publication.