Aging biomarkers are measurable quantities intended to indicate how far or how fast a person is aging, independent of the calendar. The field wants them for a specific and practical reason: without one, a trial testing whether a drug slows aging must wait for participants to become ill or die, which takes decades and enormous numbers of people. A validated biomarker would compress that to a few years. None has been accepted for the purpose.
Why the field needs one
The Geroscience hypothesis holds that slowing aging would delay many chronic diseases at once. Testing it requires an endpoint. Lifespan is unambiguous and unusable in humans. Incidence of a single disease reintroduces the disease-by-disease framing the hypothesis was built to escape. A composite of several age-related diseases plus death is the compromise adopted by the TAME design, and it still requires thousands of participants over years.
A surrogate biomarker would change the economics. A six-month readout would let sponsors run dose-ranging studies, compare candidate compounds, and fail fast, which is how drug development works in every field that has one. The absence of a surrogate is a substantial part of why so few geroprotector candidates have entered serious human testing despite a large animal literature.
What counts as a biomarker of aging
The criteria most often cited descend from work sponsored by the American Federation for Aging Research in the late 1980s and restated since.1 A biomarker of aging should predict remaining lifespan or the onset of age-related dysfunction better than chronological age does; it should reflect an underlying process, ideally one of the Hallmarks of aging, rather than a specific disease; it should be measurable repeatedly without significant harm; and it should work in laboratory animals as well as in humans, so that it can be validated against lifespan in a species where lifespan is observable.
Modern framings add a decomposition into three tiers. Analytical validity asks whether the measurement is reproducible across laboratories, platforms and repeat draws. Clinical validity asks whether it predicts outcomes in independent cohorts. Responsiveness asks whether it moves when an intervention is applied, and, crucially, whether a change in the marker predicts a corresponding change in the outcome. Most candidates have some evidence for the first two tiers and almost none for the third.
Candidate classes
Candidates are conventionally sorted by what they measure, a catalogue set out under Biological age. Sorted instead by which validity tier each class is stuck at, the picture is more useful, because the tier a class fails is what determines the experiment that would rescue it.
Molecular measures are cheap and hard to interpret. DNA methylation dominates, in the form of epigenetic clocks, with proteomic, metabolomic, glycomic and transcriptomic panels alongside. One blood draw yields tens of thousands of features, and a model can be fitted from them to almost any outcome. That is simultaneously the attraction and the difficulty: an index selected for prediction carries no evidence that moving it moves anything, and that evidence cannot come from the data the index was fitted on.
Cellular measures are the ones the field most wants and least has. Senescent-cell burden is the direct target of Senolytics, and no accepted method exists for quantifying it in a living person, because no clean marker of Cellular senescence exists to build the method on. Assays therefore rely on surrogates for a surrogate. Clonal haematopoiesis, a readout of Stem cell exhaustion, is the exception: measurable by sequencing, predictive of cardiovascular as well as blood outcomes, and confined to one tissue.
Immune and inflammatory measures track a real axis and move for the wrong reasons. Composite immune-aging scores, thymic output measures and circulating markers linked to Inflammaging all decline with age. They also shift within days after infection, injury or vaccination, which is close to disqualifying for something meant to read a decades-long process off a single timepoint.
Functional measures are the only class with real responsiveness data. Gait speed, grip strength, chair-rise time and cardiorespiratory fitness predict mortality and disability well, and they already serve as endpoints in frailty and sarcopenia trials, so the regulatory path for them is partly mapped. The awkwardness is commercial rather than scientific: they improve reliably with training, so a sponsor whose drug is judged on grip strength is competing against exercise.
Coordination efforts
The Biomarkers of Aging Consortium, established in the early 2020s, brought together academic groups, companies and funders to standardise definitions, share data and set validation criteria. Its output includes a framework paper setting out what would be required for a biomarker to be used in longevity-intervention trials and subsequent work on validation standards.2 The consortium's contribution is less a new measure than an agreed vocabulary and a shared benchmark, which the field previously lacked.
Surrogates can misleadCardiology provides the cautionary case. Antiarrhythmic drugs suppressed ventricular ectopic beats, a plausible surrogate for sudden cardiac death, and the Cardiac Arrhythmia Suppression Trial found they increased mortality.3 A marker can be strongly associated with an outcome and still be the wrong thing to move.
The regulatory bar
Regulators do not accept surrogates because they correlate with outcomes. They require evidence that an intervention's effect on the surrogate accounts for its effect on the clinical outcome, which is a mediation claim and much harder to establish. In the United States, formal qualification runs through a biomarker qualification programme, and no aging biomarker has been submitted successfully through it.
A second obstacle is categorical. Aging is not recognised as an indication by the major regulators, so even a perfect biomarker would be measuring progress against something a label cannot name. This is why proposals in the field are usually framed around a specific age-related condition or around multimorbidity, and why the ICD classification of aging-related decline has been argued over with unusual intensity for a coding question. A working group convened around Nir Barzilai's TAME trial published a shortlist of blood-based markers judged closest to usable for geroscience trials, chosen for availability and predictive validity rather than for mechanistic depth.4 The economic case for clearing this obstacle is set out under The longevity dividend.
Validation problems
Validating a biomarker of aging against mortality is partly circular: mortality is the outcome, and a marker that predicts it may be measuring current illness rather than aging. Distinguishing the two requires long follow-up in healthy cohorts, which few datasets support.
Animal validation offers a partial escape. A marker that tracks lifespan across mouse strains and across interventions of known effect, including Caloric restriction and Rapamycin, has passed a test no human dataset can supply. Cross-species methylation clocks, which Horvath and collaborators extended across mammalian species, were built partly for this purpose. The translation gap remains: a marker calibrated in mice may not carry the same meaning in a species that lives thirty times longer.
Reliability is the mundane problem that undermines many published intervention effects. Where a measure's test-retest variation is comparable to the effect being claimed, the claim cannot be evaluated, and this has been documented for several widely used clocks and partly remedied by reconstructing them on principal components.5 The same difficulty afflicted leukocyte telomere length, the field's previous favourite marker, which proved too noisy at the individual level to be informative.
Outlook
Two paths are open. One is to keep improving molecular measures until one clears the mediation bar, which will require intervention trials large enough to demonstrate that the marker carries the treatment effect. The other is to give up on surrogates for now and use function directly, which is the choice made by XPRIZE Healthspan in defining its target as restored muscle, cognitive and immune performance. The second path is slower per trial and immune to the criticism that has attached to biological-age claims. Which route the field takes will shape whether the first credible human geroprotector result arrives from a biomarker readout or from a decade-long clinical endpoint.
See also
- Biological age
- Epigenetic clocks
- Geroscience hypothesis
- Healthspan
- Hallmarks of aging
- XPRIZE Healthspan
- Compression of morbidity
References
Footnotes
-
paperBaker, G.T. & Sprott, R.L. "Biomarkers of aging." Experimental Gerontology, 1988. ↩
-
paperMoqri, M. et al. "Biomarkers of aging for the identification and evaluation of longevity interventions." Cell, 2023. ↩
-
paperEcht, D.S. et al. "Mortality and morbidity in patients receiving encainide, flecainide, or placebo: the Cardiac Arrhythmia Suppression Trial." New England Journal of Medicine, 1991. ↩
-
paperJustice, J.N. et al. "A framework for selection of blood-based biomarkers for geroscience-guided clinical trials: report from the TAME Biomarkers Workgroup." GeroScience, 2018. ↩
-
paperHiggins-Chen, A.T. et al. "A computational solution for bolstering reliability of epigenetic clocks." Nature Aging, 2022.↩Addresses test-retest noise by rebuilding existing clocks on principal components; it does not establish that any of them measures aging.