Sometimes I wonder if science is even real. I can’t see DNA, how do I know it exists? Maybe it’s all an elaborate hoax played on me by my textbooks and teachers. But then my controls behave, and for the thousandth time it turns out that science is real, that this tower of abstraction we call modern science really does hold up under its own weight.
But I distinctly remember the times when that didn’t happen.
Early in my career, I was testing a theory that transposons are part of what drives cellular senescence in old age.1 I wanted to test it, so I tried knocking down transposon activity with CRISPRi in primary fibroblast cell lines from both young & old donors, with appropriate negative transfection controls. Then I stained for senescence with a standard β-gal assay2 and did a bunch of microscopy. It was supposed to look like this:
The standard worked, but the experimental outcome was much harder to interpret. Between microscope settings, cell segmentation, image gating, and the exact analysis technique, I could have made almost any claim I wanted. My honest, best guess was that there was no real signal at all, so I took my own advice and gave up. I couldn’t look at any (mammalian) β-gal assays the same way again, and I started doubting the entire field of aging.
It’s more than just the one assay
Ok, so I had a bad experiment, big deal. Troubleshoot, run it again, move on with the science, right? But the more I looked into it, the more I found that the β-gal assay is actually a pretty bad assay.3 More modern literature says that there isn’t a good single way to identify senescent cells since the various techniques are very context-dependent and tend to contradict each other.
This issue isn’t isolated to the β-gal assay. Senescence, and aging itself, are high-level, emergent phenotypes caused by the breakdown of (one of) the most complex signaling systems in biology – or something equally complicated. The space of plausible-seeming explanations is so huge that any result can be explained after the fact. My favorite example is hormesis, where you get U-shaped dose-response curves because a little damage activates a stress response that over-repairs and ends up benefiting the organism. It’s a real phenomenon, but it’s also a great way to explain noise or bad data.
This doesn’t mean there’s scientific misconduct. Science is hard, and humans are great at fitting data to preconceived biases. If you have a very complex underlying mechanism, you can come up with an explanation for almost any piece of data. This is described well by “The Garden of Forking Paths,” where seemingly mundane decisions about how to analyze the data can give high statistical significance to completely random data with enough degrees of freedom. This is especially true if somebody has spent months or years collecting that data (aging research is infamously slow) and wants to progress their career (publish or perish!).
These issues aren’t unique to aging. An adjacent field has famously run into the same problem.
The Cancer Replication Crisis
In a now-famous result, scientists at Amgen reported in 2012 that they tried to reproduce 53 landmark cancer papers and only confirmed 6 of them. Bayer’s parallel work had a 25% success rate, and a broader replication effort found that less than half of tested cancer papers replicated, and the effects that did replicate were typically much smaller than what had been originally published. The implication here is that some of the authors of the original studies created signal from noise with their specific analyses.
Nobody has done that kind of replication effort for aging science, though it has happened anecdotally. Sirtris was bought for $720 million in 2008 based on resveratrol and sirtuin results, then shut down after their primary data were shown to be an assay artifact. The National Institute on Aging tried to replicate 15 well-studied aging interventions, and rapamycin is the only one that definitely works.4
My bet is that aging research has a replication issue at least as bad as the cancer one, and we just don’t understand how deep the issue goes because it hasn’t been a big clinical direction.5
Nope, I quit
The Importance of Stupidity in Scientific Research is one of my favorite pieces of all time, and its primary thesis is that being wrong about your science is fine, so long as you learn something. Similarly, in Hamming’s famous You and your Research, he instructs his audience to “believe the theory enough to go ahead [but] doubt it enough to notice the errors.”
I am not confident I would learn enough true things in aging research to dedicate a career to it. It is one of the most exciting open problems in biology, and while my first published paper is an aging review on autophagy, I don’t believe in the theory enough to specialize in it. I think I’d end up spending years or decades of my life on something irrelevant. Instead, I wanted to work in a field with faster experiments and a higher likelihood of eventually discovering objective truth.
But I’ve stayed interested in aging and kept an eye on the state of the field over the years.
Recent Progress
You might have noticed that most of the papers I’ve linked so far are pre-2015. That’s largely because I’m more familiar with that era of literature, but also because the field of aging has shifted since then.
The biggest change has been epigenetic clocks. Aging is correlated with genetic dysregulation through drifting epigenetic markers, and by checking methylation patterns you can get a sense of how far along that process is.6 Knowing how old you are might not sound like a big deal, but a lot of the bad science I’ve mentioned thus far resulted from poor measurements. Being able to quickly test which interventions make cells or animals younger is a game changer.
But notice I said correlated, not causative. Epigenetic tests are a step in the right direction, but they’re not without issues. Different cell types have very different epigenetics, so the ratio of different cell types is a big confounding variable for epigenetic sequencing. They also have issues with low reliability, though the richer data probably means that’s solvable.7 It’s an open problem, and time will tell if this is the be-all assay its proponents claim.
The hot new therapeutic approach is partial reprogramming, taking a page from the induced pluripotent stem cell toolbox to reset epigenetic markers without turning everything into stem cells. Doing that to a living body and keeping it alive is a tricky problem, and I think this tweet speaks more to the issues here than anything I can say:

Everything in the public record is the stuff that gave p < 0.05, which worked out great for cancer in the 2010s.
I’ve heard several hot takes about what’s next – better reprogramming, small molecules, peptides, gene editing, or some AI-fueled combination of them all. My take is that the next step is probably some kind of whole-cell model or AI-driven big-data simulation that predicts results ahead of time, because that’s the only way I can imagine bypassing the Garden of Forking Paths when your signaling network looks like this:8
But that’s the kind of science that’s very hard to be confident in for the near future.
The Kind of Science I Enjoy
That’s why I do the science I do now. At Pioneer Labs we engineer microbes for Mars, and after a few months of being unable to reproduce published data on salt tolerance, we developed a simple and reliable assay that seems to work for nearly all of the stressors we’re interested in. We do big high-throughput library experiments, and we make sure to use the right statistics9 to get real data rather than noise. I am regularly reassured that our science is real, and that this tower of knowledge we’ve spent three years building will hold its own weight.
Transposable elements are ancient parasitic sequences that make up about half the human genome, and they’re mostly inactive. But there’s some good evidence that they wake up as you get older and the silencing machinery breaks down, and when they go crazy it looks like cancer and the cell gets shut down via the senescence pathway. Senescent cells cause a lot of inflammation and are responsible for some of the health outcomes in aging, and so preventing aging-related senescence might increase healthspans.
I know this sounds weird, but it is a standard assay that’s supposed to be high-signal and easy to run. β-galactosidase activity is present in senescent fibroblasts but absent from pre-senescent and quiescent ones, and rises with donor age in human skin cells. There’s a Nature Protocols paper and commercial kits, which I used (The image is from the kit material).
Really, the assay is measuring lysosomal β-galactosidase activity, and that can change for a lot more than senescence. It goes up with cell confluence and serum starvation, and is high in perfectly healthy post-mitotic neurons.
Standard disclaimer: rapamycin is an immunosuppressant and has a lot of nasty side effects, including messing up your metabolism. Most reasonable physicians will tell you not to take it. A few other compounds showed mixed and sex-dependent results.
Altos Labs launched in 2022 with $3 billion to work on aging and might be doing a lot of this replication. As far as I know, they haven’t published anything like that. And companies have even less reason to publish negative results than academics do.
The original paper on this was Horvath 2013, but it’s advanced a lot since then. To the point where there are multiple companies whose primary competitive advantage is their secret scoring algorithms, which disagree with each other by a surprising amount.
The nice thing about doing whole-genome epigenetic sequencing is it gives you a lot of data, and so it’s probably possible to control for noise as long as the underlying signal is in there. Here’s one proposal for that, and a more recent preprint that was able to separate technical & biological reproducibility. This is an active area of research, and I’m hopeful.
Technically, this is the signaling network for colorectal cancer from this paper, but it shares a lot of space with a similar graph for aging, which is, if anything, more complex.
It is not trivial to get quantitative statistics from big library experiments, and I highly recommend reading the RNAseq literature (start with DESeq2) to understand the issues inherent in quantitative barcode sequencing. I might write a future blog post on the topic, because my mind got blown several times in the process of understanding how to run a good library experiment.



