Guide · 6 min read

Survivorship bias in a backtest, and the different problem it gets confused with

A backtest built from today's list of large caps cannot see the companies that were removed from that list along the way. That gap has a name, a rough size, and a completely different cousin that shows up in the coverage numbers on every study page.

Read one of these studies to the bottom and the disclaimer says the same thing every time: real world execution costs, slippage, and survivorship bias can materially change outcomes. It is one line, easy to skip past, and it names one of the easiest ways a backtest can look better than the market really was, without anyone doing anything wrong.

The bias itself

Build a universe from today's list of large cap US stocks and run it back to 2015, and every company that was on that kind of list in 2015 but is not on it now is invisible to the test. It was removed for a reason: bought out, delisted, or it went bankrupt. Bankruptcies and forced delistings are disproportionately the bad outcomes, and a current-list universe drops them by construction, not by anyone's choice. The stocks left over to test are, on average, the ones that survived, which is exactly what the name says.

This is not a hypothetical problem specific to one platform. It is a known property of building any historical study from a present-day list, and it shows up whenever a backtest, a fund's marketed track record, or an index's reported history is quietly built from constituents chosen after the fact rather than as of the date being tested. The honest position is that it pushes a reported return upward, by an amount that depends on the universe, the sector mix, and how much churn that decade had, and outside estimates of the size vary enough that no single number is safe to repeat as settled.

A different problem wearing similar clothes

Tenachine's MACD signal cross study ran its rule across the same 220 US large caps as every study on this site, 2015 to 2024, and only 143 of them completed a backtest. Seventy seven did not. That is a real gap, and it is worth investigating before trusting the 1.18 percent average return or the 36.98 percent win rate the completed runs produced. But it is not survivorship bias. It is a failed run: a ticker that was already inside the chosen universe and did not produce a usable result, for reasons that can be as mundane as a data gap or too few signal crosses to compute a stable statistic.

That gap is not rare and it is not small, and it varies a lot from one study to the next. Golden cross lost 27 of 220 tickers to failed runs, 12.3 percent of the universe. Bollinger squeeze breakout lost 59, 26.8 percent. Macd signal cross lost 77, the 35.0 percent above. Fifty two week high momentum lost 126 of 220, 57.3 percent, more than half its universe. None of those four numbers says anything about survivorship. They say a lot about how much of any given study's headline return you can trust before checking coverage first.

Share of the 220-stock universe lost to failed runs, by study
Share of the 220-stock universe lost to failed runs, by studyFailed runs as a share of the 220-stock universe: golden cross 12.3 percent, bollinger squeeze breakout 26.8 percent, macd signal cross 35.0 percent, 52 week high momentum 57.3 percent.Golden cross12.3%Bollinger squeeze breakout26.8%Macd signal cross35.0%52 week high momentum57.3%
Coverage ranges from about one ticker in eight lost to well over half the universe. A study passing this check tells you nothing about survivorship. A study failing it tells you the headline number rests on a smaller, unexplained sample than the universe size alone suggests.

The distinction matters because fixing one does not touch the other. A study could complete all 220 runs with no failures at all and still be built on a universe of large caps that already excludes every company that got removed from that kind of list between 2015 and today for going under. Coverage, the fraction of a chosen universe that produced a result, and survivorship, whether the companies that failed to survive were ever in the universe to begin with, are answers to two different questions. Tenachine's guide on reading a return distribution makes the coverage point on its own terms, using the golden cross study's 27 failed runs out of 220. Neither guide's coverage number says anything about survivorship, and this one should not be read as if it does.

Why it is hard to just fix

A survivorship-free universe requires point-in-time constituent data: the actual list of large cap US stocks as it stood on January 1, 2015, including every name that would later disappear, not the list as it stands today. That kind of dataset exists, but it is an institutional product, expensive, and most retail-facing research tools, this one included, are built from a current roster rather than a point-in-time one. That is not a defense of the gap. It is a fact about the cost of closing it, and a reason to treat the gap as a known, standing feature of almost any accessible backtesting tool rather than a flaw unique to one.

What to do with this

For any backtest you did not build yourself, ask whether the universe came from today's list or from a point-in-time list. Most of the time the honest answer is today's list. Then treat the reported return as biased upward by an unknown amount, rather than deciding the bias is either negligible or disqualifying, because both of those are guesses dressed as conclusions.

Expect the effect to be largest over long windows and in sectors with heavy churn, energy and financials around 2008 being the standard example, and smallest over short windows in stable sectors. Keep coverage and survivorship as two separate checks, since a study passing one tells you nothing about the other. And do not wait for a single number to fix it. The honest response is a wider margin of safety on the headline return, not a correction factor.

Why the disclaimer names it

Every study and investor page on this site carries the same disclaimer sentence naming survivorship bias directly, rather than leaving it to a footnote a reader has to go looking for. This site does not currently publish which companies were excluded from a universe or when, and that is a fair question to keep asking of any backtesting platform, this one included, before treating a study's headline number as more precise than the universe behind it allows.

The lesson is not to distrust every backtest equally. It is to read the universe construction before the return, the same way the failed-run count deserves a look before the average. Past performance is not indicative of future results, and a decade of survivors is not the same thing as a decade of the market.