Read the spread, not the average, when you judge a backtest
The single number at the top of a backtest is the least useful thing on the page. The shape underneath it tells you whether the edge is real or whether two lucky stocks carried the whole result.
The golden cross study on this site reports an average return of 9.02 percent per ticker. It ran the same 50 over 200 day moving average crossover across 220 US large caps between January 2015 and December 2024, and 193 of them completed a backtest. Nine percent over ten years is not a good result. It is also not the interesting part of the page.
The interesting part is that roughly a third of those tickers lost money, and the single biggest bucket, at 39 percent of the universe, made more than 10 percent. Both facts live inside the same average. Neither is visible from it.
What the average is actually doing
A mean is a summary that assumes the thing it summarises is roughly symmetric. Trading returns are not. A losing trade is bounded by the stop and a winning trade is not bounded at all, so the right tail runs much further than the left one. Average a set of numbers with that shape and the answer drifts toward the few large winners.
That is why the mean and the median disagree here. The mean is 9.02 percent. The median ticker made 5.51 percent. The gap between those two numbers is the part of the result contributed by the tail.
None of this means the strategy is bad. It means the headline does not tell you whether it is.
Three questions the distribution answers
How much of the universe lost money
Add the loss buckets. Here that is 32.2 percent of tickers, which is a real number to sit with before you trade the rule. A third of the time, applying this strategy to a large cap you picked in advance would have cost you money over a decade. If you were going to run it on five names rather than 193, that is the odds you were taking.
How concentrated the winners are
Drop the best five symbols and recompute. If the result survives, the edge is broad. If it collapses, the backtest found five stocks that went up a lot between 2015 and 2024, which is a fact about that decade rather than a fact about the rule. Most trend following backtests on US equities have this problem to some degree, because the index itself trended for most of the sample.
Whether the losses are shallow or deep
Here they are shallow. The worst bucket, losses beyond 10 percent, holds 2.6 percent of the universe, and the average maximum drawdown across successful runs was 9.35 percent. A strategy that loses often but loses small is a different instrument from one that loses rarely and loses badly, even when the two report the same average.
The failed runs matter too
Twenty seven of the 220 tickers did not complete a backtest. That is 12 percent of the universe missing from every number above. Before you read anything else, find out why runs failed and whether the failures are random. Missing price history is harmless. A rule that silently drops every symbol with a gap or a halt is not, because those are the symbols where the strategy would have been tested hardest.
Any backtest that reports only its successes is reporting a filtered sample, and the filter is usually correlated with the thing you are measuring. This is the same mistake as reading fund performance tables that exclude the funds that closed.
A reading order that works
- Check the universe size and how many runs completed. Everything downstream is conditional on that number.
- Read the median before the mean, and note the gap between them.
- Add up the share of the universe that lost money.
- Look at the worst bucket and the average drawdown together. Frequency and depth are separate risks.
- Ask what would remain if the top five symbols were removed.
- Only then look at the headline return.
This order is deliberately backwards from how results are usually presented, because the presentation order optimises for a decision and the reading order should optimise for doubt.
Why Tenachine publishes the buckets
Tenachine takes a strategy written in plain English, generates the code, and runs it across a universe of US stocks rather than one favourite ticker. Every study publishes the per ticker table, the sector breakdown and the return distribution alongside the headline, because a single symbol backtest can be made to say almost anything and a 200 symbol distribution cannot.
The golden cross numbers above are one such study, and they are not an endorsement of the golden cross. Nine percent across ten years, with a third of the universe losing, is a result worth knowing before you build anything on top of it. That is the whole argument for reading the spread.
Past performance does not predict future results, and a backtest is a description of one decade, not a forecast of the next.