Almost everything that looks like a trading edge is luck wearing a convincing costume. This page is about how to tell the difference — explained from scratch, with no assumed knowledge.
LONA has tested roughly 145 trading rules. One is still standing. That ratio is not pessimism, it is what honest testing does to ideas. The rest of this page is the method that killed the other 144, one rule at a time, with the real numbers it produced.
Give a hundred people a coin each and ask them to flip it five times. Three or four of them will get five heads in a row. If you then interview those three about their technique, they will have one — people always do. Nothing about their method was real, but the evidence in front of you looks exactly like skill.
Trading research is that, with money attached. You try an idea on historical prices, and some ideas will have made money by chance alone. The question is never "did this make money in the test?" — something always does. The question is "would a coin have done as well?"
Here is the first real example from our own work. We hunted for a second strategy across 26 different ideas and picked the best one. It scored 0.94 on the measure we use.
If you generate 26 completely meaningless ideas and keep the best one, it typically scores 1.11. Our best real idea scored 0.94 — worse than what pure noise usually manages.
Not "not quite significant". Below the coin. We threw all 26 away.
This is the single most important idea on this page: the more things you try, the better your best result has to be before it means anything. A score of 0.94 found by testing one idea and a 0.94 found by testing twenty-six are completely different pieces of evidence, even though the number is identical.
More data sounds like more certainty. It usually isn't, because most rows of data are repeating what the row before them already said.
On 12 August our recorder captured 158,517 snapshots of Bitcoin's order book in a single day. That sounds like an enormous sample. It isn't — each snapshot is almost identical to the one a fraction of a second earlier, so they are not independent pieces of evidence.
Once you account for how much each row repeats the last, those 158,517 rows carry about 4,723 genuinely independent observations — three per cent. A confidence figure calculated on the raw row count would be five times too confident.
Our first five findings died of exactly this mistake in a different costume: six weeks of data read as though it were a large sample. It was one market mood, sampled very finely.
Imagine judging how good a school is by interviewing its graduates. You would never hear from anyone who dropped out. Every answer you get is from someone the process already selected.
Trading data has exactly this hole. Many crypto markets are launched, trade for a while, collapse and get removed from the exchange. If your historical data only contains the markets that still exist today, you have quietly deleted the failures — and any strategy tested on it will look better than it was.
Our main dataset is 742 markets, and 148 of them are dead. They are left in, with all 97,044 days of price they produced before they died. The strategy has to survive owning things that later went to zero, because in real life it would have.
There is a sting in the tail, and it is a good example of why measuring beats assuming. We measured this bias twice. On one three-year sample, removing the dead names made the strategy look 6.2 points a year better. On a different 6.6-year sample it made it look 2.9 points worse. Opposite directions.
The spread of that measurement turned out to be larger than the measurement itself, which means there is no such thing as "the" survivorship correction — and two of our earlier findings had been subtracting one as though there were. Both were corrected.
Before trusting any result, we deliberately break the test to check the test still notices. The main one is blunt: let the strategy cheat.
We feed it tomorrow's prices today. If a strategy that can see the future doesn't return something absurd, then the machinery is broken and every honest number it produced is meaningless too.
| Control | Result | What it proves |
|---|---|---|
| Let it see tomorrow's price | +164,956% | The engine works. If this ever looked reasonable, everything else is void. |
| Hold all 20 markets, always | loses money | The edge is in choosing, not in owning crypto. |
| Shuffle the choices at random, 200× | 10% beat the real one | On that dataset, luck closes the gap one time in ten — so we do not use it as evidence. |
That last row is the point of doing controls at all. It was our own result, and the control said it wasn't good enough. So we reported it as not evidence.
Here is how self-deception actually happens. You test an idea. It doesn't quite work. You try a slightly different version. Also not quite. On the ninth attempt something works, and by then you have honestly forgotten there were eight others.
The fix is to write down, before running anything: what is being tested, how it will be measured, and exactly what result counts as a pass. Then you cannot move the goalposts, because they are already in writing.
Here is the most recent one, and it is the best example on this page — because the rule stopped a result I wanted to keep.
Our strategy buys a market when it hits a 50-day high, then holds it for five days. Why five? An earlier study noticed, in passing, that holding for two days looked better — but noticing something in a grid of results is not evidence, it is fishing. So we wrote the test down first and ran it on a completely different dataset.
The blue and orange lines agree almost perfectly. The ranking of the seven hold lengths matches at 0.95 out of 1, both peak at two days, both slope steadily down. When we first ran this we wrote that down as “the same mechanism showing up twice”.
It wasn't. Look at the dates. Hyperliquid covers 2023–26. Binance perps cover 2020–26. Every single day of the Hyperliquid sample sits inside the Binance one. Two exchanges, one period, mostly the same coins — two windows onto one history. Of course they agreed. We had checked the venues were different and never checked the years were.
The green line is Binance spot, 2017–2020 — a period neither of the others contains. It is the only genuinely independent test on the chart, and it disagrees. Its best hold is five days, not two, and the tidy downward slope is gone:
| compared | overlap in time | agreement |
|---|---|---|
| Binance perps vs Hyperliquid | complete | +0.95 |
| Spot vs Binance perps | none | +0.25 |
| Spot vs Hyperliquid | none | +0.45 |
Two datasets that overlap in time are one dataset. Before treating agreement as proof, check the calendars, not the logos.
Counting datasets is the wrong unit. Count independent years. This project has five sources and roughly two eras — which is far less evidence than five sources sounds like.
Nothing about this changes the verdict below; the two-day hold was already not adopted. What it removes is the sentence we were proudest of. That check had been available for a day and nobody thought to run it until somebody asked whether a third dataset would help.
And the effect on risk is dramatic. Here is the worst fall from a peak in each year, for both versions:
The written test required two things to pass. One passed decisively. The other — the main one — failed by a whisker: it needed a number above zero and got −0.037.
So the answer is no. Not "no, but". Just no.
Adopting it anyway would mean the written rule was decoration — and every earlier idea we rejected under the same rule would retroactively mean nothing, because the rule would only bind when it agreed with us.
A rule you wrote down is only worth something on the day it stops you. That day was 15 August 2026.
When a result is borderline, the tempting move is to analyse the same data harder — more variations, more slices, more statistics. This almost never helps, because the uncertainty comes from how much independent data you have, and re-reading the same years does not create more of it.
We once ran 70 versions of our strategy on the same three years. All 70 were positive, which sounds like overwhelming evidence. It is one observation read 70 ways.
Only new time or new information. Testing the same rule on another exchange over the same years would produce a nearly identical answer — it is the same market history wearing a different logo.
So we went looking for genuinely older data, twice. What we found is the next section, and it is not the happy answer.
That is why the most valuable thing this project does now is wait. Every day, it writes down what it would have traded, before knowing what happens. Nobody can search that record, because it has not happened yet. It needs months, not cleverness.
Rule five says prefer new time. The obvious next move, then, is to go and buy more of it — find an exchange whose data starts earlier than ours. We went looking, and the answer was not the one we wanted.
We found the best-preserved archive in crypto. Bitfinex has daily prices going back to March 2013, seven years earlier than most of our data, and unlike almost everyone else it still serves the markets that died — we tested fifty delisted pairs and all fifty came back complete.
Then we counted what was actually in those seven years.
On 3 January 2016, exactly two coins on earth traded a million dollars a day. Five by that summer. Thirteen by the start of 2017. Our rule ranks a list of markets and buys the best twenty of them — and twenty liquid markets did not exist anywhere until around the middle of 2017.
Our earliest archive starts in August 2017. That is within weeks of the earliest date this test could exist at all. There is no hidden shelf of older data we have been too lazy to fetch. The ceiling belongs to the market, not to us.
Which is a more useful sentence than “we could not find more data”, because it tells you to stop looking and start waiting.
Poloniex was the biggest home for small coins in 2016 and 2017 and looked like the obvious answer. We tested thirty of its markets that have since been delisted. Twenty-nine of them have been deleted outright — no listing, no prices, nothing.
A backtest on Poloniex would have produced a beautiful 2017 and would have been measuring one thing only: which coins are still alive in 2026. That is rule two's mistake, handed to you gift-wrapped. It is now written down as a rule of its own: an exchange that does not keep its dead is not a weaker source, it is a disqualified one.
There was still a job worth doing. One stretch of our evidence — August 2017 to 2019, which contains the whole 2018 crash and is the most informative period we own — had only ever been seen through one exchange. If that result were a quirk of one venue's fees or listings, nothing we had would have noticed.
So we ran the identical rule on Bitfinex over the identical years. Getting the list of markets took 17,576 separate requests, one for every possible three-letter ticker, because Bitfinex's own published list contains none of its dead ones. It found 268 markets, 204 of them long gone.
| Same years, two exchanges | Return/yr | Score | Buy all 20 instead |
|---|---|---|---|
| Binance · 2017–2020 | +35.7% | 1.28 | +2.2%, −91% fall |
| Bitfinex · 2017–2020 | +49.2% | 1.29 | −0.5%, −92% fall |
Two exchanges, two sets of customers, two fee schedules. The scores land at 1.28 and 1.29.
And here is the check that makes that worth anything. The obvious objection is that it is the same years, so it must be the same coins. We measured it instead of arguing about it: the two exchanges' traded lists overlap by a quarter. Fifteen markets were tradable on Bitfinex and never on Binance. Three quarters of the two universes are different.
The prize was 2018. Every previous look at that crash had too few markets to choose between — below our own minimum, so we had to write “this data cannot settle it” and mean it. Bitfinex's 2018 clears the minimum, and it says: the rule lost 4.4% in the year bitcoin fell 71%. BitMEX, which shares no markets with either exchange, measured the same crash at −5.3% while bitcoin fell 73.4%. Two independent looks, one point apart.
That Bitfinex result scores 1.79 on confidence — below 2, so within reach of luck on its own. Its worth is the agreement, not the number. Two of its four years are still too thin to count. And Bitfinex from 2021 onwards is weak — on a median of six markets, too few to be evidence either way. That last line is published because we wrote down in advance that we would publish it, whichever way it came out.
It also does not add a new era. It is a second view of years we already had. By rule five's own logic, our count of genuinely independent periods is still two.
One rule. Buy a market when it closes above every price of the previous 50 days, hold it five days, size it at a twentieth of the account, otherwise sit in cash. Roughly 16% invested on an average day.
| Tested on | Span | Return/yr | Worst fall | Confidence |
|---|---|---|---|---|
| Binance · 742 markets | 6.6 yr | +32.4% | −28% | 2.4 |
| BitMEX · bitcoin only | 10.25 yr | +27.7% | −54% | 2.39 |
| OKX · 356 markets | 6.6 yr | +26.1% | −35% | 2.1 |
| Hyperliquid · 232 markets | 3.0 yr | +20.9% | −14% | 1.5 |
“Confidence” is a t-statistic. Below 2 means the result is within reach of luck; above 2 starts to be worth something. Note the bottom row: the three-year sample is the least convincing, and it was the first one we had.
It does not mean +32% a year. Trend-following on crypto is one of the most heavily studied ideas in finance, so some of that number is the field's selection rather than ours. A realistic live expectation is meaningfully lower.
It does not mean a comfortable ride. Only 47% of trades are up after their first day. The strategy loses slightly more often than it wins — its edge is that the wins are bigger. Anyone watching will see red on most days while it behaves exactly as measured, and there is roughly a one-in-eight chance of a losing year.
None of this was designed in advance. Every line exists because a result died on it — and several of them exist because a result we had already published turned out to be wrong.