August 17, 2026 · 7 min read
How I test trading ideas before trusting them
The trading ideas that cost me the most were the ones that looked like they worked. Here are the checks I now run on every idea, and two that nearly got through.
- Research methodology
- Quantitative finance
- Statistics
I've been building trading systems for about a decade now, and a lot of that time goes into researching trading strategies. A strategy here just means a set of rules for when to buy and when to sell, like "buy when the price has dropped three days in a row." Before I trust a strategy with money, I test it on historical data, which means running its rules over past prices to see whether it would have made money. That test is called a backtest.
A strategy that loses money in a backtest doesn't cost me much. I see that it lost, I drop the idea, and I move on to the next one. The ones that cost me are the strategies that look like they work but don't. Those make it through the checks I thought to run, they come with a story that sounds believable, and they stick around long enough that I start building more work on top of them before I find out they were never real.
I keep a running list of every idea I've tested, and each one gets a verdict. It's either tradeable, real but not profitable once trading costs are included, or rejected. Most of them end up rejected. What follows are the checks that did most of that rejecting, and two ideas that got past nearly all of them before I found the problem.
Why a result that looks good is the dangerous one
A bug that crashes the program or produces obviously wrong numbers is easy to catch, because you can see it's broken. The harder problem is code that does exactly what I told it to do, but where what I told it to do answers a slightly different question than the one I meant to ask. Everything it produces is consistent, so nothing looks wrong.
The tricky part is that extra checks don't necessarily help with this. If every check I run is built on the same flawed data, every check inherits the same flaw, so running five versions of a test only tells me the result holds up across those five versions. It doesn't tell me the result is real.
The checks every idea has to pass
Each of these is on the list because a specific mistake once got past me without it.
Test on data the idea has never seen. I split the history into an earlier part where I develop the idea and a later part that I hold back until the end. A result that reverses on the held-back part is dead. One trend-following idea (the idea that prices that have been rising tend to keep rising) looked good over the full history and then lost money on the held-back period, which meant the full-history result had never been evidence of anything.
Include the real cost of trading. When you buy, you pay a slightly higher price than when you sell at the same moment, and that gap is called the spread. Add fees on top of that. This is the check that rejects my ideas most often. A mean reversion idea (betting that a price snaps back after it moves) showed up as real in every market and timeframe I tested, and in every one of them the effect was smaller than the spread. So it was real, but there was no way to actually make money from it.
Don't count the same event many times. If one big news day produces hundreds of data points, those points aren't independent, they're really one event. Counting them separately makes a result look far more certain than it is. Two of my cleanest-looking results turned out to be exactly this, the same few events counted over and over. These days, when a result looks almost too certain, I treat that as a warning sign.
Shuffle the outcomes and run it again. I scramble which outcome goes with which piece of data and rerun the test. If the effect is still there after shuffling, it came from how I set up the test and not from any real ability to predict. Measurements that mostly just track how high a price is can produce very convincing patterns this way.
Only use information that would have been available at the time. Using information from after the moment of the decision is called look-ahead, and it's where I've been burned the worst. Both stories below are about it.
Test on the trades I'd actually take. An improvement to a strategy has to be checked on the small set of trades the strategy would really make, and not just on every possible candidate. A filter can look excellent across all the candidates and turn out to be useless, or even backwards, on the ten percent I would have acted on.
A thirty-cent gap that wasn't there
Kalshi is an exchange where each contract asks a yes-or-no question and settles at either zero or a dollar, and it has cryptocurrency contracts that expire every fifteen minutes, every hour and every day. This research was on the hourly ones. At one point during research I found what looked like a thirty-cent gap between what the market was charging for these contracts and what they actually paid out. On contracts that max out at a dollar, that would be enormous.
However, upon closer inspection I found the problem was in how my code looked up prices. Price history usually comes grouped into bars, where each bar covers a stretch of time (say one minute) and records the price at the start, the high, the low, and the price at the end. Each bar is labeled with its start time. My code asked for the price at a given minute, but it was reading the end-of-bar price from the bar labeled with that minute, which is really the price from up to a minute later. So the answer I was trying to predict was already sitting in the data I was predicting from.
The issue wasn't caught by a statistical test. It was caught by manually inspecting the data for a single contract and comparing it with what actually happened. My rule now is that the price "at" a moment comes from the latest bar that had already finished by that moment. It's a one-line rule, but it took an expensive detour to learn it.
A filter that peeked at the future
A related mistake cost less, but I learned more from it. I had a rule for picking which contracts to include, and it used each contract's highest price over its entire life to make that choice. It produced a healthy-looking profit.
The problem was that a contract's highest price over its whole life isn't known until the contract is over. The rule was quietly using the future. Contracts that went on to lose tended to have price histories that gave them away, so they were being dropped from the sample after the fact. Once I removed the peek into the future, the profit went to zero.
Look-ahead rarely announces itself. It usually shows up as a filter that sounds perfectly reasonable but happens to be calculated over the wrong window of time.
Keeping a list of what failed
The habit that ties all of this together is keeping the rejections. Every idea gets a name, a short explanation of why it might work, a verdict, and the code that produced the verdict. Most of the entries say "rejected" or "real but not profitable."
That list helps me in two ways. It stops me from re-testing ideas I've already ruled out, and it keeps my expectations realistic. If I only remembered the ideas that worked, every new idea would feel promising. Because the list is mostly failures, I've learned to be skeptical of a good new result until it has made it through every one of the checks above.