How to backtest a trading strategy without fooling yourself
A step-by-step backtesting method, the six biases that make bad strategies look good, and the spreadsheet layout and formulas to run one honestly.
A backtest applies a trading rule to historical data and records what would have happened: every entry, every exit, every profit and loss. It is the cheapest way to find out that an idea does not work.
It is also the easiest way to convince yourself that a bad idea does work. A backtest is only as honest as its assumptions. This guide walks through a method that keeps it honest.
Step 1: Write the rule down completely
Before touching data, write the strategy so precisely that two people would produce identical trades from it. A complete rule specifies:
- Universe. Which instruments are eligible? ("All settled Kalshi NFL game markets from the 2023 season onward.")
- Signal. What triggers a trade, using only information available at that moment?
- Entry price. Exactly which price you pay: last trade, best ask, mid, next open?
- Exit. When and at what price you leave. For prediction markets, often "hold to settlement."
- Size. How much per trade.
- Costs. Fees and spread, per trade.
Vague rules let you make decisions with hindsight, often without noticing.
Step 2: Collect complete data
Gather every instrument in the universe, including the ones that were delisted, went to zero or settled against you. A dataset that only contains survivors produces a backtest that only contains winners.
For each observation you need the fields the rule uses, timestamped, plus the outcome. In a spreadsheet, one row per market or per day is the cleanest layout.
| Column | Example |
|---|---|
| Ticker | KXNFLGAME-... |
| Signal time | 2025-11-02 12:00 |
| Price at signal | 0.68 |
| Outcome (1 = YES) | 1 |
| Fee per contract | 0.02 |
Step 3: Apply the rule row by row
Add columns that compute, for each row, whether the rule trades and what happened. For a "buy YES on favorites priced under 70 cents and hold to settlement" rule:
F2 Trade? =IF(AND(C2>=0.5, C2<0.7), 1, 0)
G2 P&L / $1 =IF(F2=1, D2 - C2 - E2, 0)
Since a YES contract pays $1 if the event happens, the profit on one contract bought at price C2 is 1 − C2 on a win and −C2 on a loss. That simplifies to D2 − C2 when D2 is the outcome (1 or 0). Subtract the fee and you have net P&L.
Step 4: Summarize
Put the results in summary cells:
Trades =SUM(F:F)
Win rate =AVERAGEIFS(D:D, F:F, 1)
Avg price paid =AVERAGEIFS(C:C, F:F, 1)
Expectancy =AVERAGEIFS(G:G, F:F, 1)
Total P&L =SUM(G:G)
Expectancy, the average net profit per trade, is the single most important number. If it is not clearly positive after costs, nothing else matters. See Expected value and edge for how to read it.
The six biases that break backtests
1. Lookahead bias
Using information that was not available when the trade would have been made. Classic examples: using the day's closing price to decide a trade at the open, or using a final-season statistic to bet on week 3. In prediction markets, a common version is using the last traded price, which may have printed after the outcome was effectively known.
Fix: For every input, ask "could I have known this at the signal time?" Use prices from a fixed time before the event, not the last price before settlement.
2. Survivorship bias
Testing only on instruments that still exist. A stock backtest run on today's index members ignores every company that went bankrupt.
Fix: Build the universe as it was at each point in time, including everything that later disappeared.
3. Overfitting (data snooping)
Trying many variations and keeping the one that looked best. With enough parameters, any dataset can be fit perfectly, and the "strategy" is just a description of past noise.
Fix: Limit the number of variations you try, prefer simple rules with few parameters, and keep test data untouched (step 5).
4. Ignoring costs
Fees, bid-ask spreads and slippage are paid on every trade. A strategy with a 1-cent expected edge per contract and a 2-cent round-trip cost loses money.
Fix: Model the fee schedule of the venue you would actually trade on, and assume you buy at the ask and sell at the bid, not at the mid.
5. Unrealistic fills
Assuming you could trade any size at the quoted price. Thin markets move when you trade in them.
Fix: Cap position size relative to the volume or quoted depth, and treat results that depend on large fills in illiquid markets with suspicion.
6. Small samples
Twenty trades can show almost any win rate by chance. Results driven by a handful of big winners are fragile.
Fix: Report the number of trades with every result, and compute how much the result changes if you remove the best few trades.
Step 5: Split the data
Divide your history into two parts before you start:
- In-sample (design) data. Typically the earlier 60 to 70%. Explore, tune and design here.
- Out-of-sample (test) data. The rest. Run the final rule on it once.
If the out-of-sample result is much worse than in-sample, the strategy was probably fitted to noise. A stricter variant is walk-forward testing: design on years 1–3, test on year 4; design on years 2–4, test on year 5; and so on.
Step 6: Judge the result
A good backtest report answers these questions:
| Metric | What it tells you |
|---|---|
| Number of trades | Whether the sample is big enough to mean anything |
| Expectancy after costs | Whether there is an edge at all |
| Win rate and average win/loss | How the edge is delivered (many small wins vs. few big ones) |
| Maximum drawdown | The worst peak-to-trough loss you would have lived through |
| Result without top 5 trades | How dependent the result is on luck |
| Out-of-sample vs. in-sample | Whether the edge survived data it was not designed on |
Win rate on its own is misleading. A strategy that buys 90-cent favorites can win 88% of the time and still lose money, because each loss costs nine times what each win earns.
Backtesting in a spreadsheet vs. code
Code scales to millions of rows and complex logic. Spreadsheets are slower but make every step inspectable, which is exactly where backtests fail: a hidden lookahead, a missing fee, a filter that dropped the losers. For many prediction-market and daily-frequency questions, a spreadsheet holds the full dataset comfortably.
In Jordan, you can describe the rule in a sentence. Jordan collects the full market history for the series onto a sheet, adds the rule and P&L columns as formulas, and builds the summary, showing you each change for review first. Because the data is in cells, you can audit any row.
Frequently asked questions
How much historical data do I need for a backtest?
Enough trades for the result to be distinguishable from luck. As a rough guide, a few hundred trades lets you detect a meaningful edge; a few dozen rarely does. More important than calendar length is covering different conditions, such as different seasons, volatility regimes or event types.
What is a good win rate for a trading strategy?
There is no good win rate in isolation. What matters is win rate combined with the average size of wins and losses, which together give expectancy. Trend-following strategies are often profitable with win rates below 40%.
Why did my strategy work in the backtest but not live?
The usual causes are lookahead bias, overfitting, underestimated costs, fills that were not achievable, or a change in the market since the test period. Paper trading or very small live positions help reveal these gaps before they become expensive.
What is walk-forward analysis?
Walk-forward analysis repeatedly designs a strategy on one window of history and tests it on the period immediately after, then rolls both windows forward. It shows whether a strategy keeps working on genuinely unseen data rather than on one lucky split.