Get in touch

9io.ai / Blog

The look-ahead bias our backtester prevents, and the leaks it misses

Our daily backtester shows its detectors only past bars and fills stops pessimistically. What still leaks in, and a checklist for building your own.

Key takeaways

  • Hand each simulated day only the rows up to that day, so signals can’t read later bars.
  • On daily bars, assume the stop was hit first and fill gaps through the stop at the open.
  • Entering at the close that produced the signal assumes a price you couldn’t trade after seeing it.
  • Ranking stocks by end-of-test liquidity or today’s index list tilts a backtest towards winners.
  • Rules written during the test window make the replay in-sample, however strictly it slices the data.

The backtester in 9io Alpha, our own trading research platform, can’t see the future when it generates a signal. On each simulated day it gives the detector code only the bars up to and including that day. Later bars are used for one thing, working out how the trade would have exited. That shuts out the most common leak, an indicator that quietly reads tomorrow’s price.

Several quieter problems remain. Every trade enters at the close the signal needed to see. The stocks are picked using traded value from the end of the test, out of a list of companies that are still in the index today. Nothing is charged for costs. And the rules being replayed were written partway through the period they’re tested on. Some of these are look-ahead through a side door and the rest are other kinds of optimism, but all of them make results look better than a trader at the time could have done.

This post covers what the backtester gets right, where it leaks, and what we’d change, with a checklist for your own. It’s based on reading our two backtest code paths, the test suite and the git history. We didn’t run a backtest for it, and it contains no performance figures. It describes engineering and isn’t investment advice.

How the replay keeps tomorrow out of today’s signal

The backtester replays one stock at a time, one day at a time. For decision day t it slices the history to rows 0 to t with df.iloc[:t + 1] and passes that slice to scan_stock, the same function the live scanner calls. Seven rule-based detectors run inside it, looking for patterns such as breakouts, pullbacks and volatility squeezes. Every indicator they use, from RSI to Bollinger Band width, is computed from the slice. Rows after t never reach them. Only the exit simulation reads those rows, stepping through the next five bars the way a live trade would.

def replay(df, detect, simulate_exit, first_day, hold_days=5):
    """Walk one stock forward. detect() never sees a row after t."""
    trades, busy_until = [], {}
    for t in range(first_day, len(df) - 1):
        window = df.iloc[: t + 1]                      # today and earlier
        for sig in detect(window):
            if busy_until.get(sig.setup, -1) > t:
                continue                               # one open trade per setup
            fill = simulate_exit(df, t, sig, hold_days)  # reads t+1 .. t+hold_days
            if fill is not None:
                busy_until[sig.setup] = fill.exit_index
                trades.append(fill)
    return trades

A default run covers the 120 most liquid stocks over their last 150 trading days. That is about 18,000 calls to the detector code, each recomputing every indicator from scratch. A vectorised backtester computes each indicator once over the whole history and is much faster. It is also where look-ahead bugs breed, because any operation that touches a later row leaks. A negative shift such as shift(-1) does. So does a mean taken over the whole column, a centred rolling window, or a scaler fitted on the full sample.1

Freqtrade’s documentation describes this problem for its own engine, which loads the whole dataframe and calculates all indicators at once. The project ships a lookahead-analysis command that reruns the backtest on sliced data and compares indicator values and signals with the full run.2

We pay for the slicing in compute so that the guarantee comes from the shape of the loop. A detector can contain any bug and still can’t read tomorrow’s bar, because tomorrow’s bar isn’t in the object it receives. The guarantee only covers the loop, though. Decisions made before it starts, such as which stocks to test, are outside it, and several of the gaps below are of that kind.

Exit fills that assume the worst

A daily bar records the open, high, low and close. It doesn’t record whether the high came before the low. If a bar’s range covers both the stop and the target, the bar alone can’t tell you which order filled first, so we assume the stop did. Each bar after entry is checked in this order.

On a bar after entry We fill Reason
The open is at or below the stop At the open The price gapped through the stop, so the stop price never traded
The low reaches the stop At the stop The stop is checked first, so a bar that touches both is a loss
The open is at or above the target At the open A resting sell order below the market fills at the opening price
The high reaches the target At the target The target traded during the bar
None of these within five bars At the fifth bar’s close Time exit

Freqtrade makes the same stop-first choice within a candle, and it fills a stop exactly at the stop price even when the low went lower, with an adjustment for fees.3 Our older event-driven engine also fills at the stop price on a gap. Filling the gap at the open is the stricter rule. When a stock closes above the stop and opens well below it, a stop order fills near the open, and a backtest that books the stop price records a smaller loss than anyone took.

The rules still flatter us in three places. A bar that only touches the stop fills exactly at the stop, while a real stop order turns into a market order and fills at whatever price comes next. A bar with a missing open, high or low is skipped, so a stop that would have triggered on it is missed. And trades opened in the last five bars of data close early, at the last available close.

Scoring trades in R, and when the numbers count

Each trade is scored in R, its profit or loss divided by the distance from entry to stop. A trade stopped at the stop price scores −1R, and a gap through the stop scores worse. Expectancy is the average R per trade.

We count a 0R trade as a loss. That choice doesn’t move expectancy, which is the mean of the R values whatever you call them. It does lower the share of winning trades, and that share feeds the position-sizing maths and the live ranking, so we put zero on the conservative side.

Measured statistics replace our starting assumptions, estimates we took from published research before we had data of our own, only for setups with at least 25 closed trades. Twenty-five is a floor. For any proportion near one half, the standard error at 25 observations is √(0.25 ÷ 25) = 0.1, so a 95% interval runs roughly 0.2 either side of the estimate.

The statistics file is written atomically. The backtest writes the JSON to a temporary file and renames it over the old one, so the live scanner, which reads the file whenever it ranks signals, sees the previous version or the new one and never a mix. Our post on rate limits, holidays and atomic writes in our market data layer covers the same pattern for price caches, including the failure that a rename doesn’t protect against.

What “walk-forward” means in our code

Our backtester’s docstring calls it walk-forward, and the term has two common meanings. In Robert Pardo’s walk-forward analysis, you optimise parameters on an in-sample window, test the frozen parameters on the window that follows, then roll forward and repeat.4 López de Prado uses the same name for a historical simulation in which each decision only uses data from before it, and discusses that method’s pitfalls.5

Ours is the second kind. Nothing is fitted. The rules are fixed and the replay applies them day by day, so there is no in-sample or out-of-sample period, just one sample replayed in order.

The repository also holds an older, event-driven engine that the app doesn’t call. When asked to, it splits history into windows, 70% in-sample and 30% out-of-sample by default, and reports out-of-sample performance as a fraction of in-sample performance. It optimises nothing in between, so both halves run the same fixed rules and the ratio compares two arbitrary stretches of history. That engine fills stops at the stop price even when the price gaps through. Its decision function can also call analysers for market regime, news sentiment and options data, and none of them takes a date. They’re switched off by default. Switched on inside a backtest, each would return today’s state for every historical bar.

Entering at the close that produced the signal

Every trade enters at the close of the bar that generated its signal. The detectors need that close to decide, so the backtest assumes you could see the final price and trade at it in the same instant.

In the US, the official close comes from the closing auction that starts at 4:00 p.m. ET, and NYSE stops accepting most market-on-close and limit-on-close orders at 3:50 p.m.6 You can get the closing price, but only by committing before you know it.

In India, the close was the volume-weighted average price (VWAP) of the last 30 minutes of trading until 3 August 2026. Since then, stocks with derivative contracts close through an auction between 3:15 and 3:35 p.m., and every other stock still closes at the 30-minute VWAP.789 A half-hour average is a price that no single order, placed after the signal, could have been filled at.

The usual fix is to enter at the next bar’s open, which is what freqtrade’s backtester does by default.3 The alternative is to keep the close and charge a haircut for the drift between the decision and the fill. We plan to move to the next open and to start the stop and target checks on that same bar.

A universe chosen with hindsight

Two choices about which stocks to test use information from the end of the test, or from today.

The first is the liquidity cut. The backtester tests the 120 stocks with the highest average traded value, close times volume, over the last 60 bars of data. Those 60 bars sit inside the 150-day replay. Traded value rises with price, so a stock that rallied during the window is more likely to make the cut, and the sample leans towards the period’s winners before a single signal fires. Ranking point in time fixes it. On each decision day, use only the traded value up to that day, or choose the universe once, before the window starts.

The second is survivorship. The candidate list is today’s index membership, a static list of S&P 500 members from June 2026 for the US and the current Nifty 500 for India. A company dropped from the index, acquired or delisted during the window isn’t on the list, and its last months of trading aren’t in the test. Those are disproportionately bad outcomes. Shumway found that CRSP, the standard US research database, lacked correct delisting returns for most stocks delisted for bankruptcy and other negative reasons since 1962, and that the missing returns were large.10 In a follow-up on Nasdaq stocks, Shumway and Warther found that assigning −55% to missing performance-related delisting returns corrected the bias.11

Our price source makes this harder to fix. When Yahoo returns no prices for a symbol, yfinance raises an error that says the ticker is “possibly delisted”.12 A survivorship-free test needs point-in-time index membership and a price source that keeps companies after they stop trading, and we have neither yet.

No costs, and no limit on open trades

The setup backtester charges nothing. There is no commission, tax, spread or slippage. For rules that hold a trade for five days at most, that’s a large omission. Novy-Marx and Velikov found that most of the anomalies they studied with one-sided monthly turnover below 50% still earned statistically significant spreads after costs, at least when designed to reduce trading, and that few higher-turnover strategies did.13 A five-day holding period turns over far faster than that.

The older event-driven engine does model costs. It charges 5 basis points of fixed slippage plus a square-root market-impact term, capped at 50 basis points, and applies an Indian fee schedule covering brokerage, GST, securities transaction tax (STT) and stamp duty. If you build a schedule like that, check which rates apply to the trades you simulate. A position held overnight is a delivery trade, and according to one large broker’s published charges, checked on 8 October 2026, delivery trades pay STT of 0.1% on both the buy and the sell, and stamp duty of 0.015% on the buy.14 The intraday rates are much lower, and they are easy to copy into a model that holds positions overnight.

The second gap is capital. Each stock can hold one open trade per setup, but nothing limits how many are open across the market. On a day when dozens of stocks trigger, the backtest takes every one of them at equal weight. A real account can’t, and the trades it would take aren’t a random draw. The live scanner ranks signals by expected value and shows the top 40, with at most three per sector. The backtest’s statistics describe every signal the detectors produce, while a user only ever sees a ranked, sector-capped subset.

Rules written during the period they’re tested on

Our detectors were first written in June 2026 and last changed in July 2026. A backtest started today replays the most recent 150 trading days, which reach back to about the start of March 2026. For more than half of that window, whoever wrote the rules could see the charts they are now tested on. Whatever the replay reports for that stretch is in-sample, however strictly it hides tomorrow’s bar from today’s decision. Only the 60 or so sessions since the last change can be called out of sample, and that is too few to say much.

This is multiple testing in a quieter form. Bailey, Borwein, López de Prado and Zhu show that the more configurations a researcher tries, the more likely it is that the best backtest is overfit. With five years of daily data, they calculate that trying more than about 45 independent configurations almost guarantees one that looks good in sample and has an expected out-of-sample return on risk of zero.15 They also point out that analysts rarely report how many configurations they tried, so readers can’t judge the degree of overfitting. We can’t report it either, because nobody logged how many versions of each detector were tried before the current ones. Their method for estimating the probability of backtest overfitting needs exactly that trial history.16

Harvey, Liu and Zhu make a similar argument about academic factor research. With hundreds of factors already tested, they argue that a new one should clear a t-statistic of about 3.0 instead of the conventional 2.0.17 Our seven detectors across two markets produce 14 sets of statistics, and the live ranking favours whichever of them look best. The best of 14 noisy estimates is biased upwards.

The last gap is tests. Our test suite was last changed in March 2026, before the setup backtester and the data modules existed, and none of those modules has a test. The slicing guarantee holds today because of how the loop is written, and nothing would fail if a refactor broke it.

What we’d change, in order

The first change is a truncation test. Compute the signals for day t, rewrite every row after t, and compute them again. If anything for day t changes, something reads the future. It is freqtrade’s lookahead-analysis idea in unit-test form, and it’s cheap enough to run on every commit.

def test_day_t_signals_ignore_later_rows(daily_bars):
    t = 180
    before = signals_for_day(daily_bars, t)
    tampered = daily_bars.copy()
    tampered.iloc[t + 1:] *= 1.5          # rewrite the future
    assert signals_for_day(tampered, t) == before

Pointed at our universe selection, this test would fail today, because the 60-bar traded value at the end of the data changes when the future does. After the test, in order:

  1. Enter at the next bar’s open, and run the stop and target checks from that bar.
  2. Rank the universe point in time, and find a source for historical index membership and for prices of delisted companies.
  3. Port the cost model into the setup backtester, using delivery rates for overnight positions and adding a US schedule, each with the date it was checked.
  4. Simulate a portfolio with a cap on open positions and capital per trade, and report statistics for the ranked, sector-capped picks separately from all signals.
  5. Version the detectors, log every variant tried, and treat only data after the latest freeze as out of sample. Bailey and his co-authors cite Leinweber and Sisk’s “model sequestration”, announcing a strategy and publishing its results months later on data nobody had seen, as the extreme form of this.15

A checklist for your own backtester

  • Each decision receives rows 0 to t only, and a test proves it by rewriting the future and checking that day t doesn’t change.
  • No feature uses a statistic computed over the whole sample, such as a column mean, a z-score or a fitted scaler.
  • The entry price is one you could have traded after the signal was known.
  • When a daily bar touches both the stop and the target, the stop wins. Gaps through the stop fill at the open, and other stop fills include slippage.
  • The universe and every liquidity filter use only data available on the decision day.
  • Delisted and removed companies are in the data, with their last prices.
  • Costs are charged on both sides from your market’s current fee schedule, with the date it was checked.
  • Open positions and capital are limited, and you can score the signals you would actually have taken.
  • Every statistic has a minimum sample size, and you record how many rule variants you tried.
  • Rules are frozen and dated before the period you call out of sample.
  • In backtest mode, nothing in the decision path can fetch live data such as news or quotes.
  • Results files are written to a temporary file and renamed into place.

  1. Freqtrade documentation, “Common mistakes when developing strategies”, https://www.freqtrade.io/en/stable/strategy-customization/#common-mistakes-when-developing-strategies ↩

  2. Freqtrade documentation, “Lookahead analysis”, https://www.freqtrade.io/en/stable/lookahead-analysis/ ↩

  3. Freqtrade documentation, “Assumptions made by backtesting”, https://www.freqtrade.io/en/stable/backtesting/#assumptions-made-by-backtesting ↩↩

  4. Robert Pardo, “The Evaluation and Optimization of Trading Strategies”, 2nd edition, Wiley, 2008, chapter 11 (Walk-Forward Analysis), https://www.wiley-vch.de/en/areas-interest/finance-economics-law/the-evaluation-and-optimization-of-trading-strategies-978-0-470-12801-5 ↩

  5. Marcos López de Prado, “Advances in Financial Machine Learning”, Wiley, 2018, chapter 7 (Cross-Validation in Finance) and chapter 12 (Backtesting through Cross-Validation), https://www.wiley.com/en-us/Advances+in+Financial+Machine+Learning-p-9781119482086 ↩

  6. NYSE, “Auctions”, closing auction timeline, https://www.nyse.com/auctions ↩

  7. SEBI circular HO/47/11/11(3)2025-MRD-POD2/I/2765/2026, “Introduction of Closing Auction Session (CAS) in the Equity Cash Segment and certain modifications in the Pre-Open Auction Session”, 16 January 2026, https://www.sebi.gov.in/legal/circulars/jan-2026/introduction-of-closing-auction-session-cas-in-the-equity-cash-segment-and-certain-modifications-in-the-pre-open-auction-session_99122.html ↩

  8. SEBI, “Annual Report 2025-26”, chapter 4, box 4.1 (Introduction of Closing Auction Session), https://www.sebi.gov.in/reports-and-statistics/publications/aug-2026/Chapter%2004.pdf ↩

  9. Outlook Business, “Chaotic Debut: How SEBI’s Closing Auction Reform Fared In Its First Week”, 7 August 2026, https://www.outlookbusiness.com/markets/sebi-closing-auction-session-cas-first-week-performance-bse-nse-closing-volatility-traders-brokers-confusion ↩

  10. Tyler Shumway, “The Delisting Bias in CRSP Data”, Journal of Finance 52(1), 1997, https://ideas.repec.org/a/bla/jfinan/v52y1997i1p327-40.html ↩

  11. Tyler Shumway and Vincent A. Warther, “The Delisting Bias in CRSP’s Nasdaq Data and Its Implications for the Size Effect”, Journal of Finance 54(6), 1999, https://ideas.repec.org/a/bla/jfinan/v54y1999i6p2361-2379.html ↩

  12. yfinance on GitHub, “exceptions.py”, https://github.com/ranaroussi/yfinance/blob/main/yfinance/exceptions.py ↩

  13. Robert Novy-Marx and Mihail Velikov, “A Taxonomy of Anomalies and Their Trading Costs”, Review of Financial Studies 29(1), 2016, NBER working paper version, https://www.nber.org/papers/w20721 ↩

  14. Zerodha, “Charges”, equity delivery and equity intraday columns, checked 8 October 2026, https://zerodha.com/charges ↩

  15. David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, “Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance”, Notices of the AMS 61(5), 2014, https://www.ams.org/notices/201405/rnoti-p458.pdf ↩↩

  16. David H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, “The Probability of Backtest Overfitting”, Journal of Computational Finance 20(4), 2017, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2326253 ↩

  17. Campbell R. Harvey, Yan Liu and Heqing Zhu, “… and the Cross-Section of Expected Returns”, Review of Financial Studies 29(1), 2016, NBER working paper version, https://www.nber.org/papers/w20592 ↩

Frequently asked questions

What is look-ahead bias in a backtest?

It is any use of information in a simulated decision that wasn’t available at that moment. Common forms are indicators that read later bars, entries at prices only known after the signal, and stock lists compiled after the test period.

How can I test a backtester for look-ahead bias?

Compute the signals for a day t, then change every row after t and compute them again. If anything for day t changes, something reads the future. Freqtrade’s lookahead-analysis command applies the same idea to whole backtests.

Should a daily backtest assume the stop or the target was hit first?

The stop. A daily bar doesn’t record whether the high or the low came first, so the conservative assumption is the loss. If the bar opens beyond the stop, fill at the open.

Is a walk-forward backtest out of sample?

Only if the rules were fixed before the data being tested. In Pardo’s walk-forward analysis, parameters are optimised on one window and tested on the next. A day-by-day replay of rules written while looking at the same period is still in-sample.

How does survivorship bias affect a stock backtest?

Testing only today’s index members leaves out companies that were removed or delisted during the test, and those are disproportionately bad outcomes. Results look better than a trader at the time could have achieved.

How many trades are enough to trust backtest statistics?

There is no single number. We require 25 closed trades per setup before measured statistics replace our starting assumptions, and treat that as a floor. At 25 trades, the standard error of a proportion near one half is about 0.1.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.