top of page

Backtesting Trading Strategies A Systematic Guide to Avoid Curve Fitting and Bias

11 hours ago
9 min read

A trading strategy can look brilliant on five cherry-picked charts and still fail the moment real money is involved. Backtesting is the filter between an interesting idea and a tradeable system, but only if it is done with discipline.


Good backtesting answers a practical question: would this strategy have had a positive expectancy under realistic market conditions? That means testing clear rules, using clean data, accounting for costs, and resisting the urge to keep adjusting the system until the past looks perfect.


This guide explains how to backtest trading strategies in a structured way, including manual look-backs, automated testing, the metrics that matter, and the most common traps that distort results.


This article is for education only and is not financial advice. Trading involves risk, including the risk of losing capital.


Overhead view of printed candlestick charts and a handwritten backtesting plan on a wooden kitchen table.
Start with a written plan before touching the data.

Start with a testable trading idea


A backtest is only useful when the strategy rules are specific enough for someone else to repeat. Vague ideas such as “buy strong stocks” or “sell when momentum fades” create room for hindsight. The test becomes a memory exercise rather than evidence.


A good strategy definition includes:


  • Market

    The exact instrument or universe, such as FTSE 100 shares, GBP/USD, gold, or Bitcoin.


  • Timeframe

    The chart interval and holding period, such as daily candles with trades held for 5 to 20 sessions.


  • Entry rule

    The precise condition that opens a trade.


  • Exit rule

    The precise condition that closes a trade, including profit exits, stop exits, and time exits.


  • Risk rule

    Position size, maximum risk per trade, and any exposure limits.


  • Trade direction

    Long only, short only, or both.


  • Execution assumption

    Whether trades enter at the next open, next close, limit price, or stop price.


For example, “buy when the 20-day moving average crosses above the 100-day moving average” is a clear entry condition. “Buy when price looks strong after a pullback” is not.


The same applies to exits. “Sell when the trend weakens” invites guesswork. “Exit when price closes below the 20-day moving average or after 30 trading days, whichever comes first” can be tested.


Before running any test, write the rules in a short strategy specification. If the strategy changes during the backtest, record the change separately. Do not blend old and new versions into one result.


A simple specification might look like this:


Component

Rule

Market

Liquid UK large-cap shares

Timeframe

Daily chart

Direction

Long only

Entry

Close above 50-day high

Exit

Close below 20-day low

Position size

1% account risk per trade

Costs

Spread plus commission estimate

Test period

10 years, split into in-sample and out-of-sample data


This step feels basic, but it prevents many later errors. If the rules are not fixed, the backtest can become a search for a flattering story.


Choose the right method for the strategy


There are two main ways to test a strategy: manual look-backs and automated algorithmic backtesting. Both can be useful, but they answer different questions.


Manual look-backs help you understand the behaviour


A manual look-back means scrolling through historical charts and marking trades that fit the rules. It is slow, but it builds pattern recognition and highlights practical issues that a spreadsheet may hide.


Manual testing works well when:


  • The strategy includes some visual judgement.

  • The trader is still developing the rules.

  • The goal is to understand trade context.

  • The sample size does not need to be very large.

  • The market structure matters, such as gaps, liquidity, or news spikes.


The weakness is bias. It is easy to see what happened next, even when trying not to. The eye naturally gives more weight to clean setups and ignores messy ones. Manual testing also tends to produce small sample sizes, which can make random luck look like skill.


To improve manual tests, move candle by candle, hide future price action where possible, and record every qualifying trade. Do not skip the ugly trades because they “would not have been taken live” unless that filter was written into the rules before the test.


Automated backtesting measures rules at scale


Automated testing uses code or specialist software to apply rules across historical data. It is the standard method for rules-based systems because it can test hundreds or thousands of trades without emotional filtering.


Automated testing works well when:


  • Entry and exit rules are objective.

  • The strategy needs a large sample.

  • Many instruments or time periods must be tested.

  • Costs, slippage, and position sizing need consistent treatment.

  • The trader wants to compare variations.


The weakness is false precision. A backtest can produce exact-looking numbers that depend on flawed assumptions. Bad data, unrealistic fills, survivorship bias, or look-ahead logic can all create results that never could have happened in live trading.


The best process often combines both methods. Use manual review to understand how the strategy behaves. Use automated testing to measure it at scale. Then inspect a sample of automated trades manually to check whether the engine is behaving as intended.


This is where many quantitative trading tips, historical data testing, manual vs automated backtesting discussions overlap: the method matters, but the discipline around the method matters more.


Close-up view of a hand marking trade entries and exits on a printed price chart.
Manual review can reveal market context that raw statistics miss.

Build the backtest like a controlled experiment


A useful backtest separates idea generation from validation. If the same data is used to invent, tune, and approve a strategy, the results are usually too optimistic.


Split the data into clear periods


Use at least two periods:


Period

Purpose

How to use it

In-sample data

Build and tune the strategy

Test the first version and make limited rule changes

Out-of-sample data

Validate the strategy

Run once after the rules are fixed

Forward test

Observe live or demo performance

Confirm execution and behaviour without risking full capital


The out-of-sample period matters because it acts like unseen data. A strategy does not need to perform identically in both periods, but the broad behaviour should make sense. If the in-sample test shows smooth gains and the out-of-sample test collapses, the model may have learned noise.


A more advanced version is walk-forward testing. The strategy is tuned on one period, tested on the next, then rolled forward. This better reflects the way a live system would be maintained over time.


Include realistic trading costs


Costs can turn a profitable backtest into a losing strategy. Include:


  • Spread

  • Commission

  • Slippage

  • Stamp duty where relevant for UK share dealing

  • Financing or overnight costs for leveraged products

  • Borrow costs for short selling, if applicable


Short-term strategies are especially sensitive to costs. A system that makes many small trades needs very accurate assumptions. If the average trade only earns a small amount before costs, it may not survive real execution.


Use trade timing that could really happen


Many backtests accidentally assume impossible fills. A signal based on the closing price cannot also enter at that same closing price unless the rules explain how the order is placed before the close.


Common timing choices include:


  • Signal on today’s close, enter at next open.

  • Signal intraday, enter when a stop or limit is touched.

  • Signal at close, enter using a market-on-close order if the platform and market support it.


Be strict. If the strategy uses end-of-day data, assume the trade happens after that data is available.


Check the trade list, not just the equity curve


The equity curve is the headline, but the trade list reveals the mechanics. Review individual trades and confirm:


  • Entries match the written rules.

  • Exits trigger correctly.

  • Stops and targets use the right prices.

  • Position sizes change as intended.

  • No trades appear before an instrument existed.

  • No trades occur at prices outside the day’s range.


A strange trade list is a warning sign. Fix the logic before trusting the performance summary.


Track the metrics that show risk and quality


Profit alone is not enough. A strategy that made 40% with a 60% drawdown is very different from one that made 25% with a 10% drawdown. The goal is to understand both return and pain.


Use a clean tracking format for every test. Keep it consistent so different strategies and versions can be compared.


Metric

What it shows

How to read it

Net profit

Total gain after costs

Useful, but weak without risk context

Maximum drawdown

Largest peak-to-trough fall

Measures the worst historical pain

Win rate

Percentage of winning trades

Needs average win and loss to mean anything

Average win

Mean size of winning trades

Shows reward when right

Average loss

Mean size of losing trades

Shows cost when wrong

Profit factor

Gross profit divided by gross loss

Above 1 means gross gains exceeded gross losses

Expectancy

Average expected result per trade

Combines win rate, wins, and losses

Number of trades

Sample size

Low counts are less reliable

Average holding period

Time in trade

Helps match strategy to capital and temperament

Exposure

Time or capital committed

Shows how much risk was actually in the market

Return to drawdown ratio

Net return divided by maximum drawdown

Helps compare return quality


Profit factor is useful, but do not treat it as a magic number. A very high profit factor from a small sample can be less reliable than a moderate figure from hundreds of trades across different market conditions.


Win rate also needs context. A trend-following system may win less than half the time and still work if winners are much larger than losers. A mean-reversion system may win often but suffer occasional large losses. Neither profile is better by default. The question is whether the risk is acceptable and repeatable.


Use a version log as well:


Version

Change made

Test period

Reason for change

Result

V1

Original rules

2014 to 2023

Initial test

Baseline

V2

Added volatility filter

2014 to 2023

Reduce chop

Compare with V1

V3

Changed exit length

2014 to 2023

Improve drawdown

Check out-of-sample


This record stops “strategy drift”, where many tiny changes create a final system that no longer resembles the original idea.


Eye-level view of a notebook showing a trading metrics table beside a calculator and printed equity curve.
Track risk and return in the same place.

Avoid the errors that make backtests lie


Most bad backtests fail for the same reasons. The strategy may not be the problem. The test design may be giving false comfort.


Curve fitting makes the past look too perfect


Curve-fitting happens when a strategy is adjusted so closely to past data that it captures noise rather than a repeatable edge. It often appears after repeated testing of many parameters.


For example, a trader tests moving average lengths from 5 to 200 and finds that a 37-day average with a 123-day average produced the best result on one market over one period. That combination may have no real meaning. It may simply be the best fit to that historical sample.


Warning signs include:


  • Very specific parameter values with no clear logic.

  • Great performance on one instrument, poor performance elsewhere.

  • Strong in-sample results, weak out-of-sample results.

  • A small number of trades driving most of the profit.

  • Constant rule changes after every losing period.

  • Performance that depends on excluding awkward years.


To reduce curve-fitting, use simple rules, test across different market regimes, keep parameter ranges sensible, and prefer broad areas of performance over one perfect setting. If 40 to 60 days all work reasonably well, that is more convincing than one isolated best value.


Look-ahead bias uses information that was not available


Look-ahead bias occurs when a backtest uses future information to make a past decision. It can be obvious or subtle.


Common examples include:


  • Entering at today’s close using a signal that requires today’s close.

  • Using revised economic data as if it was known at the time.

  • Ranking shares using financial statement data before it was released.

  • Using the day’s high or low to decide whether a trade should have been entered earlier that same day.

  • Building a universe from today’s index members and testing it historically.


The fix is simple in principle: every decision must use only data available at that date and time. In practice, this requires careful coding and good data handling.


Survivorship bias removes the failures


Survivorship bias appears when the test only includes instruments that still exist. If a share was delisted, merged, or went bust, it may disappear from the dataset. That makes the past look safer than it was.


This is common in share strategies. Testing only current index members over a long history can overstate returns because weaker past members are missing.


Use survivorship-bias-free data where possible. If that is not available, be cautious with conclusions.


Data quality errors distort signals


Historical data can include missing candles, bad ticks, split errors, dividend adjustments, timezone issues, and incorrect highs or lows. One bad price can trigger a false trade or inflate results.


Before trusting a backtest:


  • Inspect price charts around large wins and losses.

  • Check for corporate action adjustments.

  • Compare suspicious prices with another data source.

  • Confirm timezone alignment for multi-market strategies.

  • Remove or correct clear errors, but document the change.


Ignoring liquidity creates impossible trades


A strategy may look profitable on small or illiquid instruments, but the trade size may be unrealistic. If the system assumes it can buy or sell at the recorded price without moving the market, results may be overstated.


Add liquidity filters such as minimum average volume, maximum position size as a share of daily volume, and realistic slippage assumptions.


Wide-angle view of a single trader testing a strategy on a laptop in a quiet home dining area.
Automated testing is useful only when the assumptions are realistic.

Decide whether the strategy is ready for real capital


A backtest should not end with excitement. It should end with a decision.


Ask these questions before risking money:


  • Does the strategy have fixed rules?

  • Did the backtest include realistic costs?

  • Was there a clean out-of-sample test?

  • Is the number of trades large enough to trust?

  • Did performance survive different market regimes?

  • Is the drawdown acceptable in real life?

  • Are the worst trades explainable?

  • Can the strategy be executed with the planned account size?

  • Does the edge remain after slippage and errors?

  • Has it been forward tested?


A sensible next step is a paper trade or very small live test. This checks execution, platform behaviour, spreads, order handling, and emotional pressure. Live conditions often reveal issues that historical testing misses.


Backtesting is not proof that a strategy will make money. It is a way to reject weak ideas, understand risk, and prepare for uncertainty. The value comes from the process: clear rules, honest data, realistic assumptions, and careful records.


The best backtest is not the one with the smoothest equity curve. It is the one that survives scrutiny and still looks tradeable after costs, drawdowns, mistakes, and unseen data.


 
 
 

Comments


Top Stories

Bring Trade stories straight to your inbox. Sign up for our weekly newsletter.

  • Instagram
  • Facebook
  • Twitter

© 2035 by The Global Morning. Powered and secured by Wix

bottom of page