Skip to main content
Track Robustness testing
56 / 57
Library

2020issue C1041

When mechanical historical tests decay after optimization

A visually tidy historical equity path is not, by itself, a reliable basis for expecting a mechanical trading system to keep working after live use. Treat that polish as a design warning, and keep simplicity, parameter restraint, and explicit abstention inside the same mechanical procedure.

  • Favorable historical results from a mechanical procedure are not, by themselves, a reliable basis for expecting the same procedure to keep producing favorable results after it is used live.
  • Searching tunable inputs for a near-perfect historical equity path is a common route to over-optimization, and adding extra rules until a historical test looks acceptable can curve-fit the procedure to the recorded sample.
  • A stated counterweight is to use only enough rules and parameter search to reach an adequate historical test, then stop, with abstention kept inside the same mechanical procedure.
  • Market change over spans of one, five, and ten years, plus an unknown luck component that may reverse, can still break a carefully developed procedure, so live use requires accepting that the historical result may not continue.
Entries in this reading3 entries

What a favorable historical test does not settle

A mechanical trading system is a fully specified mapping from rule inputs, market state, and execution constraints into a signal, including the choice to stand aside. System optimization searches those rule inputs against recorded market state and execution constraints so entry, exit, and abstention produce one testable signal over the system's holding period.

Favorable historical results from that mechanical procedure are not, by themselves, a reliable basis for expecting the same procedure to keep producing favorable results after it is used live.

How the recorded path becomes too tidy

Searching tunable inputs such as averaging lookback, stop size, or indicator thresholds for a near-perfect historical equity path is a common route to over-optimization. Over-optimization is continuing parameter search until the historical path looks near-perfect instead of merely adequate.

Adding extra rules and filters until a historical test looks acceptable can fit the procedure to the recorded sample. Curve-fitting is that habit of adding rules or filters until the recorded sample yields an acceptable historical path.

Mechanical backtest equity that later decays

The historical equity path climbs through most of the test, then the boxed later segment turns down. Points were read off the plotted curve in Figure 1; the article does not print a table, so the levels are approximate. A trader should treat that polished climb as a warning, not as proof the same mechanical system will keep working live.
The historical equity path climbs through most of the test, then the boxed later segment turns down. Points were read off the plotted curve in Figure 1; the article does not print a table, so the levels are approximate. A trader should treat that polished climb as a warning, not as proof the same mechanical system will keep working live.unspecified mechanical system (illustrative equity) · backtest bar or trade sequence

Raster digitization of a dark, unlabeled equity plot. Vertical scale is equity in the source chart’s units (axis ticks at 0, 500, 1000, 1500, 2000, 2500). Horizontal scale is bar/trade index with ticks at 0, 10, 20, 30, 40, 50. Peak near 2500 around index 32; boxed later window is the decay. Values are approximate to the nearest 50 equity units.

Stop at an adequate test

A stated counterweight is to use only enough rules and parameter search to reach an adequate historical test, then stop rather than chase a visually perfect path.

Robustness testing checks whether that same mechanical procedure still holds when extra rules, regime change, and an unknown luck component are not assumed to match the fitted sample.

Market change, luck, and abstention

Changes in market conditions and in the mix of participants over spans of one, five, and ten years can produce price behavior that no longer matches the sample used to build the rules.

Keeping the rule set simple and suspending trading in unusual conditions, including volatility at recorded highs or lows until activity returns toward a typical range, is offered as a way to limit damage from market change. Abstention is a tested rule that suspends trading when conditions leave the range the procedure was built to handle.

A successful historical test can contain an unknown luck component, and that component is not expected to persist and may reverse. Careful development can reduce but not remove these failure modes, so moving a historically tested procedure into live use still requires accepting that the historical result may not continue.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
56 of 57 in the Robustness testing track
20251-45 pp.Next on Robustness testingAdd a second procedure before you retune the firstOnce a first profitable procedure exists, field a second distinct entry-exit-abstention stack rather than keep refining the original one.
All readings on this track · 57 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
  52. 2019Noise-matched rules still need trend filters and robustness tests
  53. 2019Three gates for evaluating a trading system
  54. 2020Data construction as a mechanical system input
  55. 2020Hidden optimization in ported relative-strength systems
  56. 2020When mechanical historical tests decay after optimization
  57. 2025Add a second procedure before you retune the first
All 67 readings tagged Robustness testing
Also on Robustness testing5 readings