Skip to main content
Track Robustness testing
34 / 51
Library

2005issue C021-3

Grade backtested signals with holdouts and optimization plateaus

Treat a backtested result as a signal only when a walk-forward holdout, a short statistically meaningful input list, and a wide profitable optimization plateau all support the same entry, exit, and abstention procedure.

  • Hold unused historical segments out of the optimization search so finished rules can be checked on market conditions that did not choose the inputs.
  • Keep the input list short and statistically meaningful, because adding many variables makes it easier to overfit even a random price series.
  • Treat isolated spikes on an optimization curve as curve-fitting. A robust input stays profitable across a wide band of reasonable values.
  • Editorial classroom rule: the holdout, the short input list, and the wide plateau must all support the same entry, exit, and abstention procedure before the backtested result is treated as a signal.
Entries in this reading3 entries

What must agree before a result is a signal

Editorial framing: treat this archive material as a three-gate classroom audit of system design. A backtested result is treated as a signal only when a walk-forward holdout, a short statistically meaningful input list, and a wide profitable plateau on the optimization curve all support the same entry, exit, and abstention procedure.

System optimization searches rule inputs, market state, and execution constraints as one backtested procedure whose output is a signal over the system holding period. Robustness testing checks whether that same procedure still holds when inputs stay few, statistically meaningful, and profitable across a wide value range. Walk-forward analysis withholds unused historical segments from the optimization search so the signal can be judged on market conditions that did not select its inputs.

System optimization is described as searching for effective input values. Unused historical segments are to be held out of that search so the finished rules can be checked on market conditions that did not choose the inputs.

Walk-forward analysis is presented as reserving unused history, including leaving out the most recent year, while optimizing earlier data. The archive presents this check as serving the same function as observing a completed system for a year before implementation.

Keep the input list short and meaningful

A low input count is identified as a main defense against curve-fitting, because adding many variables makes it easier to overfit even a random price series.

Robustness testing is framed as preferring statistically significant, predictive market drivers over stacked smoothed price indicators, which are described as easy to fit to a single series.

Non-price candidates listed for investigation include volatility, interest rates, volume, intermarket and sector links, gaps, bar relationships, breadth, up and down volume, tape measures, put/call ratios, and cash-versus-futures relationships.

Read the optimization curve for a wide plateau

An optimization curve is a results plot across a range of input values, used to see whether profitability is isolated in spikes or persists over a broad band. Isolated spikes of profitability on that curve are treated as evidence that an input lacks predictive ability, and selecting one of those spikes is labeled curve-fitting.

A robust input is described as one whose results stay profitable across a wide band of reasonable values, with the stated design goal that the optimization curve remain above the zero line.

Reasonable input values are those that produce a commonsense trade count for the study horizon and average net per trade. That band is the reasonable input range. Overtrading is associated with slippage, and undertrading with too few observations to be statistically meaningful.

One procedure has to pass every gate

A sports-rule metaphor shows robustness shrinking as extra conditions are added. A filter that wins only when jersey numbers sum to 9, 21, or 42 is treated as no more predictive than a 9-, 21-, or 42-bar moving average.

Editorial classroom rule: the walk-forward holdout, the short input list, and the wide profitable plateau must all support the same entry, exit, and abstention procedure. If one gate fails, the backtested result is not treated as a signal.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
34 of 51 in the Robustness testing track
20061-3 pp.Next on Robustness testingReserved-sample evaluation of trading system designA finished backtest result is not treated as a sufficient evaluation of a trading system; the procedure used to obtain that result is what must be judged.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings