2005issue C021-3
Grade backtested signals with holdouts and optimization plateaus
Treat a backtested result as a signal only when a walk-forward holdout, a short statistically meaningful input list, and a wide profitable optimization plateau all support the same entry, exit, and abstention procedure.
- Hold unused historical segments out of the optimization search so finished rules can be checked on market conditions that did not choose the inputs.
- Keep the input list short and statistically meaningful, because adding many variables makes it easier to overfit even a random price series.
- Treat isolated spikes on an optimization curve as curve-fitting. A robust input stays profitable across a wide band of reasonable values.
- Editorial classroom rule: the holdout, the short input list, and the wide plateau must all support the same entry, exit, and abstention procedure before the backtested result is treated as a signal.
What must agree before a result is a signal
Editorial framing: treat this archive material as a three-gate classroom audit of system design. A backtested result is treated as a signal only when a walk-forward holdout, a short statistically meaningful input list, and a wide profitable plateau on the optimization curve all support the same entry, exit, and abstention procedure.
System optimization searches rule inputs, market state, and execution constraints as one backtested procedure whose output is a signal over the system holding period. Robustness testing checks whether that same procedure still holds when inputs stay few, statistically meaningful, and profitable across a wide value range. Walk-forward analysis withholds unused historical segments from the optimization search so the signal can be judged on market conditions that did not select its inputs.
Hold unused history out of the search
System optimization is described as searching for effective input values. Unused historical segments are to be held out of that search so the finished rules can be checked on market conditions that did not choose the inputs.
Walk-forward analysis is presented as reserving unused history, including leaving out the most recent year, while optimizing earlier data. The archive presents this check as serving the same function as observing a completed system for a year before implementation.
Keep the input list short and meaningful
A low input count is identified as a main defense against curve-fitting, because adding many variables makes it easier to overfit even a random price series.
Robustness testing is framed as preferring statistically significant, predictive market drivers over stacked smoothed price indicators, which are described as easy to fit to a single series.
Non-price candidates listed for investigation include volatility, interest rates, volume, intermarket and sector links, gaps, bar relationships, breadth, up and down volume, tape measures, put/call ratios, and cash-versus-futures relationships.
Read the optimization curve for a wide plateau
An optimization curve is a results plot across a range of input values, used to see whether profitability is isolated in spikes or persists over a broad band. Isolated spikes of profitability on that curve are treated as evidence that an input lacks predictive ability, and selecting one of those spikes is labeled curve-fitting.
A robust input is described as one whose results stay profitable across a wide band of reasonable values, with the stated design goal that the optimization curve remain above the zero line.
Reasonable input values are those that produce a commonsense trade count for the study horizon and average net per trade. That band is the reasonable input range. Overtrading is associated with slippage, and undertrading with too few observations to be statistically meaningful.
One procedure has to pass every gate
A sports-rule metaphor shows robustness shrinking as extra conditions are added. A filter that wins only when jersey numbers sum to 9, 21, or 42 is treated as no more predictive than a 9-, 21-, or 42-bar moving average.
Editorial classroom rule: the walk-forward holdout, the short input list, and the wide profitable plateau must all support the same entry, exit, and abstention procedure. If one gate fails, the backtested result is not treated as a signal.
All readings on this track · 51 readings
- 1986Degrees of freedom in trading system optimization
- 1988Walk-forward and neighborhood tests after optimization
- 1988Undisclosed rules block system robustness tests
- 1988Testing re-optimization calendars against random parameter controls
- 1989Binary search limits on multi-peak average grids
- 1989Parameter neighborhoods that survive a shift
- 1990Use profit mapping to keep a cycle and stop plateau
- 1990Why popular indicator optimization fails robustness
- 1991Retesting weighted indicator balances across horizons
- 1992Constructing forecast models with regression, walk-forward, and robustness
- 1992Diagnose regimes before you lock parameters
- 1992When stops change system timing
- 1993Walk-forward halt rules for forecast models
- 1994Walk-forward evaluation of genetic index rules
- 1995Input pruning as walk-forward system evaluation
- 1995Critiquing neural nets as incomplete trading systems
- 1996Rebuild the equity-path ratio before it ranks a designed system
- 1996Parameter grids can fit random walks
- 1996Walk-forward analysis belongs in the design of a mechanical trading system
- 1997When a holdout fails, discard the rule set
- 1997Test rewarded rule breaks before replacing the system
- 1997Walk-forward rules keep system research from rewriting live trades
- 1999Keep a channel-breakout to two lookbacks and test neighbor stability
- 1999Constant investment size in stock system evaluation
- 2000Forcing optimization maps mechanical system failure boundaries
- 2000Robust parameter selection with surface charts
- 2001A two-gate classroom test for a two-window momentum trend filter
- 2002How a two-sided continuation factor becomes a testable trend rule
- 2002Evaluating two-window trend intensity as a reversal rule
- 2003Discounting speculative bubbles in system robustness tests
- 2003Walk-forward evaluation of locked stochastic oscillator rules
- 2003Critiquing mechanical system design after extreme price regimes
- 2004Evaluating a two-window trend trigger
- 2005Grade backtested signals with holdouts and optimization plateaus
- 2006Reserved-sample evaluation of trading system design
- 2006Walk-forward critique of hindsight crossover systems
- 2008Condition-matched walk-forward evaluation for mechanical systems
- 2011Session-split evaluation of regular and overnight systems
- 2012Walk-forward evaluation as operator rehearsal
- 2013Two-window evaluation of mechanical trading systems
- 2013Walk-forward filter selection for repeated-median velocity
- 2014Walk-forward evaluation for fading-memory velocity systems
- 2015Test oscillator events before tuning rules
- 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
- 2016Walk-forward optimization without curve fitting
- 2017Optimization without overfitting in trend-system evaluation
- 2017Parameter stability is a better guide than a larger crossover grid
- 2018Point-in-time universes for system evaluation
- 2018Walk-forward robustness evaluation for optimized systems
- 2018Critiquing breakout systems through robustness tests
- 2018A critique of parameter fitting in system design