2006issue C021-3
Reserved-sample evaluation of trading system design
A finished backtest result is not treated as a sufficient evaluation of a trading system. Unused history is spent in a declared order so system-optimization, robustness-testing, and walk-forward analysis can be judged as one procedure, with unseen data held out so locked rules cannot be rewritten after every leftover outcome has already been seen.
- A finished backtest result is not treated as a sufficient evaluation of a trading system; the procedure used to obtain that result is what must be judged.
- Concept, initial rules, and degrees of freedom are written down first, and only those declared degrees of freedom may be varied later in the project.
- System-optimization and robustness-testing spend only the test sample; walk-forward analysis then applies locked degrees of freedom to unused reserved data.
- After repeated refine-test-walk-forward loops, unseen data remain as a last reserved sample so hindsight stays out of the final survive-or-abandon decision.
The procedure is what must be judged
A finished backtest result is not treated as a sufficient evaluation of a trading system. The procedure used to obtain that result is what must be judged.
Concept, initial rules, and degrees of freedom are written down first. Degrees of freedom are the rule inputs a developer is allowed to vary during a project; anything not declared stays fixed. Only those declared degrees of freedom may be varied later in the project.
A stated split of the historical pool
One stated split of the historical pool is 5% build data, 40% test data, 40% walk-forward data, and 15% unseen data, with a single symbol preferably assigned to at least three of those roles.
Build data are used only to confirm that coded orders match the intended rules before system-optimization begins. System-optimization is a search over declared degrees of freedom on a test sample to choose rule inputs before any later reserved-data evaluation.
Choose a robust set on the test sample
System-optimization on the test sample chooses a robust set of degrees of freedom. Robustness-testing is a check that acceptable rule inputs appear across a band of nearby settings rather than at a single tuned point. A wide band of acceptable settings is treated as more informative than a single tuned combination.
Walk-forward analysis on unused data
After the test-sample backtest is judged acceptable enough to continue evaluation, walk-forward analysis applies those locked degrees of freedom to unused walk-forward data. Walk-forward analysis is a locked-parameter evaluation on reserved data that asks whether the same entry, exit, and abstention rules still produce an acceptable signal after optimization has finished. If the reserved result is not acceptable, the project returns to refinement.
When reserved samples become contaminated
After five refine-test-walk-forward loops, the walk-forward sample is treated as contaminated by hindsight and is no longer trusted as an independent evaluation. Data-contamination is the loss of independence that occurs when reserved data are reused after their outcomes have already been seen.
Unseen data are a last reserved sample held out so repeated walk-forward trials do not consume every unused observation. Unseen data function as a second reserved walk-forward sample so earlier loops need not spend the last unused history. A negative result there ends the project because remaining observations are then considered spoiled.
Rejection is the expected reserved-data outcome
A development process is expected to reject most coded systems at the reserved-data stages. Keeping hindsight out of the final survive-or-abandon decision is treated as more important than producing a passing result on every project.
All readings on this track · 51 readings
- 1986Degrees of freedom in trading system optimization
- 1988Walk-forward and neighborhood tests after optimization
- 1988Undisclosed rules block system robustness tests
- 1988Testing re-optimization calendars against random parameter controls
- 1989Binary search limits on multi-peak average grids
- 1989Parameter neighborhoods that survive a shift
- 1990Use profit mapping to keep a cycle and stop plateau
- 1990Why popular indicator optimization fails robustness
- 1991Retesting weighted indicator balances across horizons
- 1992Constructing forecast models with regression, walk-forward, and robustness
- 1992Diagnose regimes before you lock parameters
- 1992When stops change system timing
- 1993Walk-forward halt rules for forecast models
- 1994Walk-forward evaluation of genetic index rules
- 1995Input pruning as walk-forward system evaluation
- 1995Critiquing neural nets as incomplete trading systems
- 1996Rebuild the equity-path ratio before it ranks a designed system
- 1996Parameter grids can fit random walks
- 1996Walk-forward analysis belongs in the design of a mechanical trading system
- 1997When a holdout fails, discard the rule set
- 1997Test rewarded rule breaks before replacing the system
- 1997Walk-forward rules keep system research from rewriting live trades
- 1999Keep a channel-breakout to two lookbacks and test neighbor stability
- 1999Constant investment size in stock system evaluation
- 2000Forcing optimization maps mechanical system failure boundaries
- 2000Robust parameter selection with surface charts
- 2001A two-gate classroom test for a two-window momentum trend filter
- 2002How a two-sided continuation factor becomes a testable trend rule
- 2002Evaluating two-window trend intensity as a reversal rule
- 2003Discounting speculative bubbles in system robustness tests
- 2003Walk-forward evaluation of locked stochastic oscillator rules
- 2003Critiquing mechanical system design after extreme price regimes
- 2004Evaluating a two-window trend trigger
- 2005Grade backtested signals with holdouts and optimization plateaus
- 2006Reserved-sample evaluation of trading system design
- 2006Walk-forward critique of hindsight crossover systems
- 2008Condition-matched walk-forward evaluation for mechanical systems
- 2011Session-split evaluation of regular and overnight systems
- 2012Walk-forward evaluation as operator rehearsal
- 2013Two-window evaluation of mechanical trading systems
- 2013Walk-forward filter selection for repeated-median velocity
- 2014Walk-forward evaluation for fading-memory velocity systems
- 2015Test oscillator events before tuning rules
- 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
- 2016Walk-forward optimization without curve fitting
- 2017Optimization without overfitting in trend-system evaluation
- 2017Parameter stability is a better guide than a larger crossover grid
- 2018Point-in-time universes for system evaluation
- 2018Walk-forward robustness evaluation for optimized systems
- 2018Critiquing breakout systems through robustness tests
- 2018A critique of parameter fitting in system design