Skip to main content
Track Robustness testing
35 / 51
Library

2006issue C021-3

Reserved-sample evaluation of trading system design

A finished backtest result is not treated as a sufficient evaluation of a trading system. Unused history is spent in a declared order so system-optimization, robustness-testing, and walk-forward analysis can be judged as one procedure, with unseen data held out so locked rules cannot be rewritten after every leftover outcome has already been seen.

  • A finished backtest result is not treated as a sufficient evaluation of a trading system; the procedure used to obtain that result is what must be judged.
  • Concept, initial rules, and degrees of freedom are written down first, and only those declared degrees of freedom may be varied later in the project.
  • System-optimization and robustness-testing spend only the test sample; walk-forward analysis then applies locked degrees of freedom to unused reserved data.
  • After repeated refine-test-walk-forward loops, unseen data remain as a last reserved sample so hindsight stays out of the final survive-or-abandon decision.
Entries in this reading3 entries

The procedure is what must be judged

A finished backtest result is not treated as a sufficient evaluation of a trading system. The procedure used to obtain that result is what must be judged.

Concept, initial rules, and degrees of freedom are written down first. Degrees of freedom are the rule inputs a developer is allowed to vary during a project; anything not declared stays fixed. Only those declared degrees of freedom may be varied later in the project.

A stated split of the historical pool

One stated split of the historical pool is 5% build data, 40% test data, 40% walk-forward data, and 15% unseen data, with a single symbol preferably assigned to at least three of those roles.

Build data are used only to confirm that coded orders match the intended rules before system-optimization begins. System-optimization is a search over declared degrees of freedom on a test sample to choose rule inputs before any later reserved-data evaluation.

Choose a robust set on the test sample

System-optimization on the test sample chooses a robust set of degrees of freedom. Robustness-testing is a check that acceptable rule inputs appear across a band of nearby settings rather than at a single tuned point. A wide band of acceptable settings is treated as more informative than a single tuned combination.

Walk-forward analysis on unused data

After the test-sample backtest is judged acceptable enough to continue evaluation, walk-forward analysis applies those locked degrees of freedom to unused walk-forward data. Walk-forward analysis is a locked-parameter evaluation on reserved data that asks whether the same entry, exit, and abstention rules still produce an acceptable signal after optimization has finished. If the reserved result is not acceptable, the project returns to refinement.

When reserved samples become contaminated

After five refine-test-walk-forward loops, the walk-forward sample is treated as contaminated by hindsight and is no longer trusted as an independent evaluation. Data-contamination is the loss of independence that occurs when reserved data are reused after their outcomes have already been seen.

Unseen data are a last reserved sample held out so repeated walk-forward trials do not consume every unused observation. Unseen data function as a second reserved walk-forward sample so earlier loops need not spend the last unused history. A negative result there ends the project because remaining observations are then considered spoiled.

Rejection is the expected reserved-data outcome

A development process is expected to reject most coded systems at the reserved-data stages. Keeping hindsight out of the final survive-or-abandon decision is treated as more important than producing a passing result on every project.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
35 of 51 in the Robustness testing track
20061-5 pp.Next on Robustness testingWalk-forward critique of hindsight crossover systemsA common design error is to certify entry and exit rules after those rules were chosen in hindsight on the same historical series later treated as proof.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings