Skip to main content
Track Robustness testing
39 / 51
Library

2012issue C0110-11

Walk-forward evaluation as operator rehearsal

A system evaluation is treated as incomplete if it shows only in-sample historical results and omits a substantial walk-forward run on previously unused data. The worked run booked every entry, exit and abstention after the cash session closed, then stressed the same unchanged rules for robustness and process-fitness.

  • An evaluation that shows only in-sample historical results is treated as incomplete unless it also includes a substantial walk-forward run on previously unused data.
  • The worked walk-forward booked signals after each weekday cash session closed, between 0930 and 1600 Eastern, under pre-set commission, slippage, contract-size and starting-balance constraints.
  • Robustness testing asks whether the same procedure stays intact after an equity peak, through an 18.7 percent decline lasting 57 trades, and at a 10000 starting balance that turns that path into a 36.3 percent loss.
  • Process-fitness asks whether a near-45-percent win rate and as many as nine consecutive losses would cause the operator to abandon or rewrite signals, or to delegate execution without override.
Entries in this reading3 entries

A walk-forward run completes the evaluation

A system evaluation is treated as incomplete if it shows only in-sample historical results and omits a substantial walk-forward run on previously unused data. In-sample testing is development and tuning on historical data the designer has already seen. Walk-forward evaluation applies a finished rule set to data withheld from design and books every entry, exit and abstention under stated costs, size limits and session hours.

Design used a six-month historical window. The same rules were then left unchanged and logged forward. That out-of-sample testing checks whether the market concept still produced a usable signal stream.

Session-close logging and stated constraints

The worked walk-forward booked signals after each weekday cash session closed, between 0930 and 1600 Eastern, under pre-set commission, slippage, contract-size and starting-balance constraints. Session-close logging records each day's signals only after the cash session ends so the evaluation uses completed bars rather than unfinished prices.

Robustness testing of path and funding

Robustness testing stresses the same procedure across holdouts, cost assumptions, starting-capital levels and adverse equity paths to see whether it remains intact. The robustness review asks whether an operator who began after an equity peak could continue through an 18.7 percent decline lasting 57 trades, rather than judging the path only by a later new high.

The same decline is shown as a 36.3 percent loss if the futures margin account started at 10000 instead of at the prior equity high, so funding level is part of the evaluation.

One robustness recipe freezes the original build parameters and withholds the first 15 percent of the series so the unused remainder can serve as an out-of-sample stress test.

Process-fitness and override

Process-fitness is whether the operator can take every signal without override through a sub-50-percent win rate, long losing streaks and multi-week drawdowns. The evaluation includes whether a near-45-percent win rate and as many as nine consecutive losses would cause the operator to abandon or rewrite signals.

If the operator cannot take every signal that conflicts with prior market beliefs, the evaluation treats delegation to a third party who will execute without override as a process choice.

Hours versus months for a similar conclusion

An automated walk-forward optimizer is described as reaching a similar viability conclusion in hours that a manual, session-by-session log took more than seven months to produce.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
39 of 51 in the Robustness testing track
201338-41 pp.Next on Robustness testingTwo-window evaluation of mechanical trading systemsA mechanical trading procedure is described as having both strengths and weaknesses, and no single system is presented as suitable for every trader or every market state.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings