Skip to main content
Track Robustness testing
4 / 51
Library

1988issue C071-7

Testing re-optimization calendars against random parameter controls

A historical evaluation compared history-based re-optimization calendars with a random-parameter-control on a channel breakout and a directional-movement index crossover. History-based search did not produce a statistically greater result than chance assignment. Editorial: treat the calendar as an extra trading rule, and treat walk-forward search as useful system-design work only when robustness-testing can show it beats a random draw from the same parameter-grid.

  • A re-optimization calendar jointly sets how much past data ranks the next parameter and how long that choice is held before the search may run again.
  • Walk-forward analysis selects a parameter on an earlier window and applies that choice only on a later interval that was not used to rank the candidates.
  • On both tested systems, history-based calendars did not differ statistically from a random-parameter-control, and they did not statistically impair mean results.
  • The in-sample-champion is not a valid stand-in for later results; the center of the full tested parameter-grid is the more defensible historical benchmark.
Entries in this reading3 entries

The calendar sits inside the procedure

System-optimization chooses a system's lookback or threshold from ranked historical results instead of committing to a setting before the test begins. The re-optimization calendar is the joint choice of how much past data ranks the next parameter and how long that choice is held before the search is allowed to run again.

That calendar decides when a new lookback may be selected and which already observed results are allowed to rank the candidates. It is part of the trading procedure, not a setting that can be left outside the test.

What the historical evaluation compared

A historical evaluation compared several history-based re-optimization calendars with a random-parameter-control across a one-parameter channel breakout and a directional-movement index crossover. The random-parameter-control assigns each evaluation period's setting by chance so the search procedure itself can be tested.

Practitioners at the time did not share a standard for how much history to use, how often to refresh parameters, or whether to search at all. Lookbacks ranged from a few years to all available data, and refresh intervals ranged from twice a year to once every several years.

Each system was searched over a parameter-grid spanning short to multi-month lookbacks in even steps. Both systems were simulated on a multi-market futures portfolio, with equal margin allocations, a fixed commission per trade, and positions taken only in the nearby contract.

Walk-forward choice and hold periods

Walk-forward analysis selects a parameter on an earlier window and applies that choice only on a later interval that was not used to rank the candidates. Every history-based calendar chose the next parameter only from already observed results, so the evaluation interval never reused the observations that ranked the candidates.

Annual-refresh calendars used several prior years of profit ranking, or all history from market inception or a fixed early start. Fixed calendars held a multi-year choice for a matching multi-year block.

Channel-breakout mean portfolio returns by re-optimization calendar, 1965–1985

Average out-of-sample portfolio returns for the channel-breakout system sit in a narrow 41–61 percent band across all ten history-based calendars, and the random-parameter control (RND, 54.14 percent) sits inside that band rather than below it. A trader should treat the re-optimization calendar as just another rule: on this test it does not lift the mean above a chance draw from the same grid. Figures are the authors’ reported mean percent returns from their Channel Breakout summary table.
Average out-of-sample portfolio returns for the channel-breakout system sit in a narrow 41–61 percent band across all ten history-based calendars, and the random-parameter control (RND, 54.14 percent) sits inside that band rather than below it. A trader should treat the re-optimization calendar as just another rule: on this test it does not lift the mean above a chance draw from the same grid. Figures are the authors’ reported mean percent returns from their Channel Breakout summary table.Channel breakout, 15-futures portfolio · 1965–1985 · 1965-01-01T00:00:00.000Z to 1985-12-31T00:00:00.000Z

Means assume 30 percent of capital is posted to initial margins. The search grid was 5–60 days in steps of five; commission $100 per trade; every calendar is evaluated out of sample on a 15-market portfolio.

Paired tests including the random control

Robustness-testing checks whether alternative search calendars, including a chance-assignment control, produce statistically distinguishable later results. Statistical comparison used tests of mean portfolio results against zero and paired-difference tests among calendars, including the random-parameter-control.

On the breakout system, paired-difference tests found no statistically significant gap among mean monthly results of the calendars, including the gap versus random assignment.

On the directional-movement system, only a few pairwise contrasts reached significance, no calendar formed a consistent ranking, and none differed statistically from random assignment.

Within the tested systems, parameter-grids, and calendars, history-based re-optimization did not produce a statistically greater result than chance assignment and did not statistically impair mean results.

What the in-sample champion cannot replace

The in-sample-champion is the single historically best setting on a window. Its own window result is not a valid estimate of later results, so that champion is not a valid stand-in for what follows. The center of the full tested parameter-grid is the more defensible historical benchmark.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
4 of 51 in the Robustness testing track
19891-3 pp.Next on Robustness testingBinary search limits on multi-peak average gridsA worked example tuned only two moving-average inputs, one exponential-average decimal for buying and one for selling, and applied the same historical database to every pair.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings