Skip to main content
Track Robustness testing
47 / 51
Library

2017issue C0946-47

Parameter stability is a better guide than a larger crossover grid

Across the windows labeled 1985, 2000, and 2008, the average of all moving-average-crossover tests is described as similar to the average of all single-trend tests, so the remaining decision is which trade-profile to accept. Editorial reading: keep the design whose already-reasonable parameter-neighborhood still works in every accepted window, and treat a slowly changing crossover surface as a warning.

  • When the average of all moving-average-crossover tests looks similar to the average of all single-trend tests, the remaining decision is which trade-profile to accept.
  • Robustness-testing asks whether three already-reasonable parameters stay usable across every evaluation window you accept as valid.
  • Selecting three parameters is easier from a 17-test single-trend test-grid than from a 119-test crossover grid, and adding more parameters complicates the solution.
  • Averaging the full test-grid sets realistic expectations. Picking a short-term lookback of 30 or 35 days after scanning the larger crossover grid is questioned as a possible overfitting step.
Entries in this reading3 entries

Similar averages leave a trade-profile choice

Across three historical windows labeled 1985, 2000, and 2008, the average of all moving-average-crossover tests is described as similar to the average of all single-trend tests. That leaves the remaining decision as which trade-profile to accept.

The single-trend profile is characterized as holding positions longer. The moving-average-crossover profile is characterized as generating more trades with smaller typical gains and smaller typical losses.

Profit-taking is described as more compatible with the higher-frequency crossover profile than with the longer-hold single-trend profile. Applying it is said to change the system profile.

Average one-trend vs crossover results by test window

In each of the 1985, 2000, and 2008 windows the average of every moving-average-crossover test sits close to the average of every single-trend test, so the remaining choice is holding time versus a busier trade profile rather than a large P&L gap. The figures are the three-column comparison printed as Figure 5.
In each of the 1985, 2000, and 2008 windows the average of every moving-average-crossover test sits close to the average of every single-trend test, so the remaining choice is holding time versus a busier trade profile rather than a large P&L gap. The figures are the three-column comparison printed as Figure 5.Test windows labeled 1985, 2000, and 2008 · 1985-01-01T00:00:00.000Z to 2008-12-31T00:00:00.000Z

Each cell is the average of all tests in that window, not the single best parameter set. The author notes losses in the 2008 heat map and stays with the single-trend design because a small already-reasonable neighborhood still worked in every window.

Read the full test-grid, not the best cell

System-optimization is presented as a visual evaluation tool for judging whether results are robust or erratic, comparing averages across systems and periods, and detecting whether a design is holding up or degrading.

Averaging all tests is presented as a way to form realistic expectations, in contrast to selecting the single best combination from a large battery of tests.

Three parameters should stay usable in every window

A robustness-testing check offered in the material is whether three parameters remain usable across all evaluation windows. The single-trend design is described as meeting that test, while the crossover design is described as slowly changing.

Selecting three parameters is described as easier from a 17-test single-trend test-grid than from a 119-test crossover test-grid. Adding more parameters is framed as complicating the solution and inviting overfitting.

Choosing a short-term lookback of 30 or 35 days after inspecting the larger crossover grid is explicitly questioned as a possible overfitting step.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
47 of 51 in the Robustness testing track
20188-11 pp.Next on Robustness testingPoint-in-time universes for system evaluationThe first run applied the demonstration rules to the Nasdaq 100 membership that existed at test time, not to the membership that existed on each historical signal date.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings