Skip to main content
Track Robustness testing
46 / 51
Library

2017issue C0918-21

Optimization without overfitting in trend-system evaluation

Overfitting occurs when parameters are tuned so tightly to historical patterns that the same rules fail on later, different data. Evaluation is stronger when ranges are chosen in advance, more history is used, and robustness is judged by how widely tests remain profitable rather than by a single peak.

  • Overfitting occurs when parameters are tuned so tightly to historical patterns that the same rules fail on later, different data.
  • Choose parameter ranges in advance from the expected holding style. Scanning every possible value and then keeping only the profitable ones is a first step toward overfitting.
  • Prefer more history so evaluation includes bull and bear regimes, price shocks, and a larger trade count, and retest the same range on three successive windows.
  • Treat a result as robust when a large share of tests remain profitable, and treat a new rule as generalized only when the average of all tests improves and most individual tests improve.
Entries in this reading3 entries

Overfitting in system optimization

Overfitting occurs when parameters are tuned so tightly to historical patterns that the same rules fail on later, different data.

Scanning every possible value and then keeping only the profitable ones is described as a first step toward overfitting.

Use more history, not less

Using more history is preferred because it adds bull and bear regimes, price shocks, and a larger trade count for evaluation.

Commit the range before the scan

Parameter ranges should be chosen in advance from the expected holding style, such as 40 to 120 days for macrotrend work and 5 to 40 days for short-term work.

Trend rules used for the tests

A one-parameter trend test can use the slope of a moving average to stay long when the line is rising and short when it is declining, avoiding many false signals from price-versus-line crossings.

A Moving-average crossover is long when the faster average is above the slower one and short when the faster average is below.

Retest the same range on later windows

Testing the same range over three successive windows, each about half as long as the previous, is used to check whether the edge still appears in later history.

A robust result is defined as a large share of tests remaining profitable, not as a single peak profit, with a trend example of as much as 70% successful tests.

When a new rule is treated as generalized

A new rule is treated as generalized only if the average of all tests improves and most individual tests improve. Improvement confined to one pocket while others worsen is treated as overfitting.

A volatility filter on new entries

A volatility filter that blocks new entries when annualized price volatility exceeds 50% and waits until it falls back below that level is presented as typically cutting return a little while cutting risk much more.

Copper crossover average profit by slow moving-average length

Row averages from the copper moving-average crossover heatmaps. The 1985 window stays profitable from 40 through 120 days and peaks near an 80-day slow average. Later starts sit lower, the best slow period shortens toward 65 days, and from 2008 the long end of the same range turns into losses.
Row averages from the copper moving-average crossover heatmaps. The 1985 window stays profitable from 40 through 120 days and peaks near an 80-day slow average. Later starts sit lower, the best slow period shortens toward 65 days, and from 2008 the long end of the same range turns into losses.Copper futures · Daily · 1985-01-01T00:00:00.000Z

Each point is the published row average across fast averages of 5, 10, 15, 20, 25, 30 and 35 days. Tests use copper back-adjusted futures, long when the fast average is above the slow average and short when it is below, with futures size set to $25,000 divided by the 20-day ATR times the big-point value.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
46 of 51 in the Robustness testing track
201746-47 pp.Next on Robustness testingParameter stability is a better guide than a larger crossover gridWhen the average of all moving-average-crossover tests looks similar to the average of all single-trend tests, the remaining decision is which trade-profile to accept.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings