Skip to main content
Track Robustness testing
15 / 51
Library

1995issue C121-9

Input pruning as walk-forward system evaluation

Shrinking the input set is the main system-optimization lever when extra capacity fits the training window and fails on unseen data. Each pruned bundle should be walked forward and kept only if robustness screens still agree on the survivors.

  • Too many inputs, hidden nodes, or connection weights can fit the training window and fail on unseen data, so shrinking the input set is the main system-optimization lever.
  • Walk-forward analysis should judge each pruned bundle on unseen observations and use holdouts longer than the forecast horizon, because later weeks can decay toward chance.
  • Robustness testing asks whether surviving inputs persist after collinearity transforms, alternate deletion screens, and more than one training run.
  • Linear correlation with the target is a weak prune, and a linear T-stat ranking is not interchangeable with a connection-weight ranking after the same 12-of-24 cut.
Entries in this reading3 entries

Shrink the input set as one procedure

Too many inputs, hidden nodes, or connection weights can fit the training window while producing unreliable forecasts on unseen data. Shrinking the input set is the main system-optimization lever.

An interior capacity optimum

In a one-input, one-output comparison, three hidden nodes generalized better on unseen data than one node or nine nodes. The comparison implies an interior capacity optimum rather than a bigger-is-better rule.

How the protocol trained and stopped

The evaluation protocol trained a 24-input model of a nine-week percent change on 240 weekly records and scored it on a 60-week random holdout drawn from 300 records, without checking that the two samples had similar distributions.

Training inspected test-set mean square error every 100 learning events and stopped after 10,000 events with no new test-set minimum, then used that lowest-error network for later deletion trials.

Deletion screens and collinear pairs

Pruning by linear correlation with the target is a weak system screen because a nonlinear model can still extract usable information from inputs that look almost unrelated to the output.

Connection-weight ranking, replacing an input with its average, and leave-one-input-out search are alternative optimization loops: delete a candidate, retrain, and continue until one input remains.

When input pairs are highly correlated, combining the pair is preferred to deleting one member. The evaluation treated absolute correlation above 0.80 as the threshold for that robustness fix.

Screens that do not agree

After the same 12-of-24 cut, a linear T-stat ranking and a connection-weight ranking kept only seven inputs in common, so the two screens are not interchangeable robustness tests.

A cheaper composite loop is offered as a practical substitute for exhaustive systematic search: keep the strongest weights, transform collinear pairs, batch-delete weak weights, then run sensitivity analysis.

Holdouts, retraining, and the trading objective

Walk-forward blocks should be longer than the forecast horizon because later weeks can decay toward chance. A single training run before each deletion can make the search path unstable. Lowest test-set error need not rank the same systems as a trading objective.

Holdout MSE after each input-pruning rule

Treat each pruning rule as a different trading procedure: the same 24 weekly transforms, trained to forecast the nine-week percent change in S&P 500 cash and scored on a 60-week holdout, do not keep the same survivors. The comprehensive blend posts the lowest test-set MSE (0.000800) with 11 inputs; correlation is the worst holdout (0.001070) while still keeping 16. Training error is lower for every rule, so extra capacity can look fine in-sample and fail unseen. Values are the Testing minimum MSE and Training MSE columns of the article summary table.
Treat each pruning rule as a different trading procedure: the same 24 weekly transforms, trained to forecast the nine-week percent change in S&P 500 cash and scored on a 60-week holdout, do not keep the same survivors. The comprehensive blend posts the lowest test-set MSE (0.000800) with 11 inputs; correlation is the worst holdout (0.001070) while still keeping 16. Training error is lower for every rule, so extra capacity can look fine in-sample and fail unseen. Values are the Testing minimum MSE and Training MSE columns of the article summary table.S&P 500 cash · weekly

One training run preceded each deletion. Systematic testing was stopped at 21 inputs; the author said more trials would likely have cut error further. Mean-square error was the stop rule, not trading profit. The table lists 16 correlation inputs (the prose says 17) and a T-statistic test MSE of 0.001020 (the prose quotes 0.00094).

Educational research material, not investment advice. Historical source context does not establish present-day performance.
15 of 51 in the Robustness testing track
19951-5 pp.Next on Robustness testingCritiquing neural nets as incomplete trading systemsTry many input and modeling approaches rather than hunt for one privileged neural-network recipe.
All readings on this track · 51 readings
  1. 1986Degrees of freedom in trading system optimization
  2. 1988Walk-forward and neighborhood tests after optimization
  3. 1988Undisclosed rules block system robustness tests
  4. 1988Testing re-optimization calendars against random parameter controls
  5. 1989Binary search limits on multi-peak average grids
  6. 1989Parameter neighborhoods that survive a shift
  7. 1990Use profit mapping to keep a cycle and stop plateau
  8. 1990Why popular indicator optimization fails robustness
  9. 1991Retesting weighted indicator balances across horizons
  10. 1992Constructing forecast models with regression, walk-forward, and robustness
  11. 1992Diagnose regimes before you lock parameters
  12. 1992When stops change system timing
  13. 1993Walk-forward halt rules for forecast models
  14. 1994Walk-forward evaluation of genetic index rules
  15. 1995Input pruning as walk-forward system evaluation
  16. 1995Critiquing neural nets as incomplete trading systems
  17. 1996Rebuild the equity-path ratio before it ranks a designed system
  18. 1996Parameter grids can fit random walks
  19. 1996Walk-forward analysis belongs in the design of a mechanical trading system
  20. 1997When a holdout fails, discard the rule set
  21. 1997Test rewarded rule breaks before replacing the system
  22. 1997Walk-forward rules keep system research from rewriting live trades
  23. 1999Keep a channel-breakout to two lookbacks and test neighbor stability
  24. 1999Constant investment size in stock system evaluation
  25. 2000Forcing optimization maps mechanical system failure boundaries
  26. 2000Robust parameter selection with surface charts
  27. 2001A two-gate classroom test for a two-window momentum trend filter
  28. 2002How a two-sided continuation factor becomes a testable trend rule
  29. 2002Evaluating two-window trend intensity as a reversal rule
  30. 2003Discounting speculative bubbles in system robustness tests
  31. 2003Walk-forward evaluation of locked stochastic oscillator rules
  32. 2003Critiquing mechanical system design after extreme price regimes
  33. 2004Evaluating a two-window trend trigger
  34. 2005Grade backtested signals with holdouts and optimization plateaus
  35. 2006Reserved-sample evaluation of trading system design
  36. 2006Walk-forward critique of hindsight crossover systems
  37. 2008Condition-matched walk-forward evaluation for mechanical systems
  38. 2011Session-split evaluation of regular and overnight systems
  39. 2012Walk-forward evaluation as operator rehearsal
  40. 2013Two-window evaluation of mechanical trading systems
  41. 2013Walk-forward filter selection for repeated-median velocity
  42. 2014Walk-forward evaluation for fading-memory velocity systems
  43. 2015Test oscillator events before tuning rules
  44. 2016Walk-forward evaluation of a five-parameter parabolic stop-and-reversal
  45. 2016Walk-forward optimization without curve fitting
  46. 2017Optimization without overfitting in trend-system evaluation
  47. 2017Parameter stability is a better guide than a larger crossover grid
  48. 2018Point-in-time universes for system evaluation
  49. 2018Walk-forward robustness evaluation for optimized systems
  50. 2018Critiquing breakout systems through robustness tests
  51. 2018A critique of parameter fitting in system design
All 58 readings tagged Robustness testing
Also on Robustness testing5 readings