Skip to main content
Track Coin-flip no-skill baseline
4 / 6
Library

2005issue C031-4

Evaluating systems with walk-forward analysis, robustness testing, and coin-flip baselines

A whole-sample test can look acceptable in aggregate while almost all of the positive contribution sits in a late minority of the bars. Reserved-sample confirmation and ranking against occupancy-matched coin-flip peers are both required before a rule set is treated as more than an in-sample artifact.

  • A whole-sample test can look acceptable in aggregate while almost all of the positive contribution sits in a late minority of the bars, and a full-sample grid search is not a robustness test.
  • Walk-forward analysis judges entry, exit, and abstention rules on confirmation slices that did not choose those rules; recycling a used confirmation slice collapses the procedure back into whole-sample testing.
  • Robustness testing re-runs the same rules on time slices or return-reshuffled series; a grid-preferred parameter set can fail on 100 series that share the sample's statistical character rather than its exact price path.
  • Encoding the occupancy signature, such as 23 percent long, 48 percent short, and 29 percent flat, lets a coin-flip no-skill baseline rank the original signals against occupancy-matched chance.
Entries in this reading3 entries

A whole-sample test can hide the path

A whole-sample test is a single contiguous-history evaluation that reports one aggregate result and cannot show how outcomes are distributed through time. A whole-sample test of a two-average crossover on a long contiguous index history can look acceptable in aggregate while the same rules, scored on ten equal-sized portions, show that almost all of the positive contribution sits in a late minority of the bars.

Searching every daily-bar crossover pair with lengths from three through twelve on the same full sample can reverse the ranking of nearby parameter sets, so a grid search by itself is not a robustness test.

9/12 MA crossover points by DJIA sample slice

The full-sample 9/12 daily crossover looks like a roughly 2,500-point DJIA gain, but the ten equal bar slices show almost all of that profit arriving in I and J. Slices A–E are only modestly positive and F–H lose several hundred points each, so the whole-sample result hides a late-sample concentration. Values are the exact Points Returned entries from the article’s Figure 2 table.
The full-sample 9/12 daily crossover looks like a roughly 2,500-point DJIA gain, but the ten equal bar slices show almost all of that profit arriving in I and J. Slices A–E are only modestly positive and F–H lose several hundred points each, so the whole-sample result hides a late-sample concentration. Values are the exact Points Returned entries from the article’s Figure 2 table.DJIA · Daily bars

Same 9/12 daily crossover on 80 years of DJIA bars split into ten contiguous slices; the last slice is slightly shorter (18,000–19,765) because the sample ends at bar 19,765.

Walk-forward analysis separates development from confirmation

Walk-forward analysis is a split of historical bars into development and confirmation slices so entry, exit, and abstention rules are judged on data that did not choose those rules. Once a confirmation slice has been used, recycling it for further tuning collapses the procedure back into whole-sample testing.

Even when remaining history is large enough for many development iterations, confirming a rule set only on a reserved subset can leave the procedure untested against a full range of market behaviors. This is sample selectivity: the risk that a reserved confirmation subset omits market behaviors the rules will later face, even when the remaining history is large.

Robustness testing checks concentration and path dependence

Robustness testing re-runs the same rule set on time slices or return-reshuffled series to test whether an apparent result is concentrated, unstable, or an artifact of one path.

A low-friction robustness procedure converts the original series to percent changes, randomly reorders those changes, and rebuilds price paths so the same rules can be retested on many statistically related series without repeating the identical path. A parameter set preferred by a contiguous-history grid can fail when the same rules are applied to 100 return-reshuffled series that share the original sample's statistical character rather than its exact price path.

A coin-flip baseline keeps the occupancy signature

Comparisons with passive holding or short-term interest, and ratios that scale return against drawdown, do not rank systems that share a similar long, short, and flat footprint against one another.

A coin-flip no-skill baseline is a no-skill comparison that keeps a system's mix of long, short, and flat bars and randomly reassigns those states, asking whether the original signals beat occupancy-matched chance. Encoding a system as the fraction of bars spent long, short, and flat, illustrated as 23 percent, 48 percent, and 29 percent, lets an evaluator build many occupancy-matched signal sequences by randomly reordering that three-state column. That mix is the occupancy signature: the share of bars a system spends long, short, or flat, used to generate look-alike signal sequences for peer ranking.

Expanding that occupancy-matched peer test to 1,000 similar systems on real market data placed a hypothetical candidate worse than 77 percent of its coin-flip twins.

Confirmation and peer ranking both remain required

After these checks, reserved-sample confirmation and ranking against occupancy-matched peers are both required before a rule set is treated as more than an in-sample artifact.

Educational research material, not investment advice. Historical source context does not establish present-day performance.
4 of 6 in the Coin-flip no-skill baseline track
201560-64 pp.Next on Coin-flip no-skill baselineTrade-tape entropy versus a coin-flip no-skill baselineInformation entropy scores uncertainty in a sequence of symbols without using what those symbols mean, so the same formula can treat bits as coin tosses or as wins and losses.
All readings on this track · 6 readings
  1. 1986Skill score versus a coin-flip forecast baseline
  2. 1991Evaluating a trailing stop against a coin-flip entry
  3. 2004Evaluating trend rules against no-skill baselines
  4. 2005Evaluating systems with walk-forward analysis, robustness testing, and coin-flip baselines
  5. 2015Trade-tape entropy versus a coin-flip no-skill baseline
  6. 2017A coin-flip timed exit as the skill floor for trend and mean-reversion
All 6 readings tagged Coin-flip no-skill baseline
Also on Coin-flip no-skill baseline5 readings