2004issue C091-4
Evaluating trend rules against no-skill baselines
A published trend forecast can be audited in three public steps. Freeze the sampling interval and lookback, write a coin-flip or cash-neutral dummy, then name the hypothesis test or Monte Carlo reshuffle that would have to reject that dummy before the rule is treated as skill.
- Freeze the sampling interval and lookback before scoring a trend rule. Intraday tests were judged weaker than daily tests because a printed quote does not establish executable size or price, and small rate moves then dominate whether a rule appears successful.
- Write the no-skill dummy in public. A borrow-one-currency, hold-another design was given a near-zero cash-neutral comparison, while an equity rule that spends time in a riskless asset needs a different comparison than buy-and-hold.
- A coin-flip no-skill baseline is the first hurdle. A historical outcome is not treated as skill until a hypothesis test or Monte Carlo reshuffle can reject that dummy over the same interval and lookback.
- After more elaborate generated indicators, a short-long double moving-average pair was treated as the default comparison when the currency and sample are unknown. The preferred question is why the same evaluation design is informative in some markets and not others.
Mechanical rules as a search, then a test
Interest in mechanical price rules was described as following an early-1990s seminar. The later form was a selection-and-recombination search that grows candidate rules rather than writing them by hand.
Long-sample moving-average exercises on dollar exchange rates were framed as a hypothesis-testing problem: whether conventional risk adjustments and likely trading costs could account for the observed outcome.
Freeze the sampling interval and lookback
Intraday rule tests were judged weaker than daily tests because a printed quote does not establish executable size or price, and small rate moves then dominate whether a rule appears successful.
After examining more elaborate generated indicators, a short-long double moving-average pair, illustrated as a 5-day window with a 150-day window, was treated as the default comparison model when the currency and sample are unknown. A double moving-average is a two-window trend forecast that compares a short lookback mean of ordered prices with a long lookback mean.
Editorial: the first gate is to lock that interval and those lookbacks before any score is computed. Changing the clock after the rule is chosen is not an evaluation.
Write the dummy in public
A coin-flip no-skill baseline is a dummy forecast with no claimed informational edge. It is the first hurdle a candidate rule must clear before a historical outcome is treated as skill.
A borrow-one-currency, hold-another design was given a near-zero no-skill benchmark. That comparison is a zero-return currency benchmark: a cash-neutral comparison that treats a funded long-short currency position as having an expected result near zero before costs and risk adjustments.
An equity rule that spends time in a riskless asset was said to need a different comparison than buy-and-hold because its market exposure is not constant. Editorial: the dummy has to match the exposure the rule actually takes, and it has to be written where a reader can see it.
Name the test that would reject the dummy
Hypothesis testing is a formal comparison of an explicit quantitative baseline with a held-out or otherwise out-of-sample result on ordered market observations.
A Monte Carlo simulation is a resampling check that asks how often shuffled or otherwise no-skill paths of ordered price, volume, or breadth observations would match a candidate forecast over a stated sampling interval and lookback.
Editorial: the third gate is to say which of those checks would be required to reject the published dummy. Until that rejection is specified, the indicator has not been allowed to look skillful.
What a measured result is not
Overlap between official intervention dates and rule outcomes was distinguished from a causal account. Most of the measured result was placed in regional sessions before the relevant intervention.
Implied volatility was evaluated as a forecast of realized volatility and was described as typically higher, and more variable, than the realized series. That implied-versus-realized-volatility design scores bias, scale, and average level against the later realized series. The pattern was often linked to risks that simple delta hedges do not remove.
The preferred research question was why the same evaluation design is informative in some markets and not others, rather than ranking additional rule specifications.
All readings on this track · 6 readings
- 1986Skill score versus a coin-flip forecast baseline
- 1991Evaluating a trailing stop against a coin-flip entry
- 2004Evaluating trend rules against no-skill baselines
- 2005Evaluating systems with walk-forward analysis, robustness testing, and coin-flip baselines
- 2015Trade-tape entropy versus a coin-flip no-skill baseline
- 2017A coin-flip timed exit as the skill floor for trend and mean-reversion