Strategy research methods

Trading strategy overfitting: diagnose the evidence behind a curve

Distinguish a weak backtest from evidence of overfitting, inspect lost folds and count the searches that produced the winning strategy.

As of 2026-10-10

Check the evidence before copying a high-ranked result

Open the existing backtest explorer and inspect a completed run's configuration, dates, capital and costs. Then ask how its rules were selected. A high score in the archive is already observed evidence; using it to choose your next rule makes the archived dates development data for that choice.

Save the exact entry, exit, sizing and indicator settings before choosing later assessment dates. Preserve the archive result as a source, not a newly unseen test. If trial history or selected-winner metadata is absent, mark it unknown rather than inferring a one-trial experiment.

Audit the population behind the published corpus

The available archive contains 222,010 recorded runs, with its newest record dated 2026-09-06. The October reanalysis selects a daily stock cohort without options, excludes optimizer-labeled runs and keeps finite-Sharpe records lasting 365 through 2,000 days.

The selected records do not cover every attempted strategy, and the archive does not report how many jobs failed. Author origin is unknown. The archive does not identify which optimizer candidates were chosen before a later test, so it cannot establish optimizer superiority or clean unseen performance.

Available source runs222010The retained archive, not the live database
Optimizer-labeled source runs74518Recorded label across the source; excluded from this ranking cohort
Eligible daily stock runs102781After asset, cadence, duration, origin-label and metric filters
Incomplete or unsupported records excluded12233Typed normalization rejects missing or unsupported execution inputs
Repeated execution records collapsed47698One representative per normalized rule/action and recorded execution contract
Distinct available recorded execution contracts42850Normalized typed rules/actions and recorded execution inputs
Zero-return contracts retained3013Zero return is not proof of zero fills

See how choosing the start date changes the story

The same five-dividend-grower rules beat the stored SPY return in the 2025 window and lagged it in the longer 2014 window. Publishing only the recent curve would hide that difference. These inspected periods are historical examples, not an untouched final test or proof of overfitting.

Each run starts with $10,000 and reinvests dividends. The strategy pays 0.1% stock transaction fees; the stored SPY baseline pays no candidate commissions. Historical point-in-time SP500 membership is unverified. Different assets, costs and exposure limit the comparison.

2025-01-0244.80%33.99%24.33%49$24.15
2014-01-02340.90%419.23%24.16%422$331.98

Ask how the winning rules were chosen

Overfitting occurs when a strategy captures patterns particular to the development sample instead of a relationship that carries to other data. A strong backtest does not tell you how much searching preceded it. Two identical curves can represent very different evidence if one was the only predeclared rule and the other survived thousands of revisions.

Write down what changed: indicator periods, assets, dates, entry families, exits, allocation and risk limits. Include manual changes and discarded prompts. Searching the calendar or deleting an inconvenient asset also selects on results.

Check competing explanations

These are diagnostic possibilities. None follows from the headline return alone.

Training rises while validation worsensThe search fits sample-specific noiseAll generations, frozen validation and candidate ledger
Few trades drive a high scoreToo little independent evidenceClosed trades, holding duration, exposure and full path
Every close-based fill occurs at that same closeInformation timing errorSignal availability timestamp and eligible fill timestamp
Only the best names remain in the universeSelection or survivorship biasEligible-as-of universe and rejected names

Audit a claimed winner

Use this request with the actual stored experiment artifacts. If the artifacts omit failed trials, label the trial count unknown instead of treating it as one.

Prompt
Show the full trial count and rejected candidates.
List every use of the final test dates before design freeze.
Report training, validation and outer-test statistics separately.
Show all test folds, modeled fees and capital assumptions.
Mark windows that overlap and warnings that relax selection floors.

Respond to failure without polishing the story

If later performance fails your frozen requirement, reject or redesign the hypothesis. Keep the failed version in the record. A simpler rule can be easier to explain and reproduce, but fewer parameters alone do not guarantee it generalizes.

A fair report includes the original question, how the candidate was selected, the loss windows and what remains untested. A historical fixed-book example cannot promise a return for the reader or establish a live strategy.

Continue exploring