Check the evidence before copying a high-ranked result
Open the existing backtest explorer and inspect a completed run's configuration, dates, capital and costs. Then ask how its rules were selected. A high score in the archive is already observed evidence; using it to choose your next rule makes the archived dates development data for that choice.
Save the exact entry, exit, sizing and indicator settings before choosing later assessment dates. Preserve the archive result as a source, not a newly unseen test. If trial history or selected-winner metadata is absent, mark it unknown rather than inferring a one-trial experiment.
Audit the population behind the published corpus
The available archive contains 222,010 recorded runs, with its newest record dated 2026-09-06. The October reanalysis selects a daily stock cohort without options, excludes optimizer-labeled runs and keeps finite-Sharpe records lasting 365 through 2,000 days.
The selected records do not cover every attempted strategy, and the archive does not report how many jobs failed. Author origin is unknown. The archive does not identify which optimizer candidates were chosen before a later test, so it cannot establish optimizer superiority or clean unseen performance.
| Available source runs | 222010 | The retained archive, not the live database |
| Optimizer-labeled source runs | 74518 | Recorded label across the source; excluded from this ranking cohort |
| Eligible daily stock runs | 102781 | After asset, cadence, duration, origin-label and metric filters |
| Incomplete or unsupported records excluded | 12233 | Typed normalization rejects missing or unsupported execution inputs |
| Repeated execution records collapsed | 47698 | One representative per normalized rule/action and recorded execution contract |
| Distinct available recorded execution contracts | 42850 | Normalized typed rules/actions and recorded execution inputs |
| Zero-return contracts retained | 3013 | Zero return is not proof of zero fills |
No matching rows. Clear the filter to see all records.
See how choosing the start date changes the story
The same five-dividend-grower rules beat the stored SPY return in the 2025 window and lagged it in the longer 2014 window. Publishing only the recent curve would hide that difference. These inspected periods are historical examples, not an untouched final test or proof of overfitting.
Each run starts with $10,000 and reinvests dividends. The strategy pays 0.1% stock transaction fees; the stored SPY baseline pays no candidate commissions. Historical point-in-time SP500 membership is unverified. Different assets, costs and exposure limit the comparison.
| 2025-01-02 | 44.80% | 33.99% | 24.33% | 49 | $24.15 |
| 2014-01-02 | 340.90% | 419.23% | 24.16% | 422 | $331.98 |
No matching rows. Clear the filter to see all records.
Ask how the winning rules were chosen
Overfitting occurs when a strategy captures patterns particular to the development sample instead of a relationship that carries to other data. A strong backtest does not tell you how much searching preceded it. Two identical curves can represent very different evidence if one was the only predeclared rule and the other survived thousands of revisions.
Write down what changed: indicator periods, assets, dates, entry families, exits, allocation and risk limits. Include manual changes and discarded prompts. Searching the calendar or deleting an inconvenient asset also selects on results.
Check competing explanations
These are diagnostic possibilities. None follows from the headline return alone.
| Training rises while validation worsens | The search fits sample-specific noise | All generations, frozen validation and candidate ledger |
| Few trades drive a high score | Too little independent evidence | Closed trades, holding duration, exposure and full path |
| Every close-based fill occurs at that same close | Information timing error | Signal availability timestamp and eligible fill timestamp |
| Only the best names remain in the universe | Selection or survivorship bias | Eligible-as-of universe and rejected names |
No matching rows. Clear the filter to see all records.
Audit a claimed winner
Use this request with the actual stored experiment artifacts. If the artifacts omit failed trials, label the trial count unknown instead of treating it as one.
Show the full trial count and rejected candidates.
List every use of the final test dates before design freeze.
Report training, validation and outer-test statistics separately.
Show all test folds, modeled fees and capital assumptions.
Mark windows that overlap and warnings that relax selection floors.Respond to failure without polishing the story
If later performance fails your frozen requirement, reject or redesign the hypothesis. Keep the failed version in the record. A simpler rule can be easier to explain and reproduce, but fewer parameters alone do not guarantee it generalizes.
A fair report includes the original question, how the candidate was selected, the loss windows and what remains untested. A historical fixed-book example cannot promise a return for the reader or establish a live strategy.