Strategy research methods

Prevent strategy overfitting with a frozen research protocol

Build a research ledger before searching: freeze rules, trial budget, dates, baselines and a final holdout, then record what happens when a gate fails.

As of 2026-10-10

A research ledger that catches a moved goalpost

Before opening a result, save the candidate rules, evaluation dates, modeled costs and the requirement you will use to reject it. The example below is an illustrative ledger entry, not an observed backtest or a recommended drawdown limit.

The declared limit is 20% drawdown. A hypothetical 24% result fails. Editing the limit to 25% after seeing that result cannot make version 1 pass; it creates version 2 and requires a later unseen assessment. Run the example to see both decisions, then replace the illustration with your saved result and predeclared protocol.

Python
protocol = {"version": "v1", "maximum_drawdown_percent": 20.0}
illustrative_result = {"maximum_drawdown_percent": 24.0}
passes_original = illustrative_result["maximum_drawdown_percent"] <= protocol["maximum_drawdown_percent"]
revised_protocol = {"version": "v2", "maximum_drawdown_percent": 25.0}
print({"passes_original": passes_original, "revised_protocol_needs_new_holdout": revised_protocol != protocol})

Write the rejection rule before seeing the result

A research protocol limits the choices a promising chart can tempt you to change. Start with a single hypothesis, such as whether a GOOG trend filter reduces drawdown relative to an equal-capital baseline after the same modeled costs. Define the rule and the comparison before running it.

Keep capital, interval, fees, share rounding, dividends and calendar constant across candidates. Freeze how you select a validation winner and what counts as insufficient activity. A result that misses a floor stays a miss even if the system returns a best-available candidate.

A proposed protocol you can fill in

This is a proposed protocol, not a claim these settings have passed. Choose numeric floors appropriate to your intended use before observing outcomes.

HypothesisGOOG trend filter reduces drawdown against equal-cost baselineBefore initial runs
Candidate registrySMA short window 20/50/100; long window 200Before search
Search budgetThree declared candidates; all manual revisions loggedBefore search
Selection ruleValidation objective plus declared activity/drawdown floorsBefore validation
Final holdoutDates outside every fold and every optimizationBefore search
Failure responseReject, or start a new version with a new holdoutBefore final test

Separate development, selection and final assessment

Use training to build candidates and validation to make choices. Walk-forward outer results reveal how those choices behaved later in each fold. If you inspect outer results to choose the final assembled book, that choice still needs a separate final holdout.

The Public Portfolio Challenge runbook freezes gates and keeps a final lockbox outside folds and search. Its concrete discipline is transferable: record the cutoff, preserve the assembled book and touch the lockbox once after design freeze. The runbook is an example process, not a claim this GOOG demonstration followed that live campaign.

The first fold's information boundary

Training2016-01-01T05:00:00Z to 2021-10-03T03:59:59.999ZDevelop or fit candidates
Embargo2021-10-03T04:00:00Z to 2021-10-17T03:59:59.999ZExcluded gap between training and validation
Validation2021-10-17T04:00:00Z to 2023-03-30T03:59:59.999ZChoose a candidate before its outer test
Outer test2023-03-30T04:00:00Z to 2023-12-07T04:59:59.999ZEvaluate the frozen candidate

Dates come from the actual plan. UTC boundaries reflect the generated market calendar; preserve them instead of hand-rounding an adjacent day into both windows.

Keep an experiment manifest

Copy this into a versioned research note. Fill every bracket from your own plan; a manifest with missing gates cannot retrospectively establish that the gates were frozen.

version: trend-goog-v1
hypothesis: [written before the run]
candidates: [all declared rule variants]
fixed: capital, dates, fees, interval, shares, dividends
selection: [objective, constraints, tie-breakers]
trial_budget: [declared count; include manual changes]
final_holdout: [start/end; excluded from all prior runs]
failure_policy: [reject or new version]
artifacts: [exact configs, results, rejected trials]

Use failed windows to end a test, not move its finish line

If a fold misses the requirement, show it beside the aggregate. If the final holdout fails, do not call a retuned rerun a fresh test. A new rule needs a new experiment version and a later unseen assessment.

Repeatability means someone can reconstruct the same rules, calendar and costs. It does not mean that repeated simulation calls are independent new evidence. Preserve raw results and compare changes against the original fixed baseline.

Continue exploring