A research ledger that catches a moved goalpost
Before opening a result, save the candidate rules, evaluation dates, modeled costs and the requirement you will use to reject it. The example below is an illustrative ledger entry, not an observed backtest or a recommended drawdown limit.
The declared limit is 20% drawdown. A hypothetical 24% result fails. Editing the limit to 25% after seeing that result cannot make version 1 pass; it creates version 2 and requires a later unseen assessment. Run the example to see both decisions, then replace the illustration with your saved result and predeclared protocol.
protocol = {"version": "v1", "maximum_drawdown_percent": 20.0}
illustrative_result = {"maximum_drawdown_percent": 24.0}
passes_original = illustrative_result["maximum_drawdown_percent"] <= protocol["maximum_drawdown_percent"]
revised_protocol = {"version": "v2", "maximum_drawdown_percent": 25.0}
print({"passes_original": passes_original, "revised_protocol_needs_new_holdout": revised_protocol != protocol})Write the rejection rule before seeing the result
A research protocol limits the choices a promising chart can tempt you to change. Start with a single hypothesis, such as whether a GOOG trend filter reduces drawdown relative to an equal-capital baseline after the same modeled costs. Define the rule and the comparison before running it.
Keep capital, interval, fees, share rounding, dividends and calendar constant across candidates. Freeze how you select a validation winner and what counts as insufficient activity. A result that misses a floor stays a miss even if the system returns a best-available candidate.
A proposed protocol you can fill in
This is a proposed protocol, not a claim these settings have passed. Choose numeric floors appropriate to your intended use before observing outcomes.
| Hypothesis | GOOG trend filter reduces drawdown against equal-cost baseline | Before initial runs |
| Candidate registry | SMA short window 20/50/100; long window 200 | Before search |
| Search budget | Three declared candidates; all manual revisions logged | Before search |
| Selection rule | Validation objective plus declared activity/drawdown floors | Before validation |
| Final holdout | Dates outside every fold and every optimization | Before search |
| Failure response | Reject, or start a new version with a new holdout | Before final test |
No matching rows. Clear the filter to see all records.
Separate development, selection and final assessment
Use training to build candidates and validation to make choices. Walk-forward outer results reveal how those choices behaved later in each fold. If you inspect outer results to choose the final assembled book, that choice still needs a separate final holdout.
The Public Portfolio Challenge runbook freezes gates and keeps a final lockbox outside folds and search. Its concrete discipline is transferable: record the cutoff, preserve the assembled book and touch the lockbox once after design freeze. The runbook is an example process, not a claim this GOOG demonstration followed that live campaign.
The first fold's information boundary
| Training | 2016-01-01T05:00:00Z to 2021-10-03T03:59:59.999Z | Develop or fit candidates |
| Embargo | 2021-10-03T04:00:00Z to 2021-10-17T03:59:59.999Z | Excluded gap between training and validation |
| Validation | 2021-10-17T04:00:00Z to 2023-03-30T03:59:59.999Z | Choose a candidate before its outer test |
| Outer test | 2023-03-30T04:00:00Z to 2023-12-07T04:59:59.999Z | Evaluate the frozen candidate |
No matching rows. Clear the filter to see all records.
Dates come from the actual plan. UTC boundaries reflect the generated market calendar; preserve them instead of hand-rounding an adjacent day into both windows.
Keep an experiment manifest
Copy this into a versioned research note. Fill every bracket from your own plan; a manifest with missing gates cannot retrospectively establish that the gates were frozen.
version: trend-goog-v1
hypothesis: [written before the run]
candidates: [all declared rule variants]
fixed: capital, dates, fees, interval, shares, dividends
selection: [objective, constraints, tie-breakers]
trial_budget: [declared count; include manual changes]
final_holdout: [start/end; excluded from all prior runs]
failure_policy: [reject or new version]
artifacts: [exact configs, results, rejected trials]Use failed windows to end a test, not move its finish line
If a fold misses the requirement, show it beside the aggregate. If the final holdout fails, do not call a retuned rerun a fresh test. A new rule needs a new experiment version and a later unseen assessment.
Repeatability means someone can reconstruct the same rules, calendar and costs. It does not mean that repeated simulation calls are independent new evidence. Preserve raw results and compare changes against the original fixed baseline.