This appendix describes the research process in general terms. Concrete strategy settings, timeranges, parameter grids, thresholds, and execution assumptions are documented separately in Appendix: Configuration. Score formulas are documented in Appendix: Scores.
The process starts by testing manually selected reference parameter sets over the broad diagnostic period. This establishes a sanity-check baseline before interpreting any optimization output.
The baseline step records equity curves, benchmark comparison, trade lists, and standard return/risk metrics.
Purpose: confirm that the strategy logic behaves sensibly and create reference candidates that remain visible throughout the research.
The full parameter grid is tested on the research period. Instead of selecting only one best candidate, the process selects metric winners: one candidate per target metric or target metric blend. This creates a diverse research set containing profit-focused, risk-focused, stability-focused, and balanced candidates.
The same winner search is also repeated after optional quality filters. Unfiltered winners preserve extreme specialists; filtered winners surface cleaner candidates that satisfy minimum quality gates.
Purpose: build a broad candidate universe without forcing all strategy quality into one score too early.
For each metric-winner candidate, trade-level adverse and favorable excursions are inspected. Candidate-specific stop-loss and take-profit grids are derived from the candidate's own trade behavior.
Several exit modes are tested, such as keeping the original strategy exit, replacing it with SL/TP exits, or combining both. The output is a set of SL/TP-modified candidates that improve at least one target metric on the research period.
Purpose: discover whether risk-management overlays can improve promising candidates, while treating those improvements as hypotheses that must still survive later validation.
Base candidates and SL/TP-modified candidates are tested across rolling windows and fixed out-of-sample periods. The rolling windows measure temporal consistency; the fixed OOS periods provide independent validation blocks.
The robustness score combines rolling-window behavior and fixed-OOS behavior. The shortlist for later steps is assembled from top robustness candidates, metric winners, filtered winners, static references, and robust SL/TP candidates.
Purpose: reduce the candidate universe to a manageable shortlist while penalizing candidates that only work in isolated periods.
Shortlisted candidates are evaluated on manually defined market-regime intervals. Each candidate is tested in distinct market contexts such as uptrends, downtrends, sharp pumps, sharp dumps, and sideways periods.
The report shows both per-regime behavior and an aggregated regime score.
Purpose: expose whether a candidate is broadly usable or depends on one favorable market environment.
The shortlisted candidates receive a composite score that combines robustness, regime behavior, and parameter stability. The composite score is a ranking heuristic, not a probability of future profitability.
Parameter stability is estimated by checking whether nearby parameter combinations also perform reasonably well. A broad plateau is preferred over a sharp isolated optimum.
Purpose: rank the shortlist while keeping the underlying components visible.
The final shortlist is backtested over full, yearly, and monthly diagnostic periods and compared with a buy-and-hold benchmark. The report includes equity curves, benchmark-relative metrics, trade details, and candidate-level diagnostic visualizations.
This step is diagnostic and does not change the composite score.
Purpose: inspect whether ranked candidates remain trader-readable and benchmark-aware.
Full-period trades are analyzed under leverage assumptions. The main output is a leverage-grid diagnostic: multiple leverage levels are tested for each shortlisted candidate to estimate the range that remains inside hard risk guardrails.
The hard guardrails are minimum trade sample, MAE P99 below approximate liquidation distance, leveraged drawdown below the accepted limit, and worst leveraged trade loss below the accepted limit. The report also computes a soft leverage suitability score from MAE safety, drawdown safety, tail-loss safety, and robustness.
If a reference leverage is supplied, the step also shows a fixed-point diagnostic at that leverage. The reference view is optional and should be read as a check, while the grid summary is the primary view.
This step is diagnostic and does not change the composite score.
Purpose: determine whether a candidate that looks attractive without leverage remains acceptable under leveraged risk.
Shortlisted candidates are re-tested across a grid of execution slippage assumptions. Each grid point applies the specified slippage per order while keeping the strategy parameters, fees, position sizing, and evaluation period unchanged.
The analysis estimates how quickly candidate profitability decays as execution costs rise and identifies the first tested slippage level where a candidate becomes unprofitable.
This step is diagnostic and does not change the composite score.
Purpose: check whether a candidate's edge is large enough to survive less favorable fills.
This is an independent train/test procedure. For each train window, the full parameter grid is evaluated and the best parameters are selected according to a chosen selection metric. Those selected train parameters are then tested on the following unseen test window.
The analysis is repeated for multiple selection metrics so the report can compare which selection logic transfers best through time.
This step is independent from the main shortlist and does not change the composite score.
Purpose: test whether the strategy framework can be repeatedly re-optimized and still transfer to future windows.
The final page consolidates key counts, top candidates, warnings, benchmark checks, and high-level findings. It is a review and navigation layer rather than a separate research test.
Purpose: give a compact executive overview after all detailed step reports.