Appendix: Scores

This appendix documents score formulas used by the reports. Scores are ranking heuristics, not probabilities of future profit. For example, a score of 0.80 should not be read as an 80% probability of future success; it only ranks candidates inside this research process.

Score Usage Map

This table separates scores that drive the main ranking from diagnostic scores that are shown for inspection only.

Score / MetricUsed InAffects Final Ranking?RangeHow To Read It
efficiency_scoreStep 02, rolling/OOS/regime windows, WFA metric selectionNo, not directly0..1General single-period quality score; higher is better.
balanced_scoreStep 02 metric winners, WFA metric selectionNo0..1Geometric balance across return, drawdown, profit factor, and Sharpe.
profit_quality_scoreStep 02 metric winners, WFA metric selectionNoUnboundedDrawdown-adjusted absolute-profit heuristic; compare within one run only.
robustness_scoreStep 04, Step 06Yes0..1Combined rolling plus fixed-OOS robustness.
regimeStep 05, Step 06Yes0..1Weighted market-regime performance score.
parameter_stabilityStep 06Yes0..1Plateau / local-neighborhood stability score.
strategy_composite_scoreStep 06, Step 11 summaryMain final ranking0..1Final composite ranking score.
statistical_confidence_scoreStep 06 diagnosticsNo0..1Sample-quality diagnostic based on trades, Wilson win-rate bound, expectancy, and OOS positivity.
trade_distribution_scoreStep 06 diagnosticsNo0..1Trade-level distribution diagnostic for concentration, tails, median, skew, and sample size.
sltp_robustness_scoreStep 03/04 SL/TP validationNo, except for SL/TP shortlist inclusion0..1Checks whether an SL/TP overlay survives validation relative to its own baseline.
leverage_scoreStep 08No0..1Soft leverage suitability score; hard leverage status should be checked separately.
adaptability_scoreStep 10No0..1Classic train/test WFA adaptability diagnostic.

Common Scale Notes

Score Dependency Map

strategy_composite_score
  = robustness_score + regime + parameter_stability

robustness_score
  = rolling_robustness_score + oos_robustness_score

regime
  = weighted per_regime_score + positive_regime_ratio + covered_weight

parameter_stability
  = local-neighborhood quality + good-neighbor ratio + density + center position

diagnostic-only scores
  = statistical_confidence_score, trade_distribution_score, leverage_score, adaptability_score

Metric Winner / Parameter Optimization Scores

These scores are mostly used in Step 02 to find metric winners and diagnostic candidates. They rank candidates by different definitions of quality, not by robustness.

Efficiency Score

efficiency_score =
  0.30 * return_quality_score
+ 0.30 * risk_adjusted_score
+ 0.25 * drawdown_score
+ 0.10 * expectancy_score
+ 0.05 * trades_score

return_quality_score = 0.65 * profit_score + 0.35 * profit_factor_score
drawdown_score = 0.55 * drawdown_depth_score + 0.25 * drawdown_duration_score + 0.20 * underwater_score
positive_bounded_score(x) = max(x, 0) / (1 + max(x, 0))
signed_bounded_score(x) = 0.5 + 0.5 * x / (1 + abs(x))

Sub-scores:

Interpretation:

RangeMeaning
> 0.65Strong candidate on the tested period: good return/risk balance, acceptable drawdown, and enough trades.
0.45 - 0.65Usable/interesting but needs OOS, regime, and trade-distribution confirmation.
< 0.45Weak quality score; usually too low unless there is a specific reason to inspect it.

Balanced Score

balanced_score = geometric_mean(
  profit_score,
  drawdown_depth_score,
  drawdown_duration_score,
  profit_factor_score,
  sharpe_score
)

Sub-scores: this is a geometric mean, so one very weak component pulls the whole score down. It intentionally rewards candidates that are decent across return, drawdown, profit factor, and Sharpe rather than excellent in only one dimension. Balanced Score uses its own bounded component transforms; its profit_factor_score is not identical to the component used inside Efficiency Score.

Interpretation:

RangeMeaning
> 0.60Balanced across several dimensions.
0.35 - 0.60Acceptable but has at least one weaker dimension.
< 0.35Not balanced; one or more components are poor.

Profit Quality Score

profit_quality_score = profit_draw * log(profit_factor + 1) * log(min(10, expectancy) + 2) * log(1.2 + win_rate) * trade_penalty

profit_draw = profit_total_abs - (max_drawdown * profit_total_abs) * (1 - 0.055)
trade_penalty = max(1 - abs(trades - 50) / 50, 0.1) when trades < 50, otherwise 1

Sub-scores:

Interpretation:

RangeMeaning
> 0Passes this drawdown-adjusted profitability heuristic.
approx. 0Marginal; profit and drawdown adjustment roughly cancel out.
< 0Losses or drawdown adjustment dominate.

Caveat: this score is unbounded and capital-size dependent. It is useful for ranking candidates within the same run, but it should not be compared directly to bounded 0..1 scores.

Sharpe Metric

sharpe = mean(period_return) / std(period_return) * sqrt(bars_per_year)
bars_per_year = (365 * 24 * 3600) / median_bar_seconds  # derived from actual data interval

This is a raw risk-adjusted performance metric, not a bounded score.

Interpretation:

RangeMeaning
> 1.0Good risk-adjusted equity performance.
0.5 - 1.0Acceptable but not strong.
< 0.5Weak risk-adjusted performance.

Trade-Based Sharpe Metric

sharpe_trade_based = mean_period_pnl / std(individual_trade_returns) * sqrt(bars_per_year)

This is a raw trade-distribution metric, not a bounded score.

Interpretation:

RangeMeaning
HigherBetter consistency across individual trades.
Low / negativeTrade outcomes are noisy or negative after normalization.

Sortino Metric

sortino = mean(period_return) / std(negative_period_return) * sqrt(bars_per_year)

This is a raw downside-adjusted performance metric, not a bounded score.

Interpretation:

RangeMeaning
> 1.5Good downside-adjusted performance.
0.75 - 1.5Acceptable but not strong.
< 0.75Weak downside-adjusted performance.

Trade-Based Sortino Metric

sortino_trade_based = mean_period_pnl / std(losing_trade_returns) * sqrt(bars_per_year)

This is a raw trade-distribution metric, not a bounded score.

Interpretation:

RangeMeaning
HigherBetter consistency after focusing only on losing-trade volatility.
Low / negativeLosing trades are too volatile relative to daily PnL.

Max Drawdown Relative Score

max_drawdown_relative_score = profit_total / max_drawdown^2

This metric is unbounded. It is intentionally aggressive: very small drawdown can produce very large values.

Interpretation:

RangeMeaning
HigherBetter return per squared unit of drawdown.
Very highCan be good, but inspect whether max_drawdown is tiny due to too few trades.
<= 0No positive drawdown-adjusted return.

Robustness Related Scores

These scores evaluate whether a candidate remains acceptable across rolling windows and fixed OOS periods.

Combined Robustness Score

robustness_score = (0.50 * rolling_robustness_score + 0.50 * oos_robustness_score) / 1.00

Interpretation:

RangeMeaning
> 0.70Robust across both rolling and fixed OOS validation.
0.55 - 0.70Promising but needs detailed inspection.
0.40 - 0.55Borderline; usually not enough without strong qualitative reasons.
< 0.40Poor robustness.

Rolling Robustness Score

rolling_robustness_score =
  0.35 * mean_window_efficiency
+ 0.25 * worst_window_efficiency
+ 0.20 * positive_window_ratio
+ 0.10 * consistency_score
+ 0.10 * drawdown_stability_score

consistency_score = clip(1 - std(window_efficiency), 0, 1)
drawdown_stability_score = clip(1 - std(window_max_drawdown), 0, 1)

Sub-scores:

Interpretation:

RangeMeaning
> 0.70Strong rolling stability across windows.
0.50 - 0.70Moderate stability; inspect worst-window and equity behavior.
< 0.50Fragile across rolling windows or too dependent on specific periods.

OOS Robustness Score

oos_robustness_score =
  0.35 * mean_oos_efficiency
+ 0.25 * worst_oos_efficiency
+ 0.20 * positive_oos_ratio
+ 0.10 * oos_consistency_score
+ 0.10 * oos_drawdown_stability_score

Sub-scores:

Interpretation:

RangeMeaning
> 0.70Strong fixed-OOS confirmation.
0.50 - 0.70Mixed but potentially acceptable; inspect each OOS period separately.
< 0.50Weak OOS confirmation; likely period-sensitive.

Regime Related Scores

These scores evaluate how candidates behave across manually defined market regimes.

Regime Score

per_regime_score =
  0.35 * profit_score
+ 0.30 * survival_score
+ 0.25 * drawdown_score
+ 0.10 * efficiency_score

regime_score = 0.75 * weighted_regime_score + 0.15 * positive_regime_ratio + 0.10 * covered_weight

Sub-scores:

Interpretation:

RangeMeaning
> 0.70Works well across the configured regime mix.
0.50 - 0.70Acceptable but may have weak regimes.
< 0.50Regime-sensitive or poor in important regimes.

Parameter Stability Related Scores

These scores estimate whether a candidate sits in a stable parameter neighborhood or a sharp isolated optimum.

Parameter Stability / Plateau Score

score_column = robustness_score
radius = 2 grid steps
good_neighbor_threshold = quantile(score_column, 0.75)

parameter_stability =
  0.40 * neighbor_quality
+ 0.30 * good_neighbor_ratio
+ 0.20 * density_score
+ 0.10 * center_score

Sub-scores:

Fallback behavior: if a proper plateau grid is unavailable, the report falls back to a simpler parameter-center score. If parameter information is not usable, the neutral fallback is 0.5.

Interpretation:

RangeMeaning
> 0.70Candidate sits in a healthy plateau; nearby parameters are also good.
0.45 - 0.70Some stability, but not a clear plateau.
< 0.45Likely sharp optimum or edge of the tested parameter space.

Final Ranking Scores

These scores combine the previous validation layers into the final ranking used by the pipeline.

Strategy Composite Score

strategy_composite_score =
  0.60 * robustness_score
+ 0.25 * regime
+ 0.15 * parameter_stability

Interpretation:

RangeMeaning
> 0.70Strong overall finalist by this pipeline's ranking logic.
0.55 - 0.70Good candidate, but inspect components before trusting it.
0.40 - 0.55Borderline.
< 0.40Weak finalist.

Diagnostic Scores

These scores do not directly replace the main ranking logic. They help diagnose sample quality, trade distribution, and hidden risks.

Statistical Confidence Score

target_trades = 100
statistical_confidence_score =
  0.35 * trade_count_score
+ 0.25 * win_rate_wilson_lower_bound
+ 0.25 * expectancy_score
+ 0.15 * positive_oos_ratio

trade_count_score = total_trades / (target_trades + total_trades)

Sub-scores:

Interpretation:

RangeMeaning
> 0.65Reasonable sample support and OOS reliability.
0.45 - 0.65Some support, but sample may still be thin.
< 0.45Weak statistical confidence; be careful with conclusions.

Trade Distribution Score

trade_distribution_score =
  0.30 * concentration_score
+ 0.25 * median_score
+ 0.20 * tail_score
+ 0.15 * skew_score
+ 0.10 * sample_score

concentration_score = 1 - top_5_profit_share
tail_score penalizes 5% expected shortfall relative to mean absolute trade PnL

Sub-scores:

Interpretation:

RangeMeaning
> 0.65Healthy trade distribution; profit is less concentrated and tails are less problematic.
0.45 - 0.65Mixed distribution; inspect largest wins/losses.
< 0.45Profit may be concentrated or tail risk may dominate.

SL/TP Related Scores

These scores validate whether SL/TP modifications found on the research period survive on independent periods.

SL/TP Robustness Score

SL/TP robustness validates Step 03 SL/TP candidates over independent OOS and rolling periods. OOS periods: ['20250101-20250701', '20250701-20260101']; rolling timerange: 20220101-20250101.

sltp_robustness_score =
  0.45 * sltp_rolling_survival_ratio
+ 0.35 * sltp_oos_survival_ratio
+ 0.20 * worst_score

sltp_improved = true when SL/TP metric value beats the baseline metric value
worst_score = 1 if worst segment improvement >= 0 else 0

Sub-scores:

Caveat: this score only says whether the SL/TP overlay improves its own baseline on the selected target metric. A robust SL/TP improvement can still be attached to a weak baseline strategy, so inspect the resulting full-period and OOS metrics as well.

Interpretation:

RangeMeaning
> 0.70SL/TP improvement survives well across validation segments.
0.50 - 0.70Promising but mixed; inspect OOS and rolling details.
< 0.50SL/TP improvement is not robust enough.

Leverage Related Scores

These scores estimate whether a candidate has enough adverse-excursion buffer for the configured leverage assumptions.

Leverage Score

leverage_score =
  0.35 * mae_safety_score
+ 0.25 * leveraged_drawdown_score
+ 0.20 * tail_risk_score
+ 0.20 * robustness_under_leverage_score

Sub-scores:

Caveat: leverage_score is a soft ranking score. The hard leverage_status can still reject a candidate when liquidation, tail-loss, or drawdown guardrails fail.

Interpretation:

RangeMeaning
> 0.75Looks suitable for the tested leverage assumptions.
0.55 - 0.75Borderline; inspect MAE/tail-loss charts.
< 0.55Unsafe or weak leverage suitability.

Walk-Forward Related Scores

These scores belong to the classic train/test walk-forward step and measure whether optimized train parameters transfer to the next test period.

Walk-Forward Adaptability Score

adaptability_score = 0.60 * mean(score_decay_closeness) + 0.40 * positive_test_ratio
score_decay = test_decay_score / train_decay_score
score_decay_closeness = clip(1 - abs(1 - score_decay), 0, 1)
For profit_quality_score, score_decay uses signed-log transformed Profit Quality:
profit_quality_score_log = sign(x) * log(1 + abs(x)).

Sub-scores:

Caveat: this step evaluates the adaptability of the re-optimization process, not the final shortlist ranking. It does not change strategy_composite_score. A high value means train-selected parameters transferred well to the next test windows under the tested selection metric.

Interpretation:

RangeMeaning
> 0.75Good train-to-test adaptability.
0.60 - 0.75Acceptable.
0.45 - 0.60Weak / period-sensitive.
< 0.45Poor adaptability.
< 0.30Very poor adaptability.