This appendix documents score formulas used by the reports. Scores are ranking heuristics, not probabilities of future profit. For example, a score of 0.80 should not be read as an 80% probability of future success; it only ranks candidates inside this research process.
This table separates scores that drive the main ranking from diagnostic scores that are shown for inspection only.
| Score / Metric | Used In | Affects Final Ranking? | Range | How To Read It |
|---|---|---|---|---|
efficiency_score | Step 02, rolling/OOS/regime windows, WFA metric selection | No, not directly | 0..1 | General single-period quality score; higher is better. |
balanced_score | Step 02 metric winners, WFA metric selection | No | 0..1 | Geometric balance across return, drawdown, profit factor, and Sharpe. |
profit_quality_score | Step 02 metric winners, WFA metric selection | No | Unbounded | Drawdown-adjusted absolute-profit heuristic; compare within one run only. |
robustness_score | Step 04, Step 06 | Yes | 0..1 | Combined rolling plus fixed-OOS robustness. |
regime | Step 05, Step 06 | Yes | 0..1 | Weighted market-regime performance score. |
parameter_stability | Step 06 | Yes | 0..1 | Plateau / local-neighborhood stability score. |
strategy_composite_score | Step 06, Step 11 summary | Main final ranking | 0..1 | Final composite ranking score. |
statistical_confidence_score | Step 06 diagnostics | No | 0..1 | Sample-quality diagnostic based on trades, Wilson win-rate bound, expectancy, and OOS positivity. |
trade_distribution_score | Step 06 diagnostics | No | 0..1 | Trade-level distribution diagnostic for concentration, tails, median, skew, and sample size. |
sltp_robustness_score | Step 03/04 SL/TP validation | No, except for SL/TP shortlist inclusion | 0..1 | Checks whether an SL/TP overlay survives validation relative to its own baseline. |
leverage_score | Step 08 | No | 0..1 | Soft leverage suitability score; hard leverage status should be checked separately. |
adaptability_score | Step 10 | No | 0..1 | Classic train/test WFA adaptability diagnostic. |
0..1; higher is better.profit_quality_score, max_drawdown_relative_score, Sharpe, and Sortino are useful for ranking, but their absolute values are not directly comparable to bounded scores.0.70 are usually strong, 0.55-0.70 are promising, 0.40-0.55 are borderline, and values below 0.40 are weak. Some sections use stricter or looser thresholds where appropriate.strategy_composite_score = robustness_score + regime + parameter_stability robustness_score = rolling_robustness_score + oos_robustness_score regime = weighted per_regime_score + positive_regime_ratio + covered_weight parameter_stability = local-neighborhood quality + good-neighbor ratio + density + center position diagnostic-only scores = statistical_confidence_score, trade_distribution_score, leverage_score, adaptability_score
These scores are mostly used in Step 02 to find metric winners and diagnostic candidates. They rank candidates by different definitions of quality, not by robustness.
efficiency_score = 0.30 * return_quality_score + 0.30 * risk_adjusted_score + 0.25 * drawdown_score + 0.10 * expectancy_score + 0.05 * trades_score return_quality_score = 0.65 * profit_score + 0.35 * profit_factor_score drawdown_score = 0.55 * drawdown_depth_score + 0.25 * drawdown_duration_score + 0.20 * underwater_score positive_bounded_score(x) = max(x, 0) / (1 + max(x, 0)) signed_bounded_score(x) = 0.5 + 0.5 * x / (1 + abs(x))
Sub-scores:
profit_score: bounded version of total return; grows with positive return but saturates.profit_factor_score: bounded version of profit factor; rewards gross profit exceeding gross loss.risk_adjusted_score: signed bounded Sharpe-style component; negative Sharpe pushes it below 0.5.drawdown_depth_score: 1 - max_drawdown; deeper drawdown lowers the score.drawdown_duration_score: 1 - drawdown_duration_pct; long time underwater is penalized.underwater_score: 1 - underwater_mean; persistent underwater depth is penalized.expectancy_score: bounded positive expectancy; zero or negative expectancy contributes nothing.trades_score: bounded trade count support; very low trade count lowers confidence in the candidate.Interpretation:
| Range | Meaning |
|---|---|
| > 0.65 | Strong candidate on the tested period: good return/risk balance, acceptable drawdown, and enough trades. |
| 0.45 - 0.65 | Usable/interesting but needs OOS, regime, and trade-distribution confirmation. |
| < 0.45 | Weak quality score; usually too low unless there is a specific reason to inspect it. |
balanced_score = geometric_mean( profit_score, drawdown_depth_score, drawdown_duration_score, profit_factor_score, sharpe_score )
Sub-scores: this is a geometric mean, so one very weak component pulls the whole score down. It intentionally rewards candidates that are decent across return, drawdown, profit factor, and Sharpe rather than excellent in only one dimension. Balanced Score uses its own bounded component transforms; its profit_factor_score is not identical to the component used inside Efficiency Score.
Interpretation:
| Range | Meaning |
|---|---|
| > 0.60 | Balanced across several dimensions. |
| 0.35 - 0.60 | Acceptable but has at least one weaker dimension. |
| < 0.35 | Not balanced; one or more components are poor. |
profit_quality_score = profit_draw * log(profit_factor + 1) * log(min(10, expectancy) + 2) * log(1.2 + win_rate) * trade_penalty profit_draw = profit_total_abs - (max_drawdown * profit_total_abs) * (1 - 0.055) trade_penalty = max(1 - abs(trades - 50) / 50, 0.1) when trades < 50, otherwise 1
Sub-scores:
profit_draw: absolute profit reduced by a drawdown penalty.log(profit_factor + 1): rewards profit factor while compressing extreme values.log(min(10, expectancy) + 2): rewards expectancy but caps runaway values.log(1.2 + win_rate): small win-rate quality multiplier.trade_penalty: penalizes very low trade counts relative to the target of 50 trades.Interpretation:
| Range | Meaning |
|---|---|
| > 0 | Passes this drawdown-adjusted profitability heuristic. |
| approx. 0 | Marginal; profit and drawdown adjustment roughly cancel out. |
| < 0 | Losses or drawdown adjustment dominate. |
Caveat: this score is unbounded and capital-size dependent. It is useful for ranking candidates within the same run, but it should not be compared directly to bounded 0..1 scores.
sharpe = mean(period_return) / std(period_return) * sqrt(bars_per_year) bars_per_year = (365 * 24 * 3600) / median_bar_seconds # derived from actual data interval
This is a raw risk-adjusted performance metric, not a bounded score.
Interpretation:
| Range | Meaning |
|---|---|
| > 1.0 | Good risk-adjusted equity performance. |
| 0.5 - 1.0 | Acceptable but not strong. |
| < 0.5 | Weak risk-adjusted performance. |
sharpe_trade_based = mean_period_pnl / std(individual_trade_returns) * sqrt(bars_per_year)
This is a raw trade-distribution metric, not a bounded score.
Interpretation:
| Range | Meaning |
|---|---|
| Higher | Better consistency across individual trades. |
| Low / negative | Trade outcomes are noisy or negative after normalization. |
sortino = mean(period_return) / std(negative_period_return) * sqrt(bars_per_year)
This is a raw downside-adjusted performance metric, not a bounded score.
Interpretation:
| Range | Meaning |
|---|---|
| > 1.5 | Good downside-adjusted performance. |
| 0.75 - 1.5 | Acceptable but not strong. |
| < 0.75 | Weak downside-adjusted performance. |
sortino_trade_based = mean_period_pnl / std(losing_trade_returns) * sqrt(bars_per_year)
This is a raw trade-distribution metric, not a bounded score.
Interpretation:
| Range | Meaning |
|---|---|
| Higher | Better consistency after focusing only on losing-trade volatility. |
| Low / negative | Losing trades are too volatile relative to daily PnL. |
max_drawdown_relative_score = profit_total / max_drawdown^2
This metric is unbounded. It is intentionally aggressive: very small drawdown can produce very large values.
Interpretation:
| Range | Meaning |
|---|---|
| Higher | Better return per squared unit of drawdown. |
| Very high | Can be good, but inspect whether max_drawdown is tiny due to too few trades. |
| <= 0 | No positive drawdown-adjusted return. |
These scores evaluate whether a candidate remains acceptable across rolling windows and fixed OOS periods.
robustness_score = (0.50 * rolling_robustness_score + 0.50 * oos_robustness_score) / 1.00
Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | Robust across both rolling and fixed OOS validation. |
| 0.55 - 0.70 | Promising but needs detailed inspection. |
| 0.40 - 0.55 | Borderline; usually not enough without strong qualitative reasons. |
| < 0.40 | Poor robustness. |
rolling_robustness_score = 0.35 * mean_window_efficiency + 0.25 * worst_window_efficiency + 0.20 * positive_window_ratio + 0.10 * consistency_score + 0.10 * drawdown_stability_score consistency_score = clip(1 - std(window_efficiency), 0, 1) drawdown_stability_score = clip(1 - std(window_max_drawdown), 0, 1)
Sub-scores:
mean_window_efficiency: average quality across rolling windows.worst_window_efficiency: weakest rolling window; prevents one bad period from being hidden by the mean.positive_window_ratio: fraction of rolling windows with positive profit.consistency_score: rewards low dispersion of efficiency scores across windows.drawdown_stability_score: rewards stable drawdown behavior across windows.Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | Strong rolling stability across windows. |
| 0.50 - 0.70 | Moderate stability; inspect worst-window and equity behavior. |
| < 0.50 | Fragile across rolling windows or too dependent on specific periods. |
oos_robustness_score = 0.35 * mean_oos_efficiency + 0.25 * worst_oos_efficiency + 0.20 * positive_oos_ratio + 0.10 * oos_consistency_score + 0.10 * oos_drawdown_stability_score
Sub-scores:
mean_oos_efficiency: average quality across fixed OOS periods.worst_oos_efficiency: weakest fixed OOS period.positive_oos_ratio: fraction of OOS periods with positive profit.oos_consistency_score: rewards low dispersion of OOS efficiency scores.oos_drawdown_stability_score: rewards similar drawdown depth across OOS periods.Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | Strong fixed-OOS confirmation. |
| 0.50 - 0.70 | Mixed but potentially acceptable; inspect each OOS period separately. |
| < 0.50 | Weak OOS confirmation; likely period-sensitive. |
These scores evaluate how candidates behave across manually defined market regimes.
per_regime_score = 0.35 * profit_score + 0.30 * survival_score + 0.25 * drawdown_score + 0.10 * efficiency_score regime_score = 0.75 * weighted_regime_score + 0.15 * positive_regime_ratio + 0.10 * covered_weight
Sub-scores:
profit_score: bounded positive return inside one regime.survival_score: equals 1 for profitable regimes; negative returns reduce survival.drawdown_score: penalizes deep regime drawdowns.efficiency_score: carries the general quality score into the regime calculation.weighted_regime_score: average per-regime score using configured regime weights.positive_regime_ratio: fraction of regimes where profit is positive.covered_weight: how much configured regime weight had usable data.Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | Works well across the configured regime mix. |
| 0.50 - 0.70 | Acceptable but may have weak regimes. |
| < 0.50 | Regime-sensitive or poor in important regimes. |
These scores estimate whether a candidate sits in a stable parameter neighborhood or a sharp isolated optimum.
score_column = robustness_score radius = 2 grid steps good_neighbor_threshold = quantile(score_column, 0.75) parameter_stability = 0.40 * neighbor_quality + 0.30 * good_neighbor_ratio + 0.20 * density_score + 0.10 * center_score
Sub-scores:
neighbor_quality: average score of nearby parameter combinations.good_neighbor_ratio: share of nearby combinations above the configured score quantile.density_score: whether enough neighbors exist inside the configured radius.center_score: rewards candidates closer to the local neighborhood center instead of edge points.Fallback behavior: if a proper plateau grid is unavailable, the report falls back to a simpler parameter-center score. If parameter information is not usable, the neutral fallback is 0.5.
Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | Candidate sits in a healthy plateau; nearby parameters are also good. |
| 0.45 - 0.70 | Some stability, but not a clear plateau. |
| < 0.45 | Likely sharp optimum or edge of the tested parameter space. |
These scores combine the previous validation layers into the final ranking used by the pipeline.
strategy_composite_score = 0.60 * robustness_score + 0.25 * regime + 0.15 * parameter_stability
Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | Strong overall finalist by this pipeline's ranking logic. |
| 0.55 - 0.70 | Good candidate, but inspect components before trusting it. |
| 0.40 - 0.55 | Borderline. |
| < 0.40 | Weak finalist. |
These scores do not directly replace the main ranking logic. They help diagnose sample quality, trade distribution, and hidden risks.
target_trades = 100 statistical_confidence_score = 0.35 * trade_count_score + 0.25 * win_rate_wilson_lower_bound + 0.25 * expectancy_score + 0.15 * positive_oos_ratio trade_count_score = total_trades / (target_trades + total_trades)
Sub-scores:
trade_count_score: saturating sample-size support.win_rate_wilson_lower_bound: conservative lower confidence bound for observed win rate.expectancy_score: bounded positive weighted expectancy.positive_oos_ratio: fraction of OOS periods with positive return.Interpretation:
| Range | Meaning |
|---|---|
| > 0.65 | Reasonable sample support and OOS reliability. |
| 0.45 - 0.65 | Some support, but sample may still be thin. |
| < 0.45 | Weak statistical confidence; be careful with conclusions. |
trade_distribution_score = 0.30 * concentration_score + 0.25 * median_score + 0.20 * tail_score + 0.15 * skew_score + 0.10 * sample_score concentration_score = 1 - top_5_profit_share tail_score penalizes 5% expected shortfall relative to mean absolute trade PnL
Sub-scores:
concentration_score: penalizes profit made by only a few top trades.median_score: rewards positive median trade PnL.tail_score: penalizes bad expected shortfall in the worst 5% trades.skew_score: rewards positive skew and penalizes ugly left-tail distributions.sample_score: small trade samples are penalized.Interpretation:
| Range | Meaning |
|---|---|
| > 0.65 | Healthy trade distribution; profit is less concentrated and tails are less problematic. |
| 0.45 - 0.65 | Mixed distribution; inspect largest wins/losses. |
| < 0.45 | Profit may be concentrated or tail risk may dominate. |
These scores validate whether SL/TP modifications found on the research period survive on independent periods.
SL/TP robustness validates Step 03 SL/TP candidates over independent OOS and rolling periods. OOS periods: ['20250101-20250701', '20250701-20260101']; rolling timerange: 20220101-20250101.
sltp_robustness_score = 0.45 * sltp_rolling_survival_ratio + 0.35 * sltp_oos_survival_ratio + 0.20 * worst_score sltp_improved = true when SL/TP metric value beats the baseline metric value worst_score = 1 if worst segment improvement >= 0 else 0
Sub-scores:
sltp_rolling_survival_ratio: share of rolling validation windows where SL/TP improves its target metric.sltp_oos_survival_ratio: share of fixed OOS periods where SL/TP improves its target metric.worst_score: guardrail that fails when the worst segment is negative.Caveat: this score only says whether the SL/TP overlay improves its own baseline on the selected target metric. A robust SL/TP improvement can still be attached to a weak baseline strategy, so inspect the resulting full-period and OOS metrics as well.
Interpretation:
| Range | Meaning |
|---|---|
| > 0.70 | SL/TP improvement survives well across validation segments. |
| 0.50 - 0.70 | Promising but mixed; inspect OOS and rolling details. |
| < 0.50 | SL/TP improvement is not robust enough. |
These scores estimate whether a candidate has enough adverse-excursion buffer for the configured leverage assumptions.
leverage_score = 0.35 * mae_safety_score + 0.25 * leveraged_drawdown_score + 0.20 * tail_risk_score + 0.20 * robustness_under_leverage_score
Sub-scores:
mae_safety_score: liquidation buffer based on MAE quantile versus liquidation distance.leveraged_drawdown_score: drawdown safety at the tested leverage.tail_risk_score: worst-trade/tail-loss safety at the tested leverage.robustness_under_leverage_score: candidate robustness carried into leverage scoring.Caveat: leverage_score is a soft ranking score. The hard leverage_status can still reject a candidate when liquidation, tail-loss, or drawdown guardrails fail.
Interpretation:
| Range | Meaning |
|---|---|
| > 0.75 | Looks suitable for the tested leverage assumptions. |
| 0.55 - 0.75 | Borderline; inspect MAE/tail-loss charts. |
| < 0.55 | Unsafe or weak leverage suitability. |
These scores belong to the classic train/test walk-forward step and measure whether optimized train parameters transfer to the next test period.
adaptability_score = 0.60 * mean(score_decay_closeness) + 0.40 * positive_test_ratio score_decay = test_decay_score / train_decay_score score_decay_closeness = clip(1 - abs(1 - score_decay), 0, 1) For profit_quality_score, score_decay uses signed-log transformed Profit Quality: profit_quality_score_log = sign(x) * log(1 + abs(x)).
Sub-scores:
score_decay: train-to-test transfer ratio. Values near 1 mean good transfer.score_decay_closeness: bounded closeness to 1; both strong degradation and explosive improvement reduce confidence.positive_test_ratio: fraction of forward test windows with positive profit.test_score_worst: not part of adaptability_score directly, but should be inspected as worst-window failure risk.Caveat: this step evaluates the adaptability of the re-optimization process, not the final shortlist ranking. It does not change strategy_composite_score. A high value means train-selected parameters transferred well to the next test windows under the tested selection metric.
Interpretation:
| Range | Meaning |
|---|---|
| > 0.75 | Good train-to-test adaptability. |
| 0.60 - 0.75 | Acceptable. |
| 0.45 - 0.60 | Weak / period-sensitive. |
| < 0.45 | Poor adaptability. |
| < 0.30 | Very poor adaptability. |